> ## Documentation Index
> Fetch the complete documentation index at: https://docs.eversince.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# transcribe media

> The words in audio and video, with timing. One ref or url: a timestamped transcript, sentence-level (word-level with include_words); media under five minutes answers in the call, longer returns a job and a second call after it finishes returns the words. Cached per library item, so a repeat read costs nothing; a url outside the workspace is read again every call. refs, up to 20: the words already stored on each file, never transcribing; a file nobody has read comes back unread. timeline_id: transcribes every uploaded clip on the timeline whose words are unknown and writes them onto the clips, which is what captions, get_transcript and remove_words read. format lines is the compact form for reading a long file whole; phrases marks the pauses, for choosing cut points. Cloud processing.



## OpenAPI

````yaml /openapi.json post /tools/transcribe_media
openapi: 3.1.0
info:
  title: Eversince API
  version: '1'
  description: >-
    Every Eversince tool as one POST call, plus the routes beside them: the tool
    list, account and keys, uploads, webhooks and models. The prose is at
    https://docs.eversince.ai.
servers:
  - url: https://eversince.ai/api/v1
security:
  - bearer: []
tags:
  - name: Discovery
    description: The tool list, the same on the MCP server, the REST API and the CLI.
  - name: Workspace
    description: Jobs, balances, settings, templates, skills and the overview.
  - name: Library
    description: >-
      Items, search, import, reading media, comments, boards, calendars, the
      brand kit and public links.
  - name: Timeline
    description: 'Video editing: timelines, clips, captions, sound, language and rendering.'
  - name: Canvas
    description: 'Still editing: canvases, slides, layers and rendering.'
  - name: Generation
    description: Image, video and audio generation, upscaling, cutouts and models.
  - name: Research
    description: 'What platforms publish: pulling and keeping posts.'
  - name: Account
    description: Balances, keys, workspaces and script sessions.
  - name: Uploads
    description: Files into the library by presigned upload or by URL.
  - name: Webhooks
    description: Job results posted to a URL as they complete.
  - name: Models
    description: Generation models and cost estimates.
paths:
  /tools/transcribe_media:
    post:
      tags:
        - Library
      summary: transcribe media
      description: >-
        The words in audio and video, with timing. One ref or url: a timestamped
        transcript, sentence-level (word-level with include_words); media under
        five minutes answers in the call, longer returns a job and a second call
        after it finishes returns the words. Cached per library item, so a
        repeat read costs nothing; a url outside the workspace is read again
        every call. refs, up to 20: the words already stored on each file, never
        transcribing; a file nobody has read comes back unread. timeline_id:
        transcribes every uploaded clip on the timeline whose words are unknown
        and writes them onto the clips, which is what captions, get_transcript
        and remove_words read. format lines is the compact form for reading a
        long file whole; phrases marks the pauses, for choosing cut points.
        Cloud processing.
      operationId: transcribe_media
      parameters:
        - $ref: '#/components/parameters/IdempotencyKey'
        - $ref: '#/components/parameters/WorkspaceId'
      requestBody:
        required: true
        content:
          application/json:
            schema:
              type: object
              properties:
                ref:
                  type: string
                  description: An item ref, item id, or the job id that made it.
                url:
                  type: string
                  format: uri
                refs:
                  description: >-
                    Up to 20 audio or video items: their stored words, free. One
                    that cannot be read comes back as its own row beside the
                    rest.
                  minItems: 1
                  maxItems: 20
                  type: array
                  items:
                    type: string
                timeline_id:
                  description: >-
                    Transcribe the uploaded clips on this timeline that have no
                    words yet, and write the words onto them.
                  type: string
                clip_ids:
                  description: >-
                    With timeline_id: only these clips, read again from scratch
                    even if they have words, the way back from a transcript
                    speech recognition got wrong.
                  type: array
                  items:
                    type: string
                include_words:
                  description: 'ref or url: word-level timing as well as sentences.'
                  type: boolean
                format:
                  description: >-
                    sentences (default): each with start and end. lines: one
                    timed line per sentence as one string, half the size, for
                    reading a long file whole. phrases: speech broken at pauses,
                    each with the gap after it, for choosing cut points.
                  type: string
                  enum:
                    - sentences
                    - lines
                    - phrases
                start_time:
                  description: >-
                    ref or url: return only from this second. The whole file is
                    still read.
                  type: number
                  minimum: 0
                end_time:
                  description: 'ref or url: return only up to this second.'
                  type: number
                  minimum: 0
                min_gap_ms:
                  description: >-
                    format phrases: a pause this long or longer breaks one
                    phrase from the next. Default 400; 700 and up finds only the
                    breaths between takes.
                  type: integer
                  minimum: 0
                  maximum: 10000
                speakers:
                  description: >-
                    format phrases: name who is talking on each phrase, from the
                    voices the speakers tool picked out. A file nobody has been
                    picked out of has none; this never listens to find out.
                  type: boolean
                window:
                  description: >-
                    format phrases: only phrases touching this span of the file,
                    in seconds.
                  type: object
                  properties:
                    start:
                      type: number
                      minimum: 0
                    end:
                      type: number
                      minimum: 0
                  required:
                    - start
                    - end
                context:
                  description: 'Optional: one sentence shown to the user beside this action.'
                  type: string
                  maxLength: 600
      responses:
        '200':
          description: >-
            Done, or a job started for work that runs longer (follow it with
            get_jobs).
          content:
            application/json:
              schema:
                type: object
                required:
                  - ok
                  - data
                  - meta
                properties:
                  ok:
                    const: true
                  data:
                    type: object
                    additionalProperties: true
                  meta:
                    type: object
                    properties:
                      request_id:
                        type: string
                      cost:
                        $ref: '#/components/schemas/Cost'
                      cloud_processing:
                        $ref: '#/components/schemas/CloudProcessing'
                  media:
                    type: array
                    items:
                      type: object
                      additionalProperties: true
        default:
          $ref: '#/components/responses/Error'
components:
  parameters:
    IdempotencyKey:
      name: Idempotency-Key
      in: header
      required: false
      schema:
        type: string
      description: >-
        Retry a paid dispatch safely: the same key returns the job already
        dispatched instead of charging again.
    WorkspaceId:
      name: x-workspace-id
      in: header
      required: false
      schema:
        type: string
      description: The workspace this one call acts in, when it is not the key's own.
  schemas:
    Cost:
      type: object
      description: AI credits the call cost.
      additionalProperties: true
    CloudProcessing:
      type: object
      description: Cloud processing the call used, and what is left.
      properties:
        ran_on:
          const: cloud
        minutes:
          type: number
        queued_minutes:
          type: number
        left_minutes:
          type: number
        bought_minutes:
          type: number
        note:
          type: string
    Error:
      type: object
      required:
        - ok
        - error
      properties:
        ok:
          const: false
        error:
          type: object
          required:
            - code
            - message
            - retryable
          properties:
            code:
              type: string
            message:
              type: string
              description: What went wrong, in words to act on.
            retryable:
              type: boolean
            suggested_action:
              type: string
              description: The next step that works.
            cloud_processing:
              $ref: '#/components/schemas/CloudProcessing'
  responses:
    Error:
      description: The call was refused or failed.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
  securitySchemes:
    bearer:
      type: http
      scheme: bearer
      description: >-
        An API key (es_live_…) from Settings, under API keys, or an OAuth access
        token.

````