> ## Documentation Index
> Fetch the complete documentation index at: https://docs.opper.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Generate audio

> Generate audio from a text prompt with any audio generation model from `GET /v3/audio/models`: music (`type=music`) or sound effects and ambience (`type=sound`). By default runs synchronously and returns it inline as base64 (200). Set `async: true` to run it on the background worker and get a 202 with a status URL to poll instead; use it for long tracks, which can take minutes. `model` is required, and `prompt` too unless `sections` are sent; `duration_seconds`, `instrumental`, `format` (omit it for the model's native format) and `seed` are normalized, and a value a model cannot honor answers 400 naming the limit. `sections` (a timed plan: each section's `duration_ms`, `prompt` and `styles`) needs a model with the `audio_sections` capability, and `loop` one with `audio_loop`; a model without the capability answers 400 rather than ignoring the field. Everything in `parameters` is forwarded verbatim to the provider. Lyrics are returned when the provider produces them. Set `store: true` to also save the output to /v3/files and get a reusable `file_id`; nothing is stored otherwise.

Music, songs, sound effects and ambience. See the [Audio guide](/build/multimodal/audio#music-and-sound-effects) for runnable examples, timed sections and loops.


## OpenAPI

````yaml post /v3/audio/generations
openapi: 3.1.0
info:
  description: Schema-driven generative API that orchestrates LLM-powered workflows.
  title: Task API
  version: 3.0.0
servers:
  - description: Production
    url: https://api.opper.ai
  - description: Local development
    url: http://localhost:8080
security:
  - BearerAuth: []
tags:
  - description: Schema-driven function management and execution
    name: Functions
  - description: OpenAI-compatible chat completions
    name: Chat
  - description: OpenAI Responses API compatible endpoint
    name: Responses
  - description: Google-compatible interactions endpoint
    name: Interactions
  - description: Model registry and capabilities
    name: Models
  - description: Synchronous image generation
    name: Images
  - description: Text-to-speech and speech-to-text
    name: Audio
  - description: Asynchronous video generation
    name: Videos
  - description: Reusable file storage for media inputs and generated outputs
    name: Files
  - description: Async generation status and downloads
    name: Artifacts
  - description: OpenAI-compatible embeddings
    name: Embeddings
  - description: Recorded HTTP request/response generations
    name: Generations
  - description: System health and status
    name: System
  - description: Roundtable endpoint — fan out a query to multiple LLMs and combine results
    name: Roundtable
  - description: Web search, fetch, and other utility tools
    name: Tools
  - description: Caller identity, credits, and usage
    name: Account
  - description: >-
      Programmatic project and API-key management. Authenticates with an
      `op-mak-…` management token; available on the control_plane and enterprise
      plans.
    name: Management
paths:
  /v3/audio/generations:
    post:
      tags:
        - Audio
      summary: Generate audio
      description: >-
        Generate audio from a text prompt with any audio generation model from
        `GET /v3/audio/models`: music (`type=music`) or sound effects and
        ambience (`type=sound`). By default runs synchronously and returns it
        inline as base64 (200). Set `async: true` to run it on the background
        worker and get a 202 with a status URL to poll instead; use it for long
        tracks, which can take minutes. `model` is required, and `prompt` too
        unless `sections` are sent; `duration_seconds`, `instrumental`, `format`
        (omit it for the model's native format) and `seed` are normalized, and a
        value a model cannot honor answers 400 naming the limit. `sections` (a
        timed plan: each section's `duration_ms`, `prompt` and `styles`) needs a
        model with the `audio_sections` capability, and `loop` one with
        `audio_loop`; a model without the capability answers 400 rather than
        ignoring the field. Everything in `parameters` is forwarded verbatim to
        the provider. Lyrics are returned when the provider produces them. Set
        `store: true` to also save the output to /v3/files and get a reusable
        `file_id`; nothing is stored otherwise.
      operationId: createAudioGeneration
      requestBody:
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/AudioGenerationRequest'
        required: true
      responses:
        '200':
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/AudioGenerationResponse'
          description: Successful response
        '202':
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/AudioGenerationJobResponse'
          description: >-
            Accepted: async generation submitted (async: true). Poll the
            returned `status_url` (GET /v3/artifacts/{id}/status) for a
            presigned download URL when complete.
        '400':
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
          description: Bad request
        '401':
          description: Unauthorized - missing or invalid API key
        '500':
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
          description: Internal server error
components:
  schemas:
    AudioGenerationRequest:
      properties:
        async:
          description: >-
            Run generation on the background worker and return a job to poll
            (202) instead of waiting inline (200). Use for full-length tracks
            and slow models.
          type: boolean
        duration_seconds:
          description: >-
            Requested length in seconds (up to 600). Omit it to let the model
            choose. Limits differ by model, and some models have no length
            setting; a value a model cannot honor answers 400 naming its limit.
          type: number
        format:
          description: >-
            Output audio format: mp3, wav, opus or pcm. Omit it to get the
            model's native format. A format the model cannot produce answers 400
            naming the ones it can.
          type: string
        instrumental:
          description: 'Music only: generate without vocals.'
          type: boolean
        loop:
          description: >-
            Generate audio that repeats seamlessly, for ambience and hums. Needs
            a model with the audio_loop capability (GET
            /v3/audio/models?capability=audio_loop); other models answer 400.
          type: boolean
        model:
          description: >-
            Audio generation model id. List them with GET /v3/audio/models:
            ?type=music for songs and scores, ?type=sound for sound effects and
            ambience, ?capability=audio_sections or ?capability=audio_loop for
            models that take those fields. For example elevenlabs/music_v2_5
            (music) or elevenlabs/eleven_text_to_sound_v2 (sound effects).
          type: string
        parameters:
          description: >-
            Provider-specific fields, forwarded verbatim and never gated: the
            provider validates them. Use it for settings specific to one model,
            in that provider's own field names.
          type: object
        prompt:
          description: >-
            What to generate. For music: genre, mood, instruments, lyrics. For a
            sound effect: the sound itself, for example "a single hammer strike
            on an iron anvil". Required unless sections are set.
          type: string
        sections:
          description: >-
            A timed plan for music that must fit a picture: each section's
            duration_ms plus its own prompt and styles, in order. The sections
            set the total length, so duration_seconds is not sent with them.
            Needs a model with the audio_sections capability (GET
            /v3/audio/models?capability=audio_sections); other models answer
            400.
          items:
            properties:
              duration_ms:
                description: >-
                  Length of this section in milliseconds. Models set their own
                  per-section limits and answer 400 outside them.
                type: integer
              lyrics:
                description: >-
                  Words to be sung in this section. Leave it empty for an
                  instrumental section.
                type: string
              negative_styles:
                description: Styles this section should avoid, for example ["vocals"].
                items:
                  type: string
                type: array
              prompt:
                description: >-
                  What this section should sound like, in words, for example "a
                  quiet piano intro" or "strings build to a swell". It describes
                  the music and is never sung; put sung words in lyrics.
                type: string
              styles:
                description: >-
                  Styles this section should have, for example ["sparse piano",
                  "instrumental"].
                items:
                  type: string
                type: array
            required:
              - duration_ms
            type: object
          type: array
        seed:
          description: >-
            Seed for more repeatable output. Models without seed support answer
            400.
          type: integer
        store:
          description: >-
            Persist the generated audio to /v3/files and return a reusable
            file_id. Defaults to false; set true to store.
          type: boolean
      required:
        - model
      type: object
    AudioGenerationResponse:
      properties:
        audio:
          properties:
            b64_json:
              type: string
            file_id:
              type: string
            mime_type:
              type: string
            url:
              type: string
          type: object
        created:
          type: integer
        id:
          type: string
        lyrics:
          type: string
        model:
          type: string
        usage:
          properties:
            cost:
              type: number
            duration_seconds:
              type: number
          required:
            - cost
          type: object
      required:
        - id
        - model
        - created
        - audio
        - usage
      type: object
    AudioGenerationJobResponse:
      properties:
        id:
          type: string
        status_url:
          type: string
      required:
        - id
        - status_url
      type: object
    ErrorResponse:
      properties:
        error:
          properties:
            code:
              type: string
            details:
              description: Any value
            message:
              type: string
          required:
            - code
            - message
          type: object
        meta:
          type: object
      required:
        - error
      type: object
  securitySchemes:
    BearerAuth:
      bearerFormat: API Key
      description: API key authentication. Pass your API key as a Bearer token.
      scheme: bearer
      type: http

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.