Skip to main content
ElevenLabs Documentation Docs
current

Search documentation

Type to search this documentation.

On this pageOverview

Text to Dialogue quickstart

Learn how to generate immersive dialogue from text.

This guide will show you how to generate immersive, natural-sounding dialogue from text using the Text to Dialogue API.

  1. Create an API key

    Create an API key in the dashboard here, which you’ll use to securely access the API.

    Store the key as a managed secret and pass it to the SDKs either as a environment variable via an .env file, or directly in your app’s configuration depending on your preference.

    .env

    JavaScript
    ELEVENLABS_API_KEY=<your_api_key_here>
  2. Install the SDK

    We’ll also use the dotenv library to load our API key from an environment variable.

    Python
    pip install elevenlabs
    pip install python-dotenv

    Install the ElevenLabs CLI. Homebrew (macOS) and Scoop (Windows) are recommended.

    Homebrew (macOS)

    Homebrew (macOS)
    brew install elevenlabs/tap/elevenlabs

    Scoop (Windows)

    Scoop (Windows)
    scoop bucket add elevenlabs https://github.com/elevenlabs/scoop-bucket
    scoop install elevenlabs

    npm

    npm
    npm install -g @elevenlabs/cli

    curl

    curl
    curl --proto '=https' --tlsv1.2 -LsSf https://github.com/elevenlabs/cli/releases/latest/download/elevenlabs-cli-installer.sh | sh

    Working with an AI coding assistant? Run elevenlabs generate-skills in your project to write a SKILL.md for every command group into skills/, so your assistant knows the CLI's full surface without you pasting docs. Use --output-dir to put them elsewhere. This reads the CLI's own embedded API definition, so it needs no API key and works offline — and it stays in step with whichever CLI version you have installed.

    Then authenticate — this opens your browser to authorize the CLI:

    Bash
    elevenlabs auth login
  3. Make the API request

    Create a new file named example.py or example.mts, depending on your language of choice, and add the following code. Add audio tags inside each text value to guide that speaker’s delivery. The voice_id selects the speaker voice for the same input item.

    Python
    # example.py
    import os
    
    from dotenv import load_dotenv
    from elevenlabs.client import ElevenLabs
    from elevenlabs.play import play
    
    load_dotenv()
    
    elevenlabs = ElevenLabs(
      api_key=os.getenv("ELEVENLABS_API_KEY"),
    )
    
    audio = elevenlabs.text_to_dialogue.convert(
        inputs=[
            {
                "text": "[cheerfully] Hello, how are you?",
                "voice_id": "9BWtsMINqrJLrRacOk9x",
            },
            {
                "text": "[stuttering] I'm... I'm doing well, thank you.",
                "voice_id": "IKne3meq5aSn9XLyUdCD",
            }
        ]
    )
    
    play(audio)

    Then run it:

    Python
    python example.py

    You should hear the dialogue audio play.

    Pass the dialogue inputs as a JSON array and save the audio to a file:

    Bash
    elevenlabs text-to-dialogue convert \
      --output dialogue.mp3 \
      --inputs '[
        { "text": "[cheerfully] Hello, how are you?", "voice_id": "9BWtsMINqrJLrRacOk9x" },
        { "text": "[stuttering] I... I am doing well, thank you.", "voice_id": "IKne3meq5aSn9XLyUdCD" }
      ]'

    Open dialogue.mp3 to hear the result.

For incremental dialogue over a long-lived connection with Eleven v3 models, follow Realtime Text to Dialogue. To compare this WebSocket with the standard TTS WebSocket, see Text to Speech vs Text to Dialogue WebSockets.

Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu