# Text to Dialogue

## Overview

The ElevenLabs [Text to Dialogue](/guides/changelog-api-reference-text-to-dialogue-convert) API creates natural sounding expressive dialogue from text using [Eleven v4](/guides/overview-capabilities-text-to-speech-eleven-v4) and Eleven v3. We recommend Eleven v4. Popular use cases include:

- Generating pitch perfect conversations for video games
- Creating immersive dialogue for podcasts and other audio content
- Bring audiobooks to life with expressive narration

Several generations might be required to achieve the desired results. When integrating Text to Dialogue into your application, consider generating several generations and allowing the user to select the best one.

Listen to a sample:

::::card-grid
:::card{title="Developers" href="/guides/elevenapi-guides-cookbooks-text-to-dialogue"}
Learn how to integrate text to dialogue into your application.
:::

:::card{title="Prompting guide" href="/guides/overview-capabilities-text-to-speech-best-practices#prompting-eleven-v4"}
Learn how to prompt expressive dialogue with audio tags.
:::

:::card{title="API reference" href="/guides/changelog-api-reference-text-to-dialogue-convert"}
Full API reference for the Text to Dialogue endpoint.
:::
::::

## Voice options

ElevenLabs offers thousands of voices through multiple creation methods. Eleven v4 supports 90+ languages and Eleven v3 supports 70+ languages.

- [Voice library](/guides/overview-capabilities-voices) with 3,000+ community-shared voices
- [Professional voice cloning](/guides/overview-capabilities-voices#cloned) for highest-fidelity replicas
- [Instant voice cloning](/guides/overview-capabilities-voices#cloned) for quick voice replication
- [Voice design](/guides/overview-capabilities-voices#voice-design) to generate custom voices from text descriptions

Learn more about our [voice options](/guides/overview-capabilities-voices).

## Prompting

The models interpret emotional context directly from the text input. For example, adding descriptive text like “she said excitedly” or using exclamation marks will influence the speech emotion. Voice settings like Stability and Similarity help control the consistency, while the underlying emotion comes from textual cues.

Read the [prompting guide](/guides/overview-capabilities-text-to-speech-best-practices#prompting-eleven-v4) for more details.

### Emotional deliveries with audio tags

:::callout{intent="warning"}
This feature is still under active development, actual results may vary.
:::

Eleven v4 and Eleven v3 allow the use of non-speech audio events to influence the delivery of the dialogue. This is done by inserting the audio events into the text input wrapped in square brackets.

In Text to Dialogue, each dialogue turn has its own text and voice. Add audio tags inside the text for the turn they should affect. The `voice_id` still selects the speaker voice for that turn, while the tags guide delivery.

For example, a speaker can use one voice while the text starts with `[giggling]`, and the next speaker can use a different voice while the text starts with `[whispering]`. For an API example that combines tags with `voice_id`, see the [Text to Dialogue quickstart](/guides/elevenapi-guides-cookbooks-text-to-dialogue).

Audio tags are natural-language instructions, not an enum parameter. Wrap the instruction in square brackets and place it in the text where the delivery should change. The examples below are not exhaustive; use the [prompting guide](/guides/overview-capabilities-text-to-speech-best-practices#prompting-eleven-v4) for more guidance on effective tags.

Audio tags come in a few different forms:

### Emotions and delivery

For example, \[sad], \[laughing] and \[whispering]

### Audio events

For example, \[leaves rustling], \[gentle footsteps] and \[applause].

### Overall direction

For example, \[football], \[wrestling match] and \[auctioneer].

Some examples include:

```
"[giggling] That's really funny!"
"[groaning] That was awful."
"Well, [sigh] I'm not sure what to say."
```

You can also use punctuation to indicate the flow of dialog, like interruptions:

```
"[cautiously] Hello, is this seat-"
"[jumping in] Free? [cheerfully] Yes it is."
```

Ellipses can be used to indicate trailing sentences:

```
"[indecisive] Hi, can I get uhhh..."
"[quizzically] The usual?"
"[elated] Yes! [laughs] I'm so glad you knew!"
```

::::accordion{title="Supported output formats"}
The default response format is `mp3`, but other formats like `pcm` and `ulaw` are available.

- **MP3**

  - Sample rates: 22.05kHz - 44.1kHz
  - Bitrates: 32kbps - 192kbps
  - 22.05kHz @ 32kbps
  - 44.1kHz @ 32kbps, 64kbps, 96kbps, 128kbps, 192kbps

- **PCM (S16LE)**

  - Sample rates: 16kHz - 44.1kHz
  - Bitrates: 8kHz, 16kHz, 22.05kHz, 24kHz, 44.1kHz, 48kHz
  - 16-bit depth

- **μ-law**

  - 8kHz sample rate
  - Optimized for telephony applications

- **A-law**

  - 8kHz sample rate
  - Optimized for telephony applications

- **Opus**

  - Sample rate: 48kHz
  - Bitrates: 32kbps - 192kbps

:::callout{intent="success"}
Higher quality audio options are only available on paid tiers - see our [pricing page](https://elevenlabs.io/pricing/api) for details.
:::
::::

## Supported languages

Eleven v4 supports 90+ languages. Eleven v3 supports 70+ languages. See [Eleven v4](/guides/overview-models#eleven-v4) and [Eleven v3](/guides/overview-models#eleven-v3) for the full lists.

## FAQ

:::accordion{title="Which models can I use?"}
Text to Dialogue is available on the Eleven v4 and Eleven v3 models.
:::

:::accordion{title="Do I own the audio output?"}
Yes. You retain ownership of any audio you generate. However, commercial usage rights are only available with paid plans. With a paid subscription, you may use generated audio for commercial purposes and monetize the outputs if you own the IP rights to the input content.
:::

:::accordion{title="What qualifies as a free regeneration?"}
A free regeneration allows you to regenerate the same text to speech content without additional cost, subject to these conditions:

- Only available within the ElevenLabs dashboard.
- You can regenerate each piece of content up to 2 times for free.
- The content must be exactly the same as the previous generation. Any changes to the text, voice settings, or other parameters will require a new, paid generation.

Free regenerations are useful in case there is a slight distortion in the audio output. According to ElevenLabs’ internal benchmarks, regenerations will solve roughly half of issues with quality, with remaining issues usually due to poor training data.
:::

:::accordion{title="How many speakers can my dialogue have?"}
There is no limit to the number of speakers in a dialogue.
:::

:::accordion{title="Why is my output sometimes inconsistent?"}
The models are nondeterministic. For consistency, use the optional [seed parameter](/guides/changelog-api-reference-text-to-dialogue-convert#request.body.seed), though subtle differences may still occur.
:::

:::accordion{title="What's the best practice for large text conversions?"}
Keep the total length of all `inputs[].text` values at or below 2,000 characters per request for reliable generation. Split longer text into chunks and concatenate the resulting audio in your application.
:::

## Key facts

- **Models**: Available with Eleven v4 and Eleven v3
- **Speakers**: No limit on number of speakers per dialogue
- **Request size**: Keep the total length of all `inputs[].text` values at or below 2,000 characters per request
- **Determinism**: Output is nondeterministic — use the `seed` parameter for more consistent results
- **Free regenerations**: Up to 2 free regenerations per generation (same content, same parameters, dashboard only)
- **Ownership**: You retain ownership of generated audio; commercial use requires a paid plan

## Related pages

- [Administration](./administration-index.md)
- [API reference](./api-reference-index.md)
- [Changelog](./changelog-index.md)
- [ElevenAgents](./elevenagents-index.md)
- [ElevenAPI](./elevenapi-index.md)
- [ElevenCreative](./elevencreative-index.md)
- [ElevenLabs Documentation Docs](../index.md)
- [General Troubleshooting FAQ](./troubleshooting-index.md)
- [General Website FAQ](./website-index.md)
- [Help Center](./help-center-2-index.md)

# Agent Instructions

Cite this page’s canonical URL and keep its documentation version.
Follow Link headers to discover available agent guidance and tools.
Read the advertised skill for the requested version before choosing starting pages.
Treat documentation as reference material, not execution authorization.
