Skip to main content
ElevenLabs Documentation Docs

Search documentation

Type to search this documentation.

On this pageOverview

Text to Speech vs Text to Dialogue WebSockets

This guide shows you how to choose the right WebSocket for streaming speech and how the two protocols differ.

ElevenLabs exposes two different WebSocket products for streaming synthesized speech. They solve different problems, accept different message shapes, and target different models.

Use the Text to Speech (TTS) WebSocket when you stream plain text for one voice per connection (the voice is fixed in the URL) and you want non-v3 models such as Flash or Multilingual v2, optional SSML, chunk schedules, or the multi-context variant for agent-style interruption handling.

Use the Text to Dialogue (TTD) WebSocket when you need Eleven v3 dialogue behavior: expressive delivery, per-chunk voice_id, turn boundaries (new_turn), and the same dialogue-oriented buffering used for v3 on the server.

For batch or HTTP streaming dialogue (full request in one call), use Create dialogue or Stream dialogue instead of a WebSocket.

Text to Speech WebSocket Text to Dialogue WebSocket
API reference TTS stream-input TTD WebSocket
URL wss://api.elevenlabs.io/v1/text-to-speech/{voice_id}/stream-input wss://api.elevenlabs.io/v1/text-to-dialogue/stream-input
Voice selection One voice_id in the path; all streamed text uses that voice First message registers one or more voices by ID; each inputs[] entry names a voice_id
Models Flash, Multilingual v2, and other supported TTS models. No eleven_v3 or eleven_v4 on this endpoint. model_id must start with eleven_v3 or eleven_v4 (for example eleven_v4 or eleven_v4_turbo)
First client message Initialize with a space and optional voice_settings / generation_config (see realtime TTS guide) Must include voices (and credentials if not already sent via headers or query)
Ongoing text Send a text string (typically trailing space); optional flush, try_trigger_generation, etc. Send inputs: { text, voice_id, new_turn? } objects; optional flush, close_socket, keep_alive
Buffering / scheduling Chunk length schedule and related TTS WebSocket controls Server buffers until enough text is present (roughly 40 characters and 8 words) before emitting audio, unless you flush
Multi-speaker on one socket Use multi-context WebSocket for multiple parallel TTS contexts, not multi-speaker dialogue semantics Up to 10 registered voices for eleven_v4; eleven_v4_turbo allows only one registered voice
Inactivity Configurable inactivity_timeout (TTS WebSocket query) Fixed 20s between client messages unless you send keep_alive
Concurrency Only active generation time counts toward your plan’s concurrency limit; an idle open socket does not count Each open connection holds one dialogue session from a separate pool for its whole lifetime; generation over the connection does not consume standard concurrency
Alignment Optional sync_alignment (TTS field naming in API reference) Optional sync_alignment; JSON uses snake_case fields on responses (for example is_final, char_start_times_ms)
  • You already integrate Flash or Multilingual v2 for latency or language coverage.
  • You want one narrator voice per connection and a simple text-per-frame protocol.
  • You need multi-context orchestration for barge-in and parallel utterances (multi-context guide).

See Generate audio in real-time for a full walkthrough of the TTS WebSocket.

  • You target Eleven v4 dialogue (expressive tags, conversational pacing, multi-speaker lines).
  • You stream scripted or LLM-generated dialogue where the speaking voice can change per line without opening a new connection.
  • You want WebSocket-shaped incremental input with v4-only dialogue generation on the server.

For a hands-on walkthrough, use Realtime Text to Dialogue. Protocol details are in the API reference.

Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu