Skip to main content
ElevenLabs Documentation Docs

Search documentation

Type to search this documentation.

On this pageOverview

Latency optimization

This guide covers the core principles for improving text-to-speech latency. For a conceptual explanation of what latency is and what contributes to it, see Understanding latency.

While there are many individual techniques, we'll group them into four principles.

  1. Use Flash models
  2. Leverage streaming
  3. Consider geographic proximity
  4. Choose appropriate voices

Flash models deliver ~75ms inference speeds, making them ideal for real-time applications. The trade-off is a slight reduction in audio quality compared to Multilingual v2.

There are three types of text-to-speech endpoints available in our API Reference:

  • Regular endpoint: Returns a complete audio file in a single response.
  • Streaming endpoint: Returns audio chunks progressively using Server-sent events.
  • Websockets endpoint: Enables bidirectional streaming for real-time audio generation.

Streaming endpoints progressively return audio as it is being generated in real-time, reducing the time-to-first-byte. This endpoint is recommended for cases where the input text is available up-front.

The text-to-speech websocket endpoint supports bidirectional streaming making it perfect for applications with real-time text input (e.g. LLM outputs).

If auto_mode is disabled, the model will wait for enough text to match the chunk schedule before starting to generate audio.

For instance, if you set a chunk schedule of 125 characters but only 50 arrive, the model stalls until additional characters come in—potentially increasing latency.

For implementation details, see the text-to-speech websocket guide.

We have observed that in some cases, voice selection can impact latency. Here's the order from fastest to slowest:

  1. Default voices (formerly premade), Synthetic voices, and Instant Voice Clones (IVC)
  2. Professional Voice Clones (PVC)

Higher audio quality output formats can increase latency. Be sure to balance your latency requirements with audio fidelity needs.

We serve our models from multiple regions to optimize latency based on your geographic location.

For example, using Flash models with Websockets, you can expect the following TTFB latencies depending on your location:

Region TTFB
North America 100-150ms
Europe 100-150ms
South East Asia 100-150ms
South Asia 150-200ms
North East Asia 150-200ms

You can check which backend region is serving your request by inspecting the x-region header in the API response. Currently used regions include: USA, Netherlands and Singapore.

To opt-out of the global routing and always use USA servers, use the api.us.elevenlabs.io base URL for your API requests:

Python
import os
from elevenlabs.client import ElevenLabs

elevenlabs = ElevenLabs(
    api_key=os.getenv("ELEVENLABS_API_KEY"),
    base_url="https://api.us.elevenlabs.io"
)
TypeScript
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import "dotenv/config";

const elevenlabs = new ElevenLabsClient({
  apiKey: process.env.ELEVENLABS_API_KEY,
  baseUrl: "https://api.us.elevenlabs.io",
});

cURL

cURL
curl -X POST -v "https://api.us.elevenlabs.io/v1/text-to-speech/{voice_id}" \
  -H "Accept: audio/mpeg" \
  -H "Content-Type: application/json" \
  -H "xi-api-key: YOUR_API_KEY" \
  -d '{
    "text": "Hello World!",
    "model_id": "eleven_flash_v2_5"
  }'
Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu