Skip to main content
ElevenLabs Documentation Docs
current

Search documentation

Type to search this documentation.

ElevenAPI

Breaking changes policy

Zero Retention Mode (Enterprise)

Zero Retention Mode (Enterprise)

IP allowlisting

Webhooks

Agent Tooling

Tools for agents to build with ElevenLabs.

Errors

Python SDK reference

JavaScript SDK reference

React SDK

JavaScript SDK

Libraries & SDKs

Secure by design

Latency optimization

This guide shows you how to reduce text-to-speech latency in your application.

Pipecat integration

Use a Pipecat pipeline as the LLM brain behind Speech Engine.

LiveKit integration

Bridge a LiveKit room into Speech Engine using a LiveKit Agents worker.

Multi-Context Websocket

Text to Speech vs Text to Dialogue WebSockets

This guide shows you how to choose the right WebSocket for streaming speech and how the two protocols differ.

Stream dialogue in real-time

Stream dialogue in real-time

Generate audio in real-time

Voice Remixing quickstart

This guide shows you how to remix an existing voice to create a new one.

Voice Design quickstart

This guide shows you how to design a voice via a text prompt using the Voice Design API.

Professional Voice Cloning quickstart

Instant Voice Cloning quickstart

This guide shows you how to clone a voice using the Clone Voice API.

Bring your own transcript

Create a dubbing project from your own transcript and supply your own translations.

Manage dubbing projects

Dub into multiple languages

Refine and regenerate a dub

Image & Video webhooks

Image & Video webhooks

References and assets

Music inpainting

Composition plans

Precise control over music generation with structured JSON

Music streaming

This guide shows you how to stream music with our Music API.

Realtime event reference

Transcript editing

This guide shows you how to apply natural-language edit instructions to committed transcripts with the Realtime Speech to Text API.

Transcripts and commit strategies

This guide shows you how to handle transcripts and commit strategies with the ElevenLabs Realtime Speech to Text API.

Server-side streaming

Server-side streaming

Client-side streaming

Client-side streaming

Vercel AI SDK

Use the ElevenLabs Provider in the Vercel AI SDK to transcribe speech from audio and video files.

Transcription Telegram Bot

Build a Telegram bot that transcribes audio and video messages in 90+languages using TypeScript with Deno in Supabase Edge Functions.

Transcript editing

This guide shows you how to edit transcripts with natural-language instructions using the Speech to Text API.

Entity detection

This guide shows you how to use entity detection with the Speech to Text API.

Keyterm prompting

This guide shows you how to use keyterm prompting with the Speech to Text API.

Asynchronous Speech to Text

This guide shows you how to use webhooks to receive asynchronous notifications when transcription tasks complete.

Multichannel speech-to-text

Multichannel speech-to-text

Sending generated audio through Twilio

Streaming and Caching with Supabase

Generate and stream speech through Supabase Edge Functions. Store speech in Supabase Storage and cache responses via built-in CDN.

Using pronunciation dictionaries

This guide shows you how to manage pronunciation dictionaries programmatically.

Stitching multiple requests

This guide shows you how to maintain voice prosody across multiple text chunks/generations.

Streaming text to speech

Voice cloning: how it works

Voice cloning: how it works

Understanding latency

What latency means in audio generation, what contributes to it, and how to reason about tradeoffs.

Understanding audio streaming

Why streaming audio generation is different from streaming files, and what that means for your application.

Forced Alignment quickstart

Learn how to use the Forced Alignment API to align text to audio.

Sound Effects quickstart

Learn how to generate sound effects using the Sound Effects API.

Dubbing quickstart

Learn how to dub audio and video files across languages using the Dubbing API.

Voice Isolator quickstart

Learn how to remove background noise from an audio file using the Voice Isolator API.

Voice Changer quickstart

Learn how to transform the voice of an audio file using the Voice Changer API.

Image & Video quickstart

Text to Dialogue quickstart

Learn how to generate immersive dialogue from text.

Music quickstart

Speech Engine quickstart

Speech to Text quickstart

Learn how to convert spoken audio into text.

How to choose the right model

This guide shows you how to choose the right ElevenLabs model for your use case.

Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu