ElevenAPI
Breaking changes policy
Zero Retention Mode (Enterprise)
Zero Retention Mode (Enterprise)
IP allowlisting
Webhooks
Agent Tooling
Tools for agents to build with ElevenLabs.
Errors
Python SDK reference
JavaScript SDK reference
React SDK
JavaScript SDK
Libraries & SDKs
Secure by design
Latency optimization
This guide shows you how to reduce text-to-speech latency in your application.
Pipecat integration
Use a Pipecat pipeline as the LLM brain behind Speech Engine.
LiveKit integration
Bridge a LiveKit room into Speech Engine using a LiveKit Agents worker.
Multi-Context Websocket
Text to Speech vs Text to Dialogue WebSockets
This guide shows you how to choose the right WebSocket for streaming speech and how the two protocols differ.
Stream dialogue in real-time
Stream dialogue in real-time
Generate audio in real-time
Voice Remixing quickstart
This guide shows you how to remix an existing voice to create a new one.
Voice Design quickstart
This guide shows you how to design a voice via a text prompt using the Voice Design API.
Professional Voice Cloning quickstart
Instant Voice Cloning quickstart
This guide shows you how to clone a voice using the Clone Voice API.
Bring your own transcript
Create a dubbing project from your own transcript and supply your own translations.
Manage dubbing projects
Dub into multiple languages
Refine and regenerate a dub
Image & Video webhooks
Image & Video webhooks
References and assets
Music inpainting
Composition plans
Precise control over music generation with structured JSON
Music streaming
This guide shows you how to stream music with our Music API.
Realtime event reference
Transcript editing
This guide shows you how to apply natural-language edit instructions to committed transcripts with the Realtime Speech to Text API.
Transcripts and commit strategies
This guide shows you how to handle transcripts and commit strategies with the ElevenLabs Realtime Speech to Text API.
Server-side streaming
Server-side streaming
Client-side streaming
Client-side streaming
Vercel AI SDK
Use the ElevenLabs Provider in the Vercel AI SDK to transcribe speech from audio and video files.
Transcription Telegram Bot
Build a Telegram bot that transcribes audio and video messages in 90+languages using TypeScript with Deno in Supabase Edge Functions.
Transcript editing
This guide shows you how to edit transcripts with natural-language instructions using the Speech to Text API.
Entity detection
This guide shows you how to use entity detection with the Speech to Text API.
Keyterm prompting
This guide shows you how to use keyterm prompting with the Speech to Text API.
Asynchronous Speech to Text
This guide shows you how to use webhooks to receive asynchronous notifications when transcription tasks complete.
Multichannel speech-to-text
Multichannel speech-to-text
Sending generated audio through Twilio
Streaming and Caching with Supabase
Generate and stream speech through Supabase Edge Functions. Store speech in Supabase Storage and cache responses via built-in CDN.
Using pronunciation dictionaries
This guide shows you how to manage pronunciation dictionaries programmatically.
Stitching multiple requests
This guide shows you how to maintain voice prosody across multiple text chunks/generations.
Streaming text to speech
Voice cloning: how it works
Voice cloning: how it works
Understanding latency
What latency means in audio generation, what contributes to it, and how to reason about tradeoffs.
Understanding audio streaming
Why streaming audio generation is different from streaming files, and what that means for your application.
Forced Alignment quickstart
Learn how to use the Forced Alignment API to align text to audio.
Sound Effects quickstart
Learn how to generate sound effects using the Sound Effects API.
Dubbing quickstart
Learn how to dub audio and video files across languages using the Dubbing API.
Voice Isolator quickstart
Learn how to remove background noise from an audio file using the Voice Isolator API.
Voice Changer quickstart
Learn how to transform the voice of an audio file using the Voice Changer API.
Image & Video quickstart
Text to Dialogue quickstart
Learn how to generate immersive dialogue from text.
Music quickstart
Speech Engine quickstart
Speech to Text quickstart
Learn how to convert spoken audio into text.
How to choose the right model
This guide shows you how to choose the right ElevenLabs model for your use case.