Get Speech Engine
GET https://api.elevenlabs.io/v1/speech-engine/
Retrieve a Speech Engine resource
Reference: https://elevenlabs.io/docs/api-reference/speech-engine/get
Servers
Section titled “Servers”https://api.elevenlabs.io(Production, default)https://api.us.elevenlabs.io(Production US)https://api.eu.residency.elevenlabs.io(Production EU)https://api.in.residency.elevenlabs.io(Production India)https://api.sg.residency.elevenlabs.io(Production Singapore)
Request
Section titled “Request”Path parameters
Section titled “Path parameters”speech_engine_id(string, required) — The speech engine ID (accepts seng_ or agent_ prefix)
Response
Section titled “Response”Successful Response
speech_engine_id(string, required) — The speech engine resource IDname(string, required) — Human-readable name for the speech enginespeech_engine(SpeechEngineConfig, required) — WebSocket connection settings for the upstream transcript serverasr(ASRConversationalConfig, required) — Automatic speech recognition configurationtts(TTSConversationalConfig-Output, required) — Text-to-speech output configurationturn(BaseTurnConfig, required) — Turn detection configurationvad(VADConfig, required) — Configuration for voice activity detectionconversation(ConversationConfig-Output, required) — Conversation-level settings including client events and duration limitsprivacy(PrivacyConfig-Output, required) — Privacy settings controlling recording, retention, and PII handlingcall_limits(AgentCallLimits, required) — Concurrency and daily conversation limits for this speech enginelanguage(string, required) — ISO language code used by the speech engine (e.g. 'en')cascade_timeout_seconds(double, required) — Time in seconds to wait for the upstream speech engine endpoint to respond before the attempt is abandoned and retried. Must be between 2 and 15 seconds.tags(list of string, required) — Arbitrary tags for categorization and filteringoverrides(SpeechEngineConversationInitiationClientDataConfig, required) — Override settings the client may set during conversation initiationmetadata(AgentMetadataDBModel, required) — Creation and update timestamps with source informationaccess_info(ResourceAccessInfo, optional, nullable) — The access information of the speech engine for the user
Errors
Section titled “Errors”422 Unprocessable Entity Error
Section titled “422 Unprocessable Entity Error”Validation Error
detail(list of ValidationError, optional)
SpeechEngineConfig
Section titled “SpeechEngineConfig”ws_url(string, required) — The WebSocket URL for the transcript serverrequest_headers(map from string to SpeechEngineConfigRequestHeaders, optional) — Headers to include in the WebSocket connection request
ASRConversationalConfig
Section titled “ASRConversationalConfig”quality(enum, optional, default: high) — The quality of the transcription- Allowed values:
high
- Allowed values:
provider(enum, optional, default: scribe_realtime) — The provider of the transcription service- Allowed values:
elevenlabs,scribe_realtime
- Allowed values:
user_input_audio_format(enum, optional, default: pcm_16000) — The format of the audio to be transcribed- Allowed values:
pcm_8000,pcm_16000,pcm_22050,pcm_24000,pcm_44100,pcm_48000,ulaw_8000
- Allowed values:
keywords(list of string, optional) — Keywords to boost prediction probability for
TTSConversationalConfig-Output
Section titled “TTSConversationalConfig-Output”model_id(enum, optional, default: eleven_flash_v2) — The model to use for TTS- Allowed values:
eleven_turbo_v2,eleven_turbo_v2_5,eleven_flash_v2,eleven_flash_v2_5,eleven_multilingual_v2,eleven_v3_conversational
- Allowed values:
voice_id(string, optional, default: cjVigY5qzO86Huf0OWal) — The voice ID to use for TTSsupported_voices(list of SupportedVoice, optional) — Additional supported voices for the agentexpressive_mode(boolean, optional, default: true) — When enabled, applies expressive audio tags prompt. Automatically disabled for non-v3 models.suggested_audio_tags(list of SuggestedAudioTag, optional) — Suggested audio tags to boost expressive speech (for eleven_v3 and eleven_v3_conversational models). The agent can still use other tags not listed here.agent_output_audio_format(enum, optional, default: pcm_16000) — The audio format to use for TTS- Allowed values:
pcm_8000,pcm_16000,pcm_22050,pcm_24000,pcm_44100,pcm_48000,ulaw_8000
- Allowed values:
stability(double, optional, default: 0.5) — The stability of generated speechspeed(double, optional, default: 1) — The speed of generated speechsimilarity_boost(double, optional, default: 0.8) — The similarity boost for generated speechtext_normalisation_type(enum, optional, default: system_prompt) — Method for converting numbers to words before converting text to speech. If set to SYSTEM_PROMPT, the system prompt will be updated to include normalization instructions. If set to ELEVENLABS, the text will be normalized after generation, incurring slight additional latency.- Allowed values:
system_prompt,elevenlabs
- Allowed values:
pronunciation_dictionary_locators(list of PydanticPronunciationDictionaryVersionLocator, optional) — The pronunciation dictionary locatorsenable_phoneme_tags(boolean, optional, default: true) — Opt-in to SSML phoneme tag handling for V3 models. When enabled, phoneme tags (inline and from pronunciation dictionaries) are parsed into inline IPA before being sent to the model.audio_effects(EffectsSpec-Output, optional, nullable) — Optional TTS effects spec: filter preset, distance (proximity EQ), and environment (convolution reverb).optimize_streaming_latency(enum, optional, deprecated) — Deprecated: this field is a no-op and is ignored.- Allowed values:
0,1,2,3,4
- Allowed values:
BaseTurnConfig
Section titled “BaseTurnConfig”turn_timeout(double, optional, default: 7) — Maximum wait time for the user's reply before re-engaging the userinitial_wait_time(double, optional, nullable) — How long the agent will wait for the user to start the conversation if the first message is empty. If not set, uses the regular turn_timeout.silence_end_call_timeout(double, optional, default: -1) — Maximum wait time since the user last spoke before terminating the callturn_eagerness(enum, optional, default: normal) — Controls how eager the agent is to respond. Low = less eager (waits longer), Standard = default eagerness, High = more eager (responds sooner)- Allowed values:
patient,normal,eager
- Allowed values:
spelling_patience(enum, optional, default: auto) — Controls if the agent should be more patient when user is spelling numbers and named entities. Auto = model based, Off = never wait extra- Allowed values:
auto,off
- Allowed values:
speculative_turn(boolean, optional, default: false) — When enabled, starts generating LLM responses during silence before full turn confidence is reached, reducing perceived latency. May increase LLM costs.retranscribe_on_turn_timeout(boolean, optional, default: false) — When enabled, if VAD detects no speech, attempts to re-transcribe accumulated audio at turn timeout. Disables silence discount billing for affected turns.turn_model(enum, optional, default: turn_v3) — Version of the turn detection model to use.- Allowed values:
turn_v2,turn_v3
- Allowed values:
interruption_ignore_terms(list of string, optional) — List of terms that should not trigger an interruption when spoken by the user (e.g. 'gotcha', 'understood'). Uses case-insensitive exact matching.interruption_ignore_term_languages(list of string, optional) — Language codes for which preset ignore-term categories have been activated. Stored explicitly so display is not inferred from term overlap.merge_with_default_ignore_terms(boolean, optional, default: false) — When enabled, the curated default terms for interruption_ignore_term_languages are used in addition to interruption_ignore_terms.transcribe_on_disabled_interruptions(boolean, optional, default: false) — When interruptions are disabled, still transcribe what the user says so it can carry into the next turn. When off, user speech during a non-interruptible turn is ignored and won't trigger a turn.
VADConfig
Section titled “VADConfig”ConversationConfig-Output
Section titled “ConversationConfig-Output”text_only(boolean, optional, default: false) — If enabled audio will not be processed and only text will be used, use to avoid audio pricing.max_duration_seconds(integer, optional, default: 600) — The maximum duration of a conversation in secondsclient_events(list of enum, optional) — The events that will be sent to the client- Allowed values:
conversation_initiation_metadata,asr_initiation_metadata,ping,audio,interruption,user_transcript,tentative_user_transcript,agent_response,agent_response_correction,client_tool_call,mcp_tool_call,mcp_connection_status,agent_tool_request,agent_tool_response,agent_tool_response_full_payload,agent_response_metadata,vad_score,agent_chat_response_part,client_error,guardrail_triggered,dtmf_request,agent_response_complete,context_usage,internal_turn_probability,internal_tentative_agent_response
- Allowed values:
file_input(FileInputConfig, optional) — Configuration for file input (image/PDF uploads) during conversations.monitoring_enabled(boolean, optional, default: false) — Enable real-time monitoring of conversations via WebSocketmonitoring_events(list of enum, optional) — The events that will be sent to monitoring connections.- Allowed values:
conversation_initiation_metadata,asr_initiation_metadata,ping,audio,interruption,user_transcript,tentative_user_transcript,agent_response,agent_response_correction,client_tool_call,mcp_tool_call,mcp_connection_status,agent_tool_request,agent_tool_response,agent_tool_response_full_payload,agent_response_metadata,vad_score,agent_chat_response_part,client_error,guardrail_triggered,dtmf_request,agent_response_complete,context_usage,internal_turn_probability,internal_tentative_agent_response
- Allowed values:
dtmf_input_settings(DTMFInputConfig, optional, nullable) — Configure DTMF (keypad) input collection during phone callsbackground_sound(BackgroundSoundConfig, optional) — Configuration for background sound during conversations.source_attribution(boolean, optional, default: false) — When enabled and knowledge base content is present, the LLM is instructed to report which sources it used.
PrivacyConfig-Output
Section titled “PrivacyConfig-Output”record_voice(boolean, optional, default: true) — Whether to record the conversationretention_days(integer, optional, default: -1) — The number of days to retain the conversation. -1 indicates there is no retention limitdelete_transcript_and_pii(boolean, optional, default: false) — Whether to delete the transcript and PIIdelete_audio(boolean, optional, default: false) — Whether to delete the audioapply_to_existing_conversations(boolean, optional, default: false) — Whether to apply the privacy settings to existing conversationszero_retention_mode(boolean, optional, default: false) — Whether to enable zero retention mode - no PII data is storedconversation_history_redaction(ConversationHistoryRedactionConfig, optional) — Config for PII redaction in the conversation history
AgentCallLimits
Section titled “AgentCallLimits”agent_concurrency_limit(integer, optional, default: -1) — The maximum number of concurrent conversations. -1 indicates that there is no maximumdaily_limit(integer, optional, default: 100000) — The maximum number of conversations per daybursting_enabled(boolean, optional, default: true) — Whether to enable bursting. If true, exceeding workspace concurrency limit will be allowed up to 3 times the limit. Calls will be charged at double rate when exceeding the limit.
SpeechEngineConversationInitiationClientDataConfig
Section titled “SpeechEngineConversationInitiationClientDataConfig”first_message(boolean, optional, default: false) — Whether the first message can be overridden by the client
AgentMetadataDBModel
Section titled “AgentMetadataDBModel”created_at_unix_secs(integer, required)updated_at_unix_secs(integer, required)created_from(enum, optional, default: unknown)- Allowed values:
cli,ui,api,template,unknown
- Allowed values:
last_updated_from(enum, optional, default: unknown)- Allowed values:
cli,ui,api,template,unknown
- Allowed values:
ResourceAccessInfo
Section titled “ResourceAccessInfo”is_creator(boolean, required) — Whether the user making the request is the creator of the agentcreator_name(string, required) — Name of the agent's creatorcreator_email(string, required) — Email of the agent's creatorrole(enum, required) — The role of the user making the request- Allowed values:
admin,editor,commenter,viewer
- Allowed values:
anonymous_access_level_override(enum, optional, nullable) — The access level for anonymous users. If None, the resource is not shared publicly.- Allowed values:
admin,editor,commenter,viewer
- Allowed values:
access_source(enum, optional, nullable) — Why the requesting user has access to this resource. 'creator' = caller is the owner. 'explicit' = caller (or one of their workspace groups) is listed in role_to_group_ids beyond the workspace-wide everyone group. 'workspace_default' = the workspace-wide everyone group is listed in role_to_group_ids (every non-anon workspace member, including admins, sees this resource). 'workspace_admin' = caller is a workspace admin and the admin seat is the only path to access; reserved for docs nobody else can see. Lets the UI disclose why an admin-bypass viewer sees a doc that wasn't explicitly shared with them.- Allowed values:
creator,explicit,workspace_admin,workspace_default
- Allowed values:
ValidationError
Section titled “ValidationError”loc(list of ValidationErrorLocItems, required)msg(string, required)type(string, required)
SpeechEngineConfigRequestHeaders
Section titled “SpeechEngineConfigRequestHeaders”SupportedVoice
Section titled “SupportedVoice”label(string, required)voice_id(string, required)description(string, optional, nullable)language(string, optional, nullable)model_family(enum, optional, nullable)- Allowed values:
turbo,flash,multilingual,v3_conversational
- Allowed values:
optimize_streaming_latency(enum, optional, nullable)- Allowed values:
0,1,2,3,4
- Allowed values:
stability(double, optional, nullable)speed(double, optional, nullable)similarity_boost(double, optional, nullable)
SuggestedAudioTag
Section titled “SuggestedAudioTag”tag(string, required) — Audio tag to use (for best performance, 1-2 words, e.g., 'happy', 'excited')description(string, optional, nullable) — Optional description of when to use this tag
PydanticPronunciationDictionaryVersionLocator
Section titled “PydanticPronunciationDictionaryVersionLocator”A locator for other documents to be able to reference a specific dictionary and it's version. This is a pydantic version of PronunciationDictionaryVersionLocatorDBModel. Required to ensure compat with the rest of the agent data models.
pronunciation_dictionary_id(string, required) — The ID of the pronunciation dictionaryversion_id(string, required, nullable) — The ID of the version of the pronunciation dictionary
EffectsSpec-Output
Section titled “EffectsSpec-Output”Filter preset, distance (proximity EQ), and environment (convolution reverb).
filter_preset_id(string, required, nullable)distance(double, required, default: 0)environment_id(string, required, nullable)background_noise_id(string, required, nullable)send_level(double, required, default: 1)seed(integer, required, nullable)
FileInputConfig
Section titled “FileInputConfig”enabled(boolean, optional, default: true) — When enabled, users may attach images or PDFs in chat when the LLM supports multimodal input.max_files_in_memory(integer, optional, default: 10) — Number of most-recent files kept in memory during a conversation. Older files are summarized and their bytes freed.max_files_per_conversation(integer, optional, default: 10) — Total files a user can upload in one conversation. Uploads are billed per file. Use -1 for no limit, or a value >= max_files_in_memory.
DTMFInputConfig
Section titled “DTMFInputConfig”Configuration for DTMF (keypad) input collection during phone calls.
dtmf_input_timeout(double, optional, default: 2) — Timeout in seconds to wait for additional DTMF digitshash_terminator(boolean, optional, default: true) — If true, pressing # immediately completes DTMF inputredact_input(boolean, optional, default: false) — If true, replace the caller's DTMF (keypad) entries with a redaction marker in the transcript, conversation log and analysis. Digits the agent repeats back or passes to a tool are not affected.
BackgroundSoundConfig
Section titled “BackgroundSoundConfig”source_type(enum, optional, nullable) — The type of background sound source.- Allowed values:
preset
- Allowed values:
source_id(enum, optional, nullable) — Identifier for the sound source.- Allowed values:
office2,office1,restaurant,city,typing,elevator1,elevator2,elevator3,elevator4
- Allowed values:
volume(double, optional, default: 0.15) — Volume level for background sound (0.01 to 1.0).crossfade_loop(boolean, optional, default: true) — Apply a crossfade at the loop boundary to avoid audible pops when the sound loops.
ConversationHistoryRedactionConfig
Section titled “ConversationHistoryRedactionConfig”enabled(boolean, optional, default: false) — Whether conversation history redaction is enabledentities(list of enum, optional) — The entities to redact from the conversation transcript, audio and analysis. Use top-level types like 'name', 'email_address', or dot notation for specific subtypes like 'name.full_name'.- Allowed values:
name,name.name_given,name.name_family,name.name_other,email_address,contact_number,dob,age,religious_belief,political_opinion,sexual_orientation,ethnicity_race,marital_status,occupation,physical_attribute,language,username,password,url,organization,financial_id,financial_id.payment_card,financial_id.payment_card.payment_card_number,financial_id.payment_card.payment_card_expiration_date,financial_id.payment_card.payment_card_cvv,financial_id.bank_account,financial_id.bank_account.bank_account_number,financial_id.bank_account.bank_routing_number,financial_id.bank_account.swift_bic_code,financial_id.financial_id_other,location,location.location_address,location.location_city,location.location_postal_code,location.location_coordinate,location.location_state,location.location_country,location.location_other,date,date_interval,unique_id,unique_id.government_issued_id,unique_id.account_number,unique_id.vehicle_id,unique_id.healthcare_number,unique_id.healthcare_number.medical_record_number,unique_id.healthcare_number.health_plan_beneficiary_number,unique_id.device_id,unique_id.unique_id_other,medical,medical.medical_condition,medical.medication,medical.medical_procedure,medical.medical_measurement,medical.medical_other
- Allowed values:
ValidationErrorLocItems
Section titled “ValidationErrorLocItems”ConvAISecretLocator
Section titled “ConvAISecretLocator”Used to reference a secret from the agent's secret store.
secret_id(string, required)
ConvAIDynamicVariable
Section titled “ConvAIDynamicVariable”Used to reference a dynamic variable.
variable_name(string, required)
Examples
Section titled “Examples”Response
{
"speech_engine_id": "seng_3701k3ttaq12ewp8b7qv5rfyszkz",
"name": "My Speech Engine",
"speech_engine": {
"ws_url": "wss://example.com/transcript",
"request_headers": {}
},
"asr": {
"quality": "high",
"provider": "elevenlabs",
"user_input_audio_format": "pcm_16000",
"keywords": []
},
"tts": {
"model_id": "eleven_flash_v2",
"voice_id": "cjVigY5qzO86Huf0OWal",
"agent_output_audio_format": "pcm_16000",
"stability": 0.5,
"speed": 1,
"similarity_boost": 0.8,
"optimize_streaming_latency": 3
},
"turn": {
"turn_timeout": 7,
"silence_end_call_timeout": -1,
"turn_eagerness": "normal",
"mode": "turn"
},
"vad": {
"background_voice_detection": false
},
"conversation": {
"max_duration_seconds": 600,
"client_events": [
"audio",
"interruption",
"agent_response",
"user_transcript"
]
},
"privacy": {
"record_voice": true,
"retention_days": -1,
"delete_transcript_and_pii": false,
"delete_audio": false,
"apply_to_existing_conversations": false,
"zero_retention_mode": false
},
"call_limits": {
"agent_concurrency_limit": -1,
"daily_limit": 100000,
"bursting_enabled": true
},
"language": "en",
"cascade_timeout_seconds": 4,
"tags": [
"production",
"v1"
],
"overrides": {
"first_message": false
},
"metadata": {
"created_at_unix_secs": 1714000000,
"updated_at_unix_secs": 1714000000,
"created_from": "api",
"last_updated_from": "api"
}
}SDK Code
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
async function main() {
const client = new ElevenLabsClient();
await client.speechEngine.get("seng_3701k3ttaq12ewp8b7qv5rfyszkz");
}
main();
from elevenlabs import ElevenLabs
client = ElevenLabs()
client.speech_engine.get(
speech_engine_id="seng_3701k3ttaq12ewp8b7qv5rfyszkz",
)
package main
import (
"fmt"
"net/http"
"io"
)
func main() {
url := "https://api.elevenlabs.io/v1/speech-engine/seng_3701k3ttaq12ewp8b7qv5rfyszkz"
req, _ := http.NewRequest("GET", url, nil)
res, _ := http.DefaultClient.Do(req)
defer res.Body.Close()
body, _ := io.ReadAll(res.Body)
fmt.Println(res)
fmt.Println(string(body))
}require 'uri'
require 'net/http'
url = URI("https://api.elevenlabs.io/v1/speech-engine/seng_3701k3ttaq12ewp8b7qv5rfyszkz")
http = Net::HTTP.new(url.host, url.port)
http.use_ssl = true
request = Net::HTTP::Get.new(url)
response = http.request(request)
puts response.read_bodyimport com.mashape.unirest.http.HttpResponse;
import com.mashape.unirest.http.Unirest;
HttpResponse<String> response = Unirest.get("https://api.elevenlabs.io/v1/speech-engine/seng_3701k3ttaq12ewp8b7qv5rfyszkz")
.asString();<?php
require_once('vendor/autoload.php');
$client = new \GuzzleHttp\Client();
$response = $client->request('GET', 'https://api.elevenlabs.io/v1/speech-engine/seng_3701k3ttaq12ewp8b7qv5rfyszkz');
echo $response->getBody();using RestSharp;
var client = new RestClient("https://api.elevenlabs.io/v1/speech-engine/seng_3701k3ttaq12ewp8b7qv5rfyszkz");
var request = new RestRequest(Method.GET);
IRestResponse response = client.Execute(request);import Foundation
let request = NSMutableURLRequest(url: NSURL(string: "https://api.elevenlabs.io/v1/speech-engine/seng_3701k3ttaq12ewp8b7qv5rfyszkz")! as URL,
cachePolicy: .useProtocolCachePolicy,
timeoutInterval: 10.0)
request.httpMethod = "GET"
let session = URLSession.shared
let dataTask = session.dataTask(with: request as URLRequest, completionHandler: { (data, response, error) -> Void in
if (error != nil) {
print(error as Any)
} else {
let httpResponse = response as? HTTPURLResponse
print(httpResponse)
}
})
dataTask.resume()