JavaScript SDK
Installation
Section titled “Installation”npm install @elevenlabs/client
# or
yarn add @elevenlabs/client
# or
pnpm install @elevenlabs/clientHere is a minimal working example that connects to Scribe and logs transcription results:
import { Scribe, RealtimeEvents } from "@elevenlabs/client";
const token = await fetchTokenFromServer();
const connection = Scribe.connect({
token,
modelId: "scribe_v2_realtime",
microphone: {
echoCancellation: true,
noiseSuppression: true,
},
});
connection.on(RealtimeEvents.PARTIAL_TRANSCRIPT, (data) => {
console.log("Partial:", data.text);
});
connection.on(RealtimeEvents.COMMITTED_TRANSCRIPT, (data) => {
console.log("Committed:", data.text);
});
// Later, close the connection
connection.close();Getting a token
Section titled “Getting a token”Scribe requires a single-use token for authentication. Create an API endpoint on your server:
// Node.js server
app.get("/scribe-token", yourAuthMiddleware, async (req, res) => {
const response = await fetch("https://api.elevenlabs.io/v1/single-use-token/realtime_scribe", {
method: "POST",
headers: {
"xi-api-key": process.env.ELEVENLABS_API_KEY,
},
});
const data = await response.json();
res.json({ token: data.token });
});// Client
const fetchToken = async () => {
const response = await fetch("/scribe-token");
const { token } = await response.json();
return token;
};Connection options
Section titled “Connection options”Scribe.connect() accepts either microphone options or manual audio options. Both share a common set of base options.
Base options
Section titled “Base options”| Property | Type | Default | Description |
|---|---|---|---|
| token | string |
Single-use token for WebSocket authentication. | |
| modelId | string |
Model ID (e.g., "scribe_v2_realtime"). |
|
| baseUri | string |
"wss://api.elevenlabs.io" |
Custom WebSocket base URI. |
| commitStrategy | CommitStrategy |
"manual" |
"manual" or "vad". |
| vadSilenceThresholdSecs | number |
1.5 |
Seconds of silence before VAD commits (0.3-3.0). |
| vadThreshold | number |
0.4 |
VAD sensitivity (0.1-0.9, lower is more sensitive). |
| minSpeechDurationMs | number |
100 |
Minimum speech duration in ms (50-2000). |
| minSilenceDurationMs | number |
100 |
Minimum silence duration in ms (50-2000). |
| languageCode | string |
ISO-639-1 or ISO-639-3 language code. Leave empty for auto-detection. | |
| includeTimestamps | boolean |
false |
Receive word-level timestamps via the COMMITTED_TRANSCRIPT_WITH_TIMESTAMPS event. |
Microphone options
Section titled “Microphone options”Pass a microphone object to stream audio directly from the user's microphone. The connection handles getUserMedia and audio encoding automatically.
const connection = Scribe.connect({
token,
modelId: "scribe_v2_realtime",
microphone: {
deviceId: "optional-device-id",
echoCancellation: true,
noiseSuppression: true,
autoGainControl: true,
},
});| Property | Type | Description |
|---|---|---|
| deviceId | string |
Specific microphone device ID. |
| echoCancellation | boolean |
Enable echo cancellation. |
| noiseSuppression | boolean |
Enable noise suppression. |
| autoGainControl | boolean |
Enable automatic gain control. |
Manual audio options
Section titled “Manual audio options”Pass audioFormat and sampleRate to send audio data manually via connection.send().
import { AudioFormat } from "@elevenlabs/client";
const connection = Scribe.connect({
token,
modelId: "scribe_v2_realtime",
audioFormat: AudioFormat.PCM_16000,
sampleRate: 16000,
});| Property | Type | Description |
|---|---|---|
| audioFormat | AudioFormat |
Audio encoding format (e.g., AudioFormat.PCM_16000). |
| sampleRate | number |
Sample rate in Hz. Must match audioFormat. |
AudioFormat enum
Section titled “AudioFormat enum”enum AudioFormat {
PCM_8000 = "pcm_8000",
PCM_16000 = "pcm_16000",
PCM_22050 = "pcm_22050",
PCM_24000 = "pcm_24000",
PCM_44100 = "pcm_44100",
PCM_48000 = "pcm_48000",
ULAW_8000 = "ulaw_8000",
}Microphone mode
Section titled “Microphone mode”Stream audio directly from the user's microphone:
import { Scribe, RealtimeEvents } from "@elevenlabs/client";
async function transcribeFromMicrophone() {
const token = await fetchToken();
const connection = Scribe.connect({
token,
modelId: "scribe_v2_realtime",
microphone: {
echoCancellation: true,
noiseSuppression: true,
autoGainControl: true,
},
});
connection.on(RealtimeEvents.PARTIAL_TRANSCRIPT, (data) => {
document.getElementById("live").textContent = data.text;
});
connection.on(RealtimeEvents.COMMITTED_TRANSCRIPT, (data) => {
const el = document.createElement("p");
el.textContent = data.text;
document.getElementById("transcripts").appendChild(el);
document.getElementById("live").textContent = "";
});
document.getElementById("stop").addEventListener("click", () => {
connection.close();
});
}Manual audio mode (file transcription)
Section titled “Manual audio mode (file transcription)”Transcribe pre-recorded audio files by sending audio data manually:
import { Scribe, RealtimeEvents, AudioFormat } from "@elevenlabs/client";
async function transcribeFile(file) {
const token = await fetchToken();
const connection = Scribe.connect({
token,
modelId: "scribe_v2_realtime",
audioFormat: AudioFormat.PCM_16000,
sampleRate: 16000,
});
connection.on(RealtimeEvents.COMMITTED_TRANSCRIPT, (data) => {
console.log("Transcript:", data.text);
});
// Decode audio file
const arrayBuffer = await file.arrayBuffer();
const audioContext = new AudioContext({ sampleRate: 16000 });
const audioBuffer = await audioContext.decodeAudioData(arrayBuffer);
// Convert to PCM16
const channelData = audioBuffer.getChannelData(0);
const pcmData = new Int16Array(channelData.length);
for (let i = 0; i < channelData.length; i++) {
const sample = Math.max(-1, Math.min(1, channelData[i]));
pcmData[i] = sample < 0 ? sample * 32768 : sample * 32767;
}
// Send in chunks
const chunkSize = 4096;
for (let offset = 0; offset < pcmData.length; offset += chunkSize) {
const chunk = pcmData.slice(offset, offset + chunkSize);
const bytes = new Uint8Array(chunk.buffer);
const base64 = btoa(String.fromCharCode(...bytes));
connection.send({ audioBase64: base64 });
await new Promise((resolve) => setTimeout(resolve, 50));
}
// Commit and close
connection.commit();
}RealtimeConnection
Section titled “RealtimeConnection”Scribe.connect() returns a RealtimeConnection instance with the following methods.
on(event, listener)
Section titled “on(event, listener)”Register an event listener. See Events for available event types.
connection.on(RealtimeEvents.COMMITTED_TRANSCRIPT, (data) => {
console.log("Committed:", data.text);
});off(event, listener)
Section titled “off(event, listener)”Remove a previously registered event listener.
const handler = (data) => console.log(data.text);
connection.on(RealtimeEvents.COMMITTED_TRANSCRIPT, handler);
// Later
connection.off(RealtimeEvents.COMMITTED_TRANSCRIPT, handler);send(data)
Section titled “send(data)”Send audio data to Scribe (manual audio mode only).
connection.send({
audioBase64: base64AudioChunk,
commit: false, // Optional: commit immediately
sampleRate: 16000, // Optional: override sample rate
previousText: "Previous transcription text", // Optional: context from a previous transcription
});commit()
Section titled “commit()”Manually commit the current transcription. Only needed when using CommitStrategy.MANUAL.
connection.commit();close()
Section titled “close()”Close the WebSocket connection and clean up resources (microphone stream, audio context).
connection.close();Events
Section titled “Events”Register event listeners using connection.on(event, listener). All events are available as constants on the RealtimeEvents enum.
Transcription events
Section titled “Transcription events”| Event | Data | Description |
|---|---|---|
| SESSION_STARTED | { session_id: string } |
Scribe session started. |
| PARTIAL_TRANSCRIPT | { text: string } |
Interim transcription result. |
| COMMITTED_TRANSCRIPT | { text: string } |
Finalized transcription result. |
| COMMITTED_TRANSCRIPT_WITH_TIMESTAMPS | { text: string; language_code?: string; words?: WordsItem[] } |
Finalized result with word-level timing. |
The WordsItem type contains word-level timing information:
interface WordsItem {
text?: string; // Word text
start?: number; // Start time in seconds
end?: number; // End time in seconds
type?: "word" | "spacing"; // Token type
speaker_id?: string; // Speaker identifier
}Connection events
Section titled “Connection events”| Event | Data | Description |
|---|---|---|
| OPEN | Event |
WebSocket connection opened. |
| CLOSE | Event |
WebSocket connection closed. |
| ERROR | Error | Event |
Generic error. |
Error events
Section titled “Error events”All error events receive { error: string }.
| Event | Description |
|---|---|
| AUTH_ERROR | Authentication error. |
| QUOTA_EXCEEDED | Usage quota exceeded. |
| COMMIT_THROTTLED | Commit request throttled. |
| TRANSCRIBER_ERROR | Transcription engine error. |
| UNACCEPTED_TERMS | Terms of service not accepted. |
| RATE_LIMITED | Rate limited. |
| INPUT_ERROR | Invalid input format. |
| QUEUE_OVERFLOW | Processing queue full. |
| RESOURCE_EXHAUSTED | Server resources at capacity. |
| SESSION_TIME_LIMIT_EXCEEDED | Maximum session time reached. |
| CHUNK_SIZE_EXCEEDED | Audio chunk too large. |
| INSUFFICIENT_AUDIO_ACTIVITY | Not enough audio activity to maintain the connection. |
Commit strategies
Section titled “Commit strategies”Control when transcriptions are committed:
import { Scribe, CommitStrategy } from '@elevenlabs/client';
// Manual (default): you control when to commit
const connection = Scribe.connect({
token,
modelId: 'scribe_v2_realtime',
audioFormat: AudioFormat.PCM_16000,
sampleRate: 16000,
commitStrategy: CommitStrategy.MANUAL,
});
// Send audio, then commit when ready
connection.send({ audioBase64: chunk });
connection.commit();
// Voice Activity Detection: Scribe detects silences and commits automatically
const connection = Scribe.connect({
token,
modelId: 'scribe_v2_realtime',
microphone: { echoCancellation: true },
commitStrategy: CommitStrategy.VAD,
});For more details, see Transcripts and commit strategies.
Complete example
Section titled “Complete example”Here is a complete example that transcribes microphone audio with VAD-based commit strategy:
import { Scribe, RealtimeEvents, CommitStrategy } from "@elevenlabs/client";
async function startTranscription() {
const token = await fetchToken();
const connection = Scribe.connect({
token,
modelId: "scribe_v2_realtime",
commitStrategy: CommitStrategy.VAD,
microphone: {
echoCancellation: true,
noiseSuppression: true,
},
});
connection.on(RealtimeEvents.SESSION_STARTED, (data) => {
console.log("Session started:", data.session_id);
});
connection.on(RealtimeEvents.PARTIAL_TRANSCRIPT, (data) => {
document.getElementById("live").textContent = data.text;
});
connection.on(RealtimeEvents.COMMITTED_TRANSCRIPT, (data) => {
const el = document.createElement("p");
el.textContent = data.text;
document.getElementById("transcripts").appendChild(el);
document.getElementById("live").textContent = "";
});
connection.on(RealtimeEvents.ERROR, (error) => {
console.error("Scribe error:", error);
});
// Stop button
document.getElementById("stop").addEventListener("click", () => {
connection.close();
});
}
document.getElementById("start").addEventListener("click", startTranscription);