Speech to Text quickstart
This guide will show you how to convert spoken audio into text using the Speech to Text API.
Using the Speech to Text API
Section titled “Using the Speech to Text API”Create an API key
Section titled “Create an API key”Create an API key in the dashboard here, which you’ll use to securely access the API.
Store the key as a managed secret and pass it to the SDKs either as a environment variable via an .env file, or directly in your app’s configuration depending on your preference.
.env
ELEVENLABS_API_KEY=<your_api_key_here>Install the SDK
Section titled “Install the SDK”We'll also use the dotenv library to load our API key from an environment variable.
pip install elevenlabs
pip install python-dotenvnpm install @elevenlabs/elevenlabs-js
npm install dotenvInstall the ElevenLabs CLI. Homebrew (macOS) and Scoop (Windows) are recommended.
Homebrew (macOS)
brew install elevenlabs/tap/elevenlabsScoop (Windows)
scoop bucket add elevenlabs https://github.com/elevenlabs/scoop-bucket
scoop install elevenlabsnpm
npm install -g @elevenlabs/clicurl
curl --proto '=https' --tlsv1.2 -LsSf https://github.com/elevenlabs/cli/releases/latest/download/elevenlabs-cli-installer.sh | shThen authenticate — this opens your browser to authorize the CLI:
elevenlabs auth loginMake the API request
Section titled “Make the API request”Create a new file named example.py or example.mts, depending on your language of choice and add the following code:
# example.py
import os
from dotenv import load_dotenv
from io import BytesIO
import requests
from elevenlabs.client import ElevenLabs
load_dotenv()
elevenlabs = ElevenLabs(
api_key=os.getenv("ELEVENLABS_API_KEY"),
)
audio_url = (
"https://storage.googleapis.com/eleven-public-cdn/audio/marketing/nicole.mp3"
)
response = requests.get(audio_url)
audio_data = BytesIO(response.content)
transcription = elevenlabs.speech_to_text.convert(
file=audio_data,
model_id="scribe_v2", # Model to use
tag_audio_events=True, # Tag audio events like laughter, applause, etc.
language_code="eng", # Language of the audio file. If set to None, the model will detect the language automatically.
diarize=True, # Whether to annotate who is speaking
)
print(transcription)// example.mts
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import "dotenv/config";
const elevenlabs = new ElevenLabsClient();
const response = await fetch(
"https://storage.googleapis.com/eleven-public-cdn/audio/marketing/nicole.mp3"
);
const audioBlob = new Blob([await response.arrayBuffer()], { type: "audio/mp3" });
const transcription = await elevenlabs.speechToText.convert({
file: audioBlob,
modelId: "scribe_v2", // Model to use
tagAudioEvents: true, // Tag audio events like laughter, applause, etc.
languageCode: "eng", // Language of the audio file. If set to null, the model will detect the language automatically.
diarize: true, // Whether to annotate who is speaking
});
console.log(transcription);Then run it:
python example.pynpx tsx example.mtsYou should see the transcription of the audio file printed to the console.
Download the sample audio, then transcribe it:
curl -O https://storage.googleapis.com/eleven-public-cdn/audio/marketing/nicole.mp3
elevenlabs speech-to-text convert \
--file nicole.mp3 \
--model-id scribe_v2 \
--tag-audio-events true \
--language-code eng \
--diarize trueThe transcription is printed to your terminal.
For medical and clinical audio, pass --model-id scribe_v2_medical.
Next steps
Section titled “Next steps”Transcribe pre-recorded audio files with speaker diarization and event tagging
Stream audio and receive transcriptions in real time
Explore all Speech to Text parameters and response formats