Skip to main content
ElevenLabs Documentation Docs

Search documentation

Type to search this documentation.

On this pageOverview

Forced Alignment quickstart

This guide will show you how to use the Forced Alignment API to align text to audio.

Create an API key in the dashboard here, which you’ll use to securely access the API.

Store the key as a managed secret and pass it to the SDKs either as a environment variable via an .env file, or directly in your app’s configuration depending on your preference.

.env

.env
ELEVENLABS_API_KEY=<your_api_key_here>

We'll also use the dotenv library to load our API key from an environment variable.

Python
pip install elevenlabs
pip install python-dotenv
TypeScript
npm install @elevenlabs/elevenlabs-js
npm install dotenv

Install the ElevenLabs CLI. Homebrew (macOS) and Scoop (Windows) are recommended.

Homebrew (macOS)

Homebrew (macOS)
brew install elevenlabs/tap/elevenlabs

Scoop (Windows)

Scoop (Windows)
scoop bucket add elevenlabs https://github.com/elevenlabs/scoop-bucket
scoop install elevenlabs

npm

npm
npm install -g @elevenlabs/cli

curl

curl
curl --proto '=https' --tlsv1.2 -LsSf https://github.com/elevenlabs/cli/releases/latest/download/elevenlabs-cli-installer.sh | sh

Then authenticate — this opens your browser to authorize the CLI:

Bash
elevenlabs auth login

Create a new file named example.py or example.mts, depending on your language of choice and add the following code:

Python
# example.py
import os
from io import BytesIO
from elevenlabs.client import ElevenLabs
import requests
from dotenv import load_dotenv

load_dotenv()

elevenlabs = ElevenLabs(
    api_key=os.getenv("ELEVENLABS_API_KEY"),
)

audio_url = (
    "https://storage.googleapis.com/eleven-public-cdn/audio/marketing/nicole.mp3"
)
response = requests.get(audio_url)
audio_data = BytesIO(response.content)

# Perform the text-to-speech conversion
transcription = elevenlabs.forced_alignment.create(
    file=audio_data,
    text="With a soft and whispery American accent, I'm the ideal choice for creating ASMR content, meditative guides, or adding an intimate feel to your narrative projects."
)

print(transcription)
TypeScript
// example.ts
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import "dotenv/config";
const elevenlabs = new ElevenLabsClient();

const response = await fetch(
    "https://storage.googleapis.com/eleven-public-cdn/audio/marketing/nicole.mp3"
);
const audioBlob = new Blob([await response.arrayBuffer()], { type: "audio/mp3" });

const transcript = await elevenlabs.forcedAlignment.create({
    file: audioBlob,
    text: "With a soft and whispery American accent, I'm the ideal choice for creating ASMR content, meditative guides, or adding an intimate feel to your narrative projects."
})

console.log(transcript);

Then run it:

Python
python example.py
TypeScript
npx tsx example.mts

You should see the transcript of the audio file with exact timestamps printed to the console.

Download the sample audio, then align it with the transcript:

Bash
curl -O https://storage.googleapis.com/eleven-public-cdn/audio/marketing/nicole.mp3

elevenlabs forced-alignment create \
  --file nicole.mp3 \
  --text "With a soft and whispery American accent, I'm the ideal choice for creating ASMR content, meditative guides, or adding an intimate feel to your narrative projects."

The alignment with per-word timestamps is printed to your terminal.

Transcribe audio to text without requiring an existing transcript

Generate the audio from text to use with forced alignment

Explore all Forced Alignment parameters and response formats

Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu