Skip to main content
ElevenLabs Documentation Docs

Search documentation

Type to search this documentation.

On this pageOverview

Using pronunciation dictionaries

Pronunciation dictionaries allow you to customize how your AI agent pronounces specific words or phrases. This is particularly useful for:

  • Correcting pronunciation of names, places, or technical terms
  • Ensuring consistent pronunciation across conversations
  • Customizing regional pronunciation variations

ElevenLabs supports both IPA and CMU alphabets.

In this example, we will create a pronunciation dictionary file for the word tomato.

This rule will use the "IPA" alphabet and update the pronunciation for tomato and Tomato with a different pronunciation. PLS files are case sensitive which is why we include it both with and without a capital "T".

dictionary.pls

dictionary.pls
<?xml version="1.0" encoding="UTF-8"?>
<lexicon version="1.0"
    xmlns="http://www.w3.org/2005/01/pronunciation-lexicon"
    xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
    xsi:schemaLocation="http://www.w3.org/2005/01/pronunciation-lexicon
        http://www.w3.org/TR/2007/CR-pronunciation-lexicon-20071212/pls.xsd"
    alphabet="ipa" xml:lang="en-US">
<lexeme>
    <grapheme>tomato</grapheme>
    <phoneme>/tə'meɪtoʊ/</phoneme>
</lexeme>
<lexeme>
    <grapheme>Tomato</grapheme>
    <phoneme>/tə'meɪtoʊ/</phoneme>
</lexeme>
</lexicon>

Create a pronunciation dictionary from a file via the SDK

Section titled “Create a pronunciation dictionary from a file via the SDK”

Create a new file named example.py or example.mts, depending on your language of choice and add the following code:

Python
from elevenlabs import ElevenLabs, PronunciationDictionaryVersionLocator
from elevenlabs.play import play

elevenlabs = ElevenLabs()

with open("dictionary.pls", "rb") as f:
    # this dictionary changes how tomato is pronounced
    pronunciation_dictionary = elevenlabs.pronunciation_dictionaries.create_from_file(
        file=f.read(), name="example"
    )

audio_1 = elevenlabs.text_to_speech.convert(
    text="Without the dictionary: tomato",
    voice_id="aMSt68OGf4xUZAnLpTU8",
    model_id="eleven_flash_v2",
)

audio_2 = elevenlabs.text_to_speech.convert(
    text="With the dictionary: tomato",
    voice_id="aMSt68OGf4xUZAnLpTU8",
    model_id="eleven_flash_v2",
    pronunciation_dictionary_locators=[
        PronunciationDictionaryVersionLocator(
            pronunciation_dictionary_id=pronunciation_dictionary.id,
            version_id=pronunciation_dictionary.version_id,
        )
    ],
)

# play the audio
play(audio_1)
play(audio_2)
TypeScript
import { ElevenLabsClient, play } from "@elevenlabs/elevenlabs-js";
import "dotenv/config";
import fs from "node:fs";

const elevenlabs = new ElevenLabsClient();

const pronunciationDictionary = await elevenlabs.pronunciationDictionaries.createFromFile({
    file: fs.createReadStream("dictionary.pls"),
});

const audio1 = await elevenlabs.textToSpeech.convert("Without the dictionary: tomato", {
    voiceId: "aMSt68OGf4xUZAnLpTU8",
    modelId: "eleven_flash_v2",
});

const audio2 = await elevenlabs.textToSpeech.convert("With the dictionary: tomato", {
    voiceId: "aMSt68OGf4xUZAnLpTU8",
    modelId: "eleven_flash_v2",
    pronunciationDictionaryLocators: [
        {
            pronunciationDictionaryId: pronunciationDictionary.id,
            versionId: pronunciationDictionary.versionId,
        },
    ],
});

play(audio1);
play(audio2);
Python
python example.py
TypeScript
npx tsx example.mts

You should hear two versions of the audio playing through your speakers, one with and one without the pronunciation dictionary.

Full pronunciation dictionary API reference.

Stream text to speech progressively for lower latency playback.

Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu