Using pronunciation dictionaries
This guide shows you how to manage pronunciation dictionaries programmatically.
Overview
Section titled “Overview”Pronunciation dictionaries allow you to customize how your AI agent pronounces specific words or phrases. This is particularly useful for:
- Correcting pronunciation of names, places, or technical terms
- Ensuring consistent pronunciation across conversations
- Customizing regional pronunciation variations
ElevenLabs supports both IPA and CMU alphabets.
Quickstart
Section titled “Quickstart”Create a pronunciation dictionary file
In this example, we will create a pronunciation dictionary file for the word
tomato.This rule will use the “IPA” alphabet and update the pronunciation for
tomatoandTomatowith a different pronunciation. PLS files are case sensitive which is why we include it both with and without a capital “T”.dictionary.pls
xml <?xml version="1.0" encoding="UTF-8"?> <lexicon version="1.0" xmlns="http://www.w3.org/2005/01/pronunciation-lexicon" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.w3.org/2005/01/pronunciation-lexicon http://www.w3.org/TR/2007/CR-pronunciation-lexicon-20071212/pls.xsd" alphabet="ipa" xml:lang="en-US"> <lexeme> <grapheme>tomato</grapheme> <phoneme>/tə'meɪtoʊ/</phoneme> </lexeme> <lexeme> <grapheme>Tomato</grapheme> <phoneme>/tə'meɪtoʊ/</phoneme> </lexeme> </lexicon>Create a pronunciation dictionary from a file via the SDK
Create a new file named
example.pyorexample.mts, depending on your language of choice and add the following code:Python from elevenlabs import ElevenLabs, PronunciationDictionaryVersionLocator from elevenlabs.play import play elevenlabs = ElevenLabs() with open("dictionary.pls", "rb") as f: # this dictionary changes how tomato is pronounced pronunciation_dictionary = elevenlabs.pronunciation_dictionaries.create_from_file( file=f.read(), name="example" ) audio_1 = elevenlabs.text_to_speech.convert( text="Without the dictionary: tomato", voice_id="aMSt68OGf4xUZAnLpTU8", model_id="eleven_flash_v2", ) audio_2 = elevenlabs.text_to_speech.convert( text="With the dictionary: tomato", voice_id="aMSt68OGf4xUZAnLpTU8", model_id="eleven_flash_v2", pronunciation_dictionary_locators=[ PronunciationDictionaryVersionLocator( pronunciation_dictionary_id=pronunciation_dictionary.id, version_id=pronunciation_dictionary.version_id, ) ], ) # play the audio play(audio_1) play(audio_2)Execute the code
Python python example.pyYou should hear two versions of the audio playing through your speakers, one with and one without the pronunciation dictionary.