How to Use Predefined Voices in Pocket‑TTS: A Complete Guide
Pocket‑TTS ships with a catalog of ready‑to‑use voice prompts—including identifiers like "cosette", "marius", and "lola"—that you can select programmatically by passing a string key to TTSModel.get_state_for_audio_prompt.
The kyutai‑labs/pocket‑tts library provides built‑in speaker embeddings for immediate text‑to‑speech generation without requiring custom voice samples. These predefined voices are stored as a constant mapping in the source code and resolve automatically to Hugging Face hosted .safetensors files containing the acoustic conditioning required for consistent voice generation.
Where Predefined Voices Are Defined in the Source Code
The authoritative list of available voices lives in the private constant _ORIGINS_OF_PREDEFINED_VOICES located in pocket_tts/utils/utils.py (lines 15‑42). This dictionary maps each voice identifier to a URL pointing to a pre‑encoded audio prompt or .safetensors blob.
When you request a voice, the helper function get_predefined_voice (lines 45‑47 in the same file) resolves the language‑specific path. For example, if you load a Spanish model, the function selects the Spanish‑compatible embedding for that voice identifier rather than the default.
The core selection logic resides in pocket_tts/models/tts_model.py. Inside the get_state_for_audio_prompt method, an elif branch at lines 53‑68 checks whether the supplied audio_conditioning argument exists as a key in _ORIGINS_OF_PREDEFINED_VOICES. If matched, the model downloads the file (if necessary) via download_if_necessary and caches the speaker state.
How to Select Predefined Voices Programmatically
You do not need to handle URLs or file paths manually. Pass the voice identifier string directly to the model’s state retrieval method.
Basic Voice Selection
Load the default model and retrieve a predefined voice state by name:
from pocket_tts import TTSModel
# Load the default model (automatically downloads weights)
model = TTSModel.load_model()
# Retrieve the "cosette" voice state
voice_state = model.get_state_for_audio_prompt("cosette")
# Generate audio
audio_tensor = model.generate_audio(voice_state, "Hello, this is a test.")
Language‑Specific Voice Loading
If you initialize a model for a specific language, the same identifier resolves to the appropriate language‑specific embedding:
# Load a Spanish model
model = TTSModel.load_model(language="es")
# Get the "lola" voice (Spanish‑specific embedding)
voice_state = model.get_state_for_audio_prompt("lola")
audio_tensor = model.generate_audio(voice_state, "Hola, ¿cómo estás?")
How to List All Available Predefined Voices
To enumerate every predefined voice identifier at runtime, import the constant directly from the utilities module:
from pocket_tts.utils.utils import _ORIGINS_OF_PREDEFINED_VOICES
print("Available predefined voices:")
for name in sorted(_ORIGINS_OF_PREDEFINED_VOICES):
print(f" • {name}")
This iterates over the keys defined in pocket_tts/utils/utils.py, giving you the exact set of strings accepted by get_state_for_audio_prompt.
Command‑Line Usage
The CLI entry point in pocket_tts/main.py exposes the same functionality via the --voice argument. When you run pocket‑tts generate --voice cosette, the script internally calls model.get_state_for_audio_prompt("cosette"), utilizing the same resolution logic described above.
Summary
- Predefined voices are stored in
_ORIGINS_OF_PREDEFINED_VOICESinsidepocket_tts/utils/utils.py. - Select a voice by passing its identifier string (e.g.,
"cosette","marius") toTTSModel.get_state_for_audio_prompt. - The method automatically downloads, caches, and loads the speaker embedding from Hugging Face.
- Language‑specific models (e.g.,
language="es") automatically receive the correct embedding variant for the requested voice. - You can list all available identifiers by importing and inspecting
_ORIGINS_OF_PREDEFINED_VOICES.
Frequently Asked Questions
What predefined voices are available in Pocket‑TTS?
The library provides a curated set of voices such as "cosette", "marius", and "lola". The complete list is defined in the _ORIGINS_OF_PREDEFINED_VOICES constant in pocket_tts/utils/utils.py. You can print the sorted keys of this dictionary at runtime to see every supported identifier.
How does Pocket‑TTS download predefined voice files?
When get_state_for_audio_prompt receives a predefined voice name, it triggers get_predefined_voice to resolve the URL. The utility then calls download_if_necessary to fetch the .safetensors file from Hugging Face and caches it locally. Subsequent calls use the cached version, so no repeated downloads occur.
Can I use predefined voices with non‑English models?
Yes. Predefined voices are language‑aware. When you load a model with a specific language code—such as TTSModel.load_model(language="es")—the get_predefined_voice helper selects the embedding variant that matches that language. The same identifier (e.g., "lola") works across different language models but loads the appropriate acoustic conditioning.
Where is the voice selection logic implemented?
The selection branch is in pocket_tts/models/tts_model.py within the get_state_for_audio_prompt method (lines 53‑68). This code checks if the input string exists in _ORIGINS_OF_PREDEFINED_VOICES and, if so, resolves it via the utilities module before loading the state.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →