How to Configure Kokoro TTS with Voice Presets and Custom Voice Files
You can configure Kokoro TTS voices at three levels—handler initialization for session defaults, runtime configuration for per-request overrides, or LLM response metadata for dynamic selection—though the system only supports predefined voice identifiers rather than arbitrary custom voice files.
The huggingface/speech-to-speech repository implements a modular text-to-speech pipeline using the Kokoro TTS engine via the KokoroTTSHandler class. Understanding the voice configuration hierarchy allows you to control speech synthesis for multilingual applications without restarting the handler or modifying the underlying model code.
Understanding the Kokoro TTS Handler Architecture
The core implementation resides in src/speech_to_speech/TTS/kokoro_handler.py. The KokoroTTSHandler class extends the base handler architecture and manages voice selection through a cascading priority system defined in the process method (lines 66-74). The handler maintains two critical mapping dictionaries: WHISPER_LANGUAGE_TO_KOKORO_LANG (lines 31-47) for language code translation, and KOKORO_LANG_DEFAULT_VOICES (lines 63-73) for default voice assignment per language.
Three Methods to Configure Kokoro TTS Voices
1. Handler Initialization (Session Defaults)
Configure the default voice when instantiating the handler by passing parameters to the setup method. This establishes a baseline voice and language code for the entire session.
from src.speech_to_speech.TTS.kokoro_handler import KokoroTTSHandler
handler = KokoroTTSHandler()
handler.setup(
should_listen=listen_event,
voice="af_heart", # American English female voice
lang_code="a", # Language code for American English
speed=1.0,
)
The voice parameter accepts pretrained identifiers such as bm_fable (British male), af_heart (American female), or zf_xiaobei (Chinese female). The lang_code parameter maps to Kokoro's internal language identifiers ("a" for American English, "b" for British English, etc.).
2. Runtime Configuration (Per-Request Overrides)
Override the session default for individual requests by including a voice specification in the runtime_config object. The handler checks runtime_config.session.audio.output.voice during processing.
{
"runtime_config": {
"session": {
"audio": {
"output": {
"voice": "ef_dora"
}
}
}
},
"text": "Hola, ¿cómo estás?"
}
This method updates self.voice within the process method without requiring handler reinitialization, making it ideal for multilingual applications where the target language changes between requests.
3. LLM Response Metadata (Dynamic Selection)
Enable the language model to select voices dynamically by returning an Audio object with a voice field in the response metadata. This takes highest priority in the voice selection cascade.
{
"response": {
"audio": {
"output": {
"voice": "ff_siwis"
}
}
},
"text": "Bonjour, je suis votre assistant."
}
When the LLM returns this structure, the handler immediately adopts the specified voice for that specific synthesis operation, allowing contextual voice switching based on conversation content or user preferences.
Voice Selection Priority and Logic
The process method implements a strict precedence hierarchy (lines 66-74 in kokoro_handler.py):
- LLM Response – If the incoming message contains
response.audio.output.voice, this value overrides all other settings. - Runtime Configuration – If no LLM voice is present but
runtime_config.session.audio.output.voiceexists, the handler adopts this value. - Handler Fallback – If neither dynamic source provides a voice, the handler retains the voice established during
setup.
This cascade ensures that system defaults persist while allowing granular control at the conversation or request level.
Language Mapping and Automatic Voice Switching
The handler includes intelligent language detection through WHISPER_LANGUAGE_TO_KOKORO_LANG, which maps Whisper-detected language codes to Kokoro's language identifiers. When the input language changes, the handler references KOKORO_LANG_DEFAULT_VOICES to automatically select an appropriate default voice for that language (implemented in _process_mlx and _process_kokoro around lines 86-100).
For example, if Whisper detects Spanish ("es"), the handler maps this to Kokoro's "e" language code and defaults to a Spanish voice preset unless explicitly overridden by one of the three configuration methods.
Limitations of Custom Voice Files
Kokoro TTS does not support loading arbitrary custom voice files. The system only accepts pretrained voice identifiers defined in KOKORO_LANG_DEFAULT_VOICES. Available options include bm_fable, af_heart, zf_xiaobei, and other bundled presets.
If you require a completely new voice, you must extend the upstream Kokoro model repository with a new pretrained voice rather than loading external voice files through this handler. The speech-to-speech repository strictly consumes existing voice identifiers from the Kokoro ecosystem.
Summary
- Configure session defaults using
KokoroTTSHandler.setup()with thevoiceandlang_codeparameters. - Override voices per-request via
runtime_config.session.audio.output.voicein the request payload. - Enable dynamic voice selection by having the LLM return
response.audio.output.voicemetadata. - Voice selection follows a strict priority: LLM response > Runtime config > Handler initialization.
- Automatic language mapping uses
WHISPER_LANGUAGE_TO_KOKORO_LANGandKOKORO_LANG_DEFAULT_VOICESto select appropriate defaults when languages change. - Custom voice files are not supported; only predefined identifiers from the Kokoro model repository are valid.
Frequently Asked Questions
How do I set a default voice for the entire session?
Call handler.setup() with the voice parameter when initializing the KokoroTTSHandler in src/speech_to_speech/TTS/kokoro_handler.py. This value persists across all synthesis operations unless overridden by runtime configuration or LLM responses.
Can I change voices dynamically for individual requests?
Yes. Include a voice field in runtime_config.session.audio.output within your request JSON. The handler's process method checks this value (lines 66-74) and updates the active voice before synthesis begins.
Does Kokoro TTS support loading custom voice files?
No. The handler only accepts pretrained voice identifiers such as af_heart or bm_fable defined in KOKORO_LANG_DEFAULT_VOICES. To use a new voice, you must add it to the upstream Kokoro model repository, as the current implementation does not expose an API for arbitrary voice file loading.
How does automatic language detection work with voice selection?
The handler uses WHISPER_LANGUAGE_TO_KOKORO_LANG (lines 31-47) to map Whisper language codes to Kokoro language identifiers. When the input language changes, the system automatically selects the default voice for that language from KOKORO_LANG_DEFAULT_VOICES (lines 63-73), though explicit voice configurations from any of the three methods always take precedence.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →