# How to Configure Kokoro TTS with Voice Presets and Custom Voice Files

> Learn to configure Kokoro TTS with voice presets. Discover how to manage voice settings at handler, runtime, and LLM metadata levels for dynamic speech synthesis.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-30

---

**You can configure Kokoro TTS voices at three levels—handler initialization for session defaults, runtime configuration for per-request overrides, or LLM response metadata for dynamic selection—though the system only supports predefined voice identifiers rather than arbitrary custom voice files.**

The `huggingface/speech-to-speech` repository implements a modular text-to-speech pipeline using the **Kokoro TTS** engine via the `KokoroTTSHandler` class. Understanding the voice configuration hierarchy allows you to control speech synthesis for multilingual applications without restarting the handler or modifying the underlying model code.

## Understanding the Kokoro TTS Handler Architecture

The core implementation resides in [`src/speech_to_speech/TTS/kokoro_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/kokoro_handler.py). The `KokoroTTSHandler` class extends the base handler architecture and manages voice selection through a cascading priority system defined in the `process` method (lines 66-74). The handler maintains two critical mapping dictionaries: `WHISPER_LANGUAGE_TO_KOKORO_LANG` (lines 31-47) for language code translation, and `KOKORO_LANG_DEFAULT_VOICES` (lines 63-73) for default voice assignment per language.

## Three Methods to Configure Kokoro TTS Voices

### 1. Handler Initialization (Session Defaults)

Configure the default voice when instantiating the handler by passing parameters to the `setup` method. This establishes a baseline voice and language code for the entire session.

```python
from src.speech_to_speech.TTS.kokoro_handler import KokoroTTSHandler

handler = KokoroTTSHandler()
handler.setup(
    should_listen=listen_event,
    voice="af_heart",          # American English female voice

    lang_code="a",             # Language code for American English

    speed=1.0,
)

```

The `voice` parameter accepts pretrained identifiers such as `bm_fable` (British male), `af_heart` (American female), or `zf_xiaobei` (Chinese female). The `lang_code` parameter maps to Kokoro's internal language identifiers ("a" for American English, "b" for British English, etc.).

### 2. Runtime Configuration (Per-Request Overrides)

Override the session default for individual requests by including a voice specification in the `runtime_config` object. The handler checks `runtime_config.session.audio.output.voice` during processing.

```json
{
  "runtime_config": {
    "session": {
      "audio": {
        "output": {
          "voice": "ef_dora"
        }
      }
    }
  },
  "text": "Hola, ¿cómo estás?"
}

```

This method updates `self.voice` within the `process` method without requiring handler reinitialization, making it ideal for multilingual applications where the target language changes between requests.

### 3. LLM Response Metadata (Dynamic Selection)

Enable the language model to select voices dynamically by returning an `Audio` object with a `voice` field in the response metadata. This takes highest priority in the voice selection cascade.

```json
{
  "response": {
    "audio": {
      "output": {
        "voice": "ff_siwis"
      }
    }
  },
  "text": "Bonjour, je suis votre assistant."
}

```

When the LLM returns this structure, the handler immediately adopts the specified voice for that specific synthesis operation, allowing contextual voice switching based on conversation content or user preferences.

## Voice Selection Priority and Logic

The `process` method implements a strict precedence hierarchy (lines 66-74 in [`kokoro_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/kokoro_handler.py)):

1. **LLM Response** – If the incoming message contains `response.audio.output.voice`, this value overrides all other settings.
2. **Runtime Configuration** – If no LLM voice is present but `runtime_config.session.audio.output.voice` exists, the handler adopts this value.
3. **Handler Fallback** – If neither dynamic source provides a voice, the handler retains the voice established during `setup`.

This cascade ensures that system defaults persist while allowing granular control at the conversation or request level.

## Language Mapping and Automatic Voice Switching

The handler includes intelligent language detection through `WHISPER_LANGUAGE_TO_KOKORO_LANG`, which maps Whisper-detected language codes to Kokoro's language identifiers. When the input language changes, the handler references `KOKORO_LANG_DEFAULT_VOICES` to automatically select an appropriate default voice for that language (implemented in `_process_mlx` and `_process_kokoro` around lines 86-100).

For example, if Whisper detects Spanish ("es"), the handler maps this to Kokoro's "e" language code and defaults to a Spanish voice preset unless explicitly overridden by one of the three configuration methods.

## Limitations of Custom Voice Files

**Kokoro TTS does not support loading arbitrary custom voice files.** The system only accepts pretrained voice identifiers defined in `KOKORO_LANG_DEFAULT_VOICES`. Available options include `bm_fable`, `af_heart`, `zf_xiaobei`, and other bundled presets.

If you require a completely new voice, you must extend the upstream Kokoro model repository with a new pretrained voice rather than loading external voice files through this handler. The `speech-to-speech` repository strictly consumes existing voice identifiers from the Kokoro ecosystem.

## Summary

- Configure session defaults using `KokoroTTSHandler.setup()` with the `voice` and `lang_code` parameters.
- Override voices per-request via `runtime_config.session.audio.output.voice` in the request payload.
- Enable dynamic voice selection by having the LLM return `response.audio.output.voice` metadata.
- Voice selection follows a strict priority: LLM response > Runtime config > Handler initialization.
- Automatic language mapping uses `WHISPER_LANGUAGE_TO_KOKORO_LANG` and `KOKORO_LANG_DEFAULT_VOICES` to select appropriate defaults when languages change.
- Custom voice files are not supported; only predefined identifiers from the Kokoro model repository are valid.

## Frequently Asked Questions

### How do I set a default voice for the entire session?

Call `handler.setup()` with the `voice` parameter when initializing the `KokoroTTSHandler` in [`src/speech_to_speech/TTS/kokoro_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/kokoro_handler.py). This value persists across all synthesis operations unless overridden by runtime configuration or LLM responses.

### Can I change voices dynamically for individual requests?

Yes. Include a `voice` field in `runtime_config.session.audio.output` within your request JSON. The handler's `process` method checks this value (lines 66-74) and updates the active voice before synthesis begins.

### Does Kokoro TTS support loading custom voice files?

No. The handler only accepts pretrained voice identifiers such as `af_heart` or `bm_fable` defined in `KOKORO_LANG_DEFAULT_VOICES`. To use a new voice, you must add it to the upstream Kokoro model repository, as the current implementation does not expose an API for arbitrary voice file loading.

### How does automatic language detection work with voice selection?

The handler uses `WHISPER_LANGUAGE_TO_KOKORO_LANG` (lines 31-47) to map Whisper language codes to Kokoro language identifiers. When the input language changes, the system automatically selects the default voice for that language from `KOKORO_LANG_DEFAULT_VOICES` (lines 63-73), though explicit voice configurations from any of the three methods always take precedence.