# How to Integrate a Custom TTS or STT Handler into the Speech-to-Speech Pipeline

> Integrate custom TTS or STT handlers into the speech-to-speech pipeline by subclassing, implementing methods, and registering handlers for enhanced voice AI applications.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-08

---

**Integrate a custom TTS or STT handler by subclassing `BaseHandler` (or `BaseSTTHandler`), implementing the `setup` and `process` generator methods, registering the handler in the selector functions within [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py), and exposing CLI arguments via a dedicated dataclass in `src/speech_to_speech/arguments_classes/`.**

The huggingface/speech-to-speech repository implements a modular handler architecture where each pipeline stage—VAD → STT → LM → LM-output-processor → TTS—operates as an independent class. To integrate a custom TTS or STT handler, you extend the base abstractions and register your implementation in the orchestration layer without touching the core queue management logic.

## Understanding the Handler Architecture

Every pipeline component inherits from `BaseHandler` defined in [`src/speech_to_speech/baseHandler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/baseHandler.py). STT handlers specifically inherit from `BaseSTTHandler`, which extends `BaseHandler` with speech-to-text-specific input/output types.

The main orchestration in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) uses two key factory functions:
- `get_stt_handler` – Instantiates the active STT implementation based on the `--stt` CLI argument
- `get_tts_handler` – Instantiates the active TTS implementation based on the `--tts` CLI argument

Each handler must implement:
- `setup(**kwargs)` – One-time initialization (model loading, warmup)
- `process(input_item)` – A generator yielding output items

## Step 1: Create a Custom Handler Class

Subclass the appropriate base class and implement the required methods. The `process` method must yield type-appropriate objects: `Transcription` instances for STT, raw audio bytes or `AUDIO_RESPONSE_DONE` for TTS.

### STT Handler Example

Create a file at [`src/speech_to_speech/STT/echo_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/echo_stt_handler.py):

```python
from __future__ import annotations
import logging
from typing import Iterator

import numpy as np
from speech_to_speech.pipeline.handler_types import STTIn, STTOut
from speech_to_speech.pipeline.messages import Transcription
from speech_to_speech.STT.base_stt_handler import BaseSTTHandler

logger = logging.getLogger(__name__)

class EchoSTTHandler(BaseSTTHandler):
    """A minimal STT that echoes the duration of the input audio."""
    def setup(self, **kwargs) -> None:
        # No heavy model loading needed

        logger.info("EchoSTTHandler initialised")

    def process(self, vad_audio: STTIn) -> Iterator[STTOut]:
        # `vad_audio.audio` is a NumPy array (float32) sampled at 16 kHz

        duration_s = len(vad_audio.audio) / 16000.0
        text = f"[Echo] Detected {duration_s:.2f}s of speech"
        yield Transcription(
            text=text,
            language_code="en",
            turn_id=vad_audio.turn_id,
            turn_revision=vad_audio.turn_revision,
            speech_stopped_at_s=vad_audio.created_at_s,
        )

```

### TTS Handler Example

Create a file at [`src/speech_to_speech/TTS/reverse_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/reverse_tts_handler.py):

```python
from __future__ import annotations
import logging
from typing import Iterator

import numpy as np
from speech_to_speech.pipeline.handler_types import TTSIn, TTSOut
from speech_to_speech.pipeline.messages import AUDIO_RESPONSE_DONE, TTSInput
from speech_to_speech.TTS.base import BaseHandler

logger = logging.getLogger(__name__)

class ReverseTTSHandler(BaseHandler[TTSIn, TTSOut]):
    """A demo TTS that reverses the incoming text and returns a silent audio chunk."""
    def setup(self, **kwargs) -> None:
        logger.info("ReverseTTSHandler initialised")

    def process(self, tts_input: TTSIn) -> Iterator[TTSOut]:
        # Reverse the text for demonstration

        if isinstance(tts_input, TTSInput):
            reversed_text = tts_input.text[::-1]
            logger.info("ReverseTTSHandler received: %s → %s", tts_input.text, reversed_text)

        # Produce a single silent 1-second chunk (16 kHz int16) just to keep the pipeline happy

        silence = np.zeros(16000, dtype=np.int16)
        yield silence.tobytes()
        # Signal that the assistant's response is finished

        yield AUDIO_RESPONSE_DONE

```

## Step 2: Define Handler Arguments

Create a dataclass extending `ModuleArguments` to expose CLI flags for your handler.

Create [`src/speech_to_speech/arguments_classes/echo_stt_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/echo_stt_arguments.py):

```python
from __future__ import annotations
from dataclasses import dataclass

from speech_to_speech.arguments_classes.module_arguments import ModuleArguments

@dataclass
class EchoSTTHandlerArguments(ModuleArguments):
    """Arguments for the EchoSTTHandler."""
    pass

```

Export the class in [`src/speech_to_speech/arguments_classes/__init__.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/__init__.py):

```python
from .echo_stt_arguments import EchoSTTHandlerArguments

```

Update the `HfArgumentParser` tuple in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) to include your new arguments class alongside existing ones like `WhisperSTTHandlerArguments`.

## Step 3: Register the Handler in the Pipeline

Modify the selector functions in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) to instantiate your handler when the corresponding CLI argument is provided.

### Registering an STT Handler

Update the `get_stt_handler` function:

```python
def get_stt_handler(..., echo_stt_handler_kwargs, ...) -> BaseHandler[STTIn, STTOut]:
    # ... existing branches ...

    elif module_kwargs.stt == "echo":
        from speech_to_speech.STT.echo_stt_handler import EchoSTTHandler
        return EchoSTTHandler(
            stop_event,
            queue_in=spoken_prompt_queue,
            queue_out=text_prompt_queue,
            setup_kwargs=vars(echo_stt_handler_kwargs),
        )

```

### Registering a TTS Handler

Update the `get_tts_handler` function:

```python
def get_tts_handler(..., reverse_tts_handler_kwargs, ...) -> BaseHandler[TTSIn, TTSOut]:
    # ... existing branches ...

    elif module_kwargs.tts == "reverse":
        from speech_to_speech.TTS.reverse_tts_handler import ReverseTTSHandler
        return ReverseTTSHandler(
            stop_event,
            queue_in=lm_processed_queue,
            queue_out=send_audio_chunks_queue,
            setup_kwargs=vars(reverse_tts_handler_kwargs),
        )

```

### Update ModuleArguments

Add your handler keys to the allowed choices in [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py):

```python
from dataclasses import dataclass, field

@dataclass
class ModuleArguments:
    stt: str = field(default="whisper", metadata={"choices": ["whisper", "faster-whisper", "paraformer", "echo"]})
    tts: str = field(default="qwen3", metadata={"choices": ["qwen3", "pocket", "reverse"]})

```

## Step 4: Launch the Pipeline with Your Custom Handler

After reinstalling the package in editable mode:

```bash
pip install -e .

```

Run the pipeline with your custom handlers:

```bash
python -m speech_to_speech.s2s_pipeline \
    --stt echo \
    --tts reverse \
    --mode local

```

The logs will confirm instantiation of `EchoSTTHandler` and `ReverseTTSHandler`, and the pipeline will process audio through your custom implementations.

## Summary

- **Extend `BaseSTTHandler`** for speech-to-text or **`BaseHandler`** for text-to-speech to ensure compatibility with the queue-based pipeline architecture.
- **Implement `setup()`** for initialization and **`process()`** as a generator yielding `Transcription` objects (STT) or audio bytes/`AUDIO_RESPONSE_DONE` (TTS).
- **Create an arguments dataclass** in `src/speech_to_speech/arguments_classes/` and export it in [`__init__.py`](https://github.com/huggingface/speech-to-speech/blob/main/__init__.py) to enable CLI configuration.
- **Register the handler** in `get_stt_handler` or `get_tts_handler` inside [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) by mapping the CLI value to your class import and instantiation.
- **Update `ModuleArguments`** to include your handler name in the `choices` metadata for the `--stt` or `--tts` flags.

## Frequently Asked Questions

### What base class should I use for a custom STT handler?

Use `BaseSTTHandler` defined in [`src/speech_to_speech/STT/base_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/base_stt_handler.py). It inherits from `BaseHandler` and provides the correct type signatures for `STTIn` and `STTOut`, ensuring your handler receives `VADAudio` inputs and produces `Transcription` outputs that the downstream Language Model stage expects.

### How do I pass configuration parameters to my custom handler?

Add fields to your arguments dataclass (e.g., `EchoSTTHandlerArguments`), then access them via `setup_kwargs` in the selector function. The pipeline passes these kwargs to your handler's `setup()` method. For example, if your dataclass has a `model_path` field, you can access it inside `setup()` as `kwargs.get("model_path")`.

### Can I integrate external model servers or cloud APIs instead of local models?

Yes. Implement the `process()` method to make HTTP requests or gRPC calls to external services. Since `process()` is a generator, you can yield results as they arrive from the API. Ensure your `setup()` method handles any required authentication or connection pooling, and yield properly formatted `Transcription` or audio byte objects to maintain pipeline flow.

### Where is the queue management handled so I don't need to implement it myself?

The `BaseHandler` class in [`src/speech_to_speech/baseHandler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/baseHandler.py) implements the `run()` method, which manages the input/output queues, handles cancellation via `stop_event`, and calls your `process()` implementation. You only need to focus on the business logic inside `setup()` and `process()`; the threading and queue orchestration is abstracted away by the base class.