# How to Add a Custom STT, LLM, or TTS Handler to the Speech-to-Speech Pipeline Architecture

> Integrate custom STT, LLM, or TTS handlers into the Hugging Face speech-to-speech pipeline. Learn to subclass BaseHandler, implement methods, and register your custom components.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-30

---

**To add a custom handler to the Hugging Face speech-to-speech pipeline, subclass `BaseHandler` (or `BaseSTTHandler` for speech-to-text), implement the `setup()` and `process()` methods, register the handler in `get_stt_handler` or `get_tts_handler` within [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py), and expose configuration options via dataclasses in `arguments_classes/`.**

The `huggingface/speech-to-speech` repository implements a modular pipeline architecture where each stage—VAD, STT, Language Model, LM-output-processor, and TTS—operates as an independent handler. This design allows you to inject custom implementations without modifying the core orchestration logic in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py). Whether you need to integrate a proprietary STT model, a local LLM, or a specialized TTS engine, the handler interface provides a standardized contract for queue-based communication between pipeline stages.

## Understanding the Handler Architecture

Every pipeline component inherits from **`BaseHandler`** (defined in [`src/speech_to_speech/baseHandler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/baseHandler.py)), which manages queue handling, cancellation signals, and logging. STT handlers specifically inherit from **`BaseSTTHandler`**, which extends `BaseHandler` with speech-to-text specific typing.

The pipeline orchestration relies on two key selector functions in **[`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py)**:
- `get_stt_handler()` – Instantiates the speech-to-text handler based on the `--stt` argument
- `get_tts_handler()` – Instantiates the text-to-speech handler based on the `--tts` argument

Each handler must implement:
- **`setup(**kwargs)`**: Initialize models, load weights, or configure runtime parameters
- **`process(input_item)`**: A generator method that yields output items (e.g., `Transcription` objects for STT, audio bytes for TTS)

## Step 1: Implement the Custom Handler Class

Create a new file in the appropriate handler directory (`src/speech_to_speech/STT/` or `src/speech_to_speech/TTS/`) and subclass the base handler.

### STT Handler Example

The following `EchoSTTHandler` demonstrates the minimal implementation for a speech-to-text handler that returns the audio duration as text:

```python

# src/speech_to_speech/STT/echo_stt_handler.py

from __future__ import annotations
import logging
from typing import Iterator

import numpy as np
from speech_to_speech.pipeline.handler_types import STTIn, STTOut
from speech_to_speech.pipeline.messages import Transcription
from speech_to_speech.STT.base_stt_handler import BaseSTTHandler

logger = logging.getLogger(__name__)

class EchoSTTHandler(BaseSTTHandler):
    """A minimal STT that echoes the duration of the input audio."""
    def setup(self, **kwargs) -> None:
        # No heavy model loading needed

        logger.info("EchoSTTHandler initialised")

    def process(self, vad_audio: STTIn) -> Iterator[STTOut]:
        # `vad_audio.audio` is a NumPy array (float32) sampled at 16 kHz

        duration_s = len(vad_audio.audio) / 16000.0
        text = f"[Echo] Detected {duration_s:.2f}s of speech"
        yield Transcription(
            text=text,
            language_code="en",
            turn_id=vad_audio.turn_id,
            turn_revision=vad_audio.turn_revision,
            speech_stopped_at_s=vad_audio.created_at_s,
        )

```

### TTS Handler Example

The following `ReverseTTSHandler` demonstrates a text-to-speech handler that reverses input text and returns silent audio:

```python

# src/speech_to_speech/TTS/reverse_tts_handler.py

from __future__ import annotations
import logging
from typing import Iterator

import numpy as np
from speech_to_speech.pipeline.handler_types import TTSIn, TTSOut
from speech_to_speech.pipeline.messages import AUDIO_RESPONSE_DONE, TTSInput
from speech_to_speech.TTS.base import BaseHandler

logger = logging.getLogger(__name__)

class ReverseTTSHandler(BaseHandler[TTSIn, TTSOut]):
    """A demo TTS that reverses the incoming text and returns a silent audio chunk."""
    def setup(self, **kwargs) -> None:
        logger.info("ReverseTTSHandler initialised")

    def process(self, tts_input: TTSIn) -> Iterator[TTSOut]:
        # Reverse the text for demonstration

        if isinstance(tts_input, TTSInput):
            reversed_text = tts_input.text[::-1]
            logger.info("ReverseTTSHandler received: %s → %s", tts_input.text, reversed_text)

        # Produce a single silent 1-second chunk (16 kHz int16) just to keep the pipeline happy

        silence = np.zeros(16000, dtype=np.int16)
        yield silence.tobytes()
        # Signal that the assistant's response is finished

        yield AUDIO_RESPONSE_DONE

```

## Step 2: Define Arguments and CLI Integration

To expose your handler via command-line arguments, create a dataclass extending `ModuleArguments`.

```python

# src/speech_to_speech/arguments_classes/echo_stt_arguments.py

from __future__ import annotations
from dataclasses import dataclass

from speech_to_speech.arguments_classes.module_arguments import ModuleArguments

@dataclass
class EchoSTTHandlerArguments(ModuleArguments):
    """Arguments for the EchoSTTHandler."""
    # Add any custom parameters here (model paths, temperature, etc.)

    pass

```

Register the new argument class in the package initialization:

```python

# src/speech_to_speech/arguments_classes/__init__.py

from .echo_stt_arguments import EchoSTTHandlerArguments

# from .reverse_tts_arguments import ReverseTTSHandlerArguments  # For TTS

```

Update **[`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py)** to include your handler in the allowed choices:

```python
@dataclass
class ModuleArguments:
    stt: str = field(default="whisper", metadata={"choices": ["whisper", "faster-whisper", "paraformer", "echo"]})
    tts: str = field(default="qwen3", metadata={"choices": ["qwen3", "pocket", "reverse"]})

```

## Step 3: Register the Handler in the Pipeline Selectors

Edit the selector functions in **[`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py)** to instantiate your handler when the corresponding argument is provided.

For STT handlers, extend `get_stt_handler()`:

```python
def get_stt_handler(..., echo_stt_handler_kwargs, ...) -> BaseHandler[STTIn, STTOut]:
    # Existing branches...

    elif module_kwargs.stt == "echo":
        from speech_to_speech.STT.echo_stt_handler import EchoSTTHandler
        return EchoSTTHandler(
            stop_event,
            queue_in=spoken_prompt_queue,
            queue_out=text_prompt_queue,
            setup_kwargs=vars(echo_stt_handler_kwargs),
        )

```

For TTS handlers, extend `get_tts_handler()`:

```python
def get_tts_handler(..., reverse_tts_handler_kwargs, ...) -> BaseHandler[TTSIn, TTSOut]:
    # Existing branches...

    elif module_kwargs.tts == "reverse":
        from speech_to_speech.TTS.reverse_tts_handler import ReverseTTSHandler
        return ReverseTTSHandler(
            stop_event,
            queue_in=lm_processed_queue,
            queue_out=send_audio_chunks_queue,
            setup_kwargs=vars(reverse_tts_handler_kwargs),
        )

```

## Step 4: Run the Pipeline with Your Custom Handler

After installing the package in editable mode (`pip install -e .`), launch the pipeline with your custom handlers:

```bash
python -m speech_to_speech.s2s_pipeline \
    --stt echo \
    --tts reverse \
    --mode local

```

The pipeline logs will confirm instantiation of `EchoSTTHandler` and `ReverseTTSHandler`, and the system will process audio through your custom implementations.

## Key Implementation Files

- **[`src/speech_to_speech/baseHandler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/baseHandler.py)** – Defines the `BaseHandler` class with queue management and the `process()` generator contract
- **[`src/speech_to_speech/STT/base_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/base_stt_handler.py)** – STT-specific base class with `STTIn` and `STTOut` type definitions
- **[`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py)** – Contains `get_stt_handler()` and `get_tts_handler()` selector functions
- **[`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py)** – Defines CLI argument schemas and allowed handler choices
- **[`src/speech_to_speech/arguments_classes/__init__.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/__init__.py)** – Registration point for new argument dataclasses

## Summary

- **Inherit from the correct base class**: Use `BaseSTTHandler` for speech-to-text and `BaseHandler` for LLM or TTS components.
- **Implement the generator pattern**: The `process()` method must yield `Transcription` objects (STT), audio bytes (TTS), or appropriate message types rather than returning single values.
- **Register in selectors**: Add instantiation logic to `get_stt_handler()` or `get_tts_handler()` in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) based on the `module_kwargs` value.
- **Expose via arguments**: Create dataclasses in `arguments_classes/` and update `ModuleArguments` to include your handler in the CLI choices.
- **Respect the queue contract**: Handlers communicate via thread-safe queues; never block indefinitely in `process()` without checking `stop_event`.

## Frequently Asked Questions

### How do I add a custom LLM handler instead of STT or TTS?

The process mirrors the TTS implementation: subclass `BaseHandler`, implement `setup()` and `process()`, and register it in the LLM selector function within [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py). LLM handlers typically yield `AssistantResponse` objects rather than audio bytes. Reference the existing Qwen or Llama handlers in `src/speech_to_speech/LM/` for the specific input/output message types required.

### Can I use multiple custom handlers simultaneously?

Yes. The pipeline supports mixing custom and built-in handlers. You can use a custom STT handler with a standard TTS handler, or vice versa, by specifying different values for `--stt` and `--tts`. Each handler operates in its own thread with independent queues, so they execute concurrently without interference.

### What is the difference between `BaseHandler` and `BaseSTTHandler`?

`BaseSTTHandler` extends `BaseHandler` with type hints specific to speech-to-text (using `STTIn` and `STTOut` type variables) and may include STT-specific utility methods. For LLM and TTS handlers, use `BaseHandler` directly. Both enforce the same `setup()` and `process()` interface, but `BaseSTTHandler` provides clearer type safety when working with audio transcription inputs.

### How do I pass model paths or hyperparameters to my custom handler?

Add fields to your `*HandlerArguments` dataclass (e.g., `model_path: str = field(default="")`). These fields automatically become available as command-line arguments (e.g., `--model_path`). Access them in your handler's `setup()` method via the `setup_kwargs` dictionary passed during instantiation in the selector function.