How to Integrate a Custom TTS or STT Handler into the Speech-to-Speech Pipeline

Integrate a custom TTS or STT handler by subclassing BaseHandler (or BaseSTTHandler), implementing the setup and process generator methods, registering the handler in the selector functions within s2s_pipeline.py, and exposing CLI arguments via a dedicated dataclass in src/speech_to_speech/arguments_classes/.

The huggingface/speech-to-speech repository implements a modular handler architecture where each pipeline stage—VAD → STT → LM → LM-output-processor → TTS—operates as an independent class. To integrate a custom TTS or STT handler, you extend the base abstractions and register your implementation in the orchestration layer without touching the core queue management logic.

Understanding the Handler Architecture

Every pipeline component inherits from BaseHandler defined in src/speech_to_speech/baseHandler.py. STT handlers specifically inherit from BaseSTTHandler, which extends BaseHandler with speech-to-text-specific input/output types.

The main orchestration in src/speech_to_speech/s2s_pipeline.py uses two key factory functions:

  • get_stt_handler – Instantiates the active STT implementation based on the --stt CLI argument
  • get_tts_handler – Instantiates the active TTS implementation based on the --tts CLI argument

Each handler must implement:

  • setup(**kwargs) – One-time initialization (model loading, warmup)
  • process(input_item) – A generator yielding output items

Step 1: Create a Custom Handler Class

Subclass the appropriate base class and implement the required methods. The process method must yield type-appropriate objects: Transcription instances for STT, raw audio bytes or AUDIO_RESPONSE_DONE for TTS.

STT Handler Example

Create a file at src/speech_to_speech/STT/echo_stt_handler.py:

from __future__ import annotations
import logging
from typing import Iterator

import numpy as np
from speech_to_speech.pipeline.handler_types import STTIn, STTOut
from speech_to_speech.pipeline.messages import Transcription
from speech_to_speech.STT.base_stt_handler import BaseSTTHandler

logger = logging.getLogger(__name__)

class EchoSTTHandler(BaseSTTHandler):
    """A minimal STT that echoes the duration of the input audio."""
    def setup(self, **kwargs) -> None:
        # No heavy model loading needed

        logger.info("EchoSTTHandler initialised")

    def process(self, vad_audio: STTIn) -> Iterator[STTOut]:
        # `vad_audio.audio` is a NumPy array (float32) sampled at 16 kHz

        duration_s = len(vad_audio.audio) / 16000.0
        text = f"[Echo] Detected {duration_s:.2f}s of speech"
        yield Transcription(
            text=text,
            language_code="en",
            turn_id=vad_audio.turn_id,
            turn_revision=vad_audio.turn_revision,
            speech_stopped_at_s=vad_audio.created_at_s,
        )

TTS Handler Example

Create a file at src/speech_to_speech/TTS/reverse_tts_handler.py:

from __future__ import annotations
import logging
from typing import Iterator

import numpy as np
from speech_to_speech.pipeline.handler_types import TTSIn, TTSOut
from speech_to_speech.pipeline.messages import AUDIO_RESPONSE_DONE, TTSInput
from speech_to_speech.TTS.base import BaseHandler

logger = logging.getLogger(__name__)

class ReverseTTSHandler(BaseHandler[TTSIn, TTSOut]):
    """A demo TTS that reverses the incoming text and returns a silent audio chunk."""
    def setup(self, **kwargs) -> None:
        logger.info("ReverseTTSHandler initialised")

    def process(self, tts_input: TTSIn) -> Iterator[TTSOut]:
        # Reverse the text for demonstration

        if isinstance(tts_input, TTSInput):
            reversed_text = tts_input.text[::-1]
            logger.info("ReverseTTSHandler received: %s → %s", tts_input.text, reversed_text)

        # Produce a single silent 1-second chunk (16 kHz int16) just to keep the pipeline happy

        silence = np.zeros(16000, dtype=np.int16)
        yield silence.tobytes()
        # Signal that the assistant's response is finished

        yield AUDIO_RESPONSE_DONE

Step 2: Define Handler Arguments

Create a dataclass extending ModuleArguments to expose CLI flags for your handler.

Create src/speech_to_speech/arguments_classes/echo_stt_arguments.py:

from __future__ import annotations
from dataclasses import dataclass

from speech_to_speech.arguments_classes.module_arguments import ModuleArguments

@dataclass
class EchoSTTHandlerArguments(ModuleArguments):
    """Arguments for the EchoSTTHandler."""
    pass

Export the class in src/speech_to_speech/arguments_classes/__init__.py:

from .echo_stt_arguments import EchoSTTHandlerArguments

Update the HfArgumentParser tuple in s2s_pipeline.py to include your new arguments class alongside existing ones like WhisperSTTHandlerArguments.

Step 3: Register the Handler in the Pipeline

Modify the selector functions in src/speech_to_speech/s2s_pipeline.py to instantiate your handler when the corresponding CLI argument is provided.

Registering an STT Handler

Update the get_stt_handler function:

def get_stt_handler(..., echo_stt_handler_kwargs, ...) -> BaseHandler[STTIn, STTOut]:
    # ... existing branches ...

    elif module_kwargs.stt == "echo":
        from speech_to_speech.STT.echo_stt_handler import EchoSTTHandler
        return EchoSTTHandler(
            stop_event,
            queue_in=spoken_prompt_queue,
            queue_out=text_prompt_queue,
            setup_kwargs=vars(echo_stt_handler_kwargs),
        )

Registering a TTS Handler

Update the get_tts_handler function:

def get_tts_handler(..., reverse_tts_handler_kwargs, ...) -> BaseHandler[TTSIn, TTSOut]:
    # ... existing branches ...

    elif module_kwargs.tts == "reverse":
        from speech_to_speech.TTS.reverse_tts_handler import ReverseTTSHandler
        return ReverseTTSHandler(
            stop_event,
            queue_in=lm_processed_queue,
            queue_out=send_audio_chunks_queue,
            setup_kwargs=vars(reverse_tts_handler_kwargs),
        )

Update ModuleArguments

Add your handler keys to the allowed choices in src/speech_to_speech/arguments_classes/module_arguments.py:

from dataclasses import dataclass, field

@dataclass
class ModuleArguments:
    stt: str = field(default="whisper", metadata={"choices": ["whisper", "faster-whisper", "paraformer", "echo"]})
    tts: str = field(default="qwen3", metadata={"choices": ["qwen3", "pocket", "reverse"]})

Step 4: Launch the Pipeline with Your Custom Handler

After reinstalling the package in editable mode:

pip install -e .

Run the pipeline with your custom handlers:

python -m speech_to_speech.s2s_pipeline \
    --stt echo \
    --tts reverse \
    --mode local

The logs will confirm instantiation of EchoSTTHandler and ReverseTTSHandler, and the pipeline will process audio through your custom implementations.

Summary

  • Extend BaseSTTHandler for speech-to-text or BaseHandler for text-to-speech to ensure compatibility with the queue-based pipeline architecture.
  • Implement setup() for initialization and process() as a generator yielding Transcription objects (STT) or audio bytes/AUDIO_RESPONSE_DONE (TTS).
  • Create an arguments dataclass in src/speech_to_speech/arguments_classes/ and export it in __init__.py to enable CLI configuration.
  • Register the handler in get_stt_handler or get_tts_handler inside s2s_pipeline.py by mapping the CLI value to your class import and instantiation.
  • Update ModuleArguments to include your handler name in the choices metadata for the --stt or --tts flags.

Frequently Asked Questions

What base class should I use for a custom STT handler?

Use BaseSTTHandler defined in src/speech_to_speech/STT/base_stt_handler.py. It inherits from BaseHandler and provides the correct type signatures for STTIn and STTOut, ensuring your handler receives VADAudio inputs and produces Transcription outputs that the downstream Language Model stage expects.

How do I pass configuration parameters to my custom handler?

Add fields to your arguments dataclass (e.g., EchoSTTHandlerArguments), then access them via setup_kwargs in the selector function. The pipeline passes these kwargs to your handler's setup() method. For example, if your dataclass has a model_path field, you can access it inside setup() as kwargs.get("model_path").

Can I integrate external model servers or cloud APIs instead of local models?

Yes. Implement the process() method to make HTTP requests or gRPC calls to external services. Since process() is a generator, you can yield results as they arrive from the API. Ensure your setup() method handles any required authentication or connection pooling, and yield properly formatted Transcription or audio byte objects to maintain pipeline flow.

Where is the queue management handled so I don't need to implement it myself?

The BaseHandler class in src/speech_to_speech/baseHandler.py implements the run() method, which manages the input/output queues, handles cancellation via stop_event, and calls your process() implementation. You only need to focus on the business logic inside setup() and process(); the threading and queue orchestration is abstracted away by the base class.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →