How to Integrate a Custom TTS or STT Handler into the Speech-to-Speech Pipeline
Integrate a custom TTS or STT handler by subclassing BaseHandler (or BaseSTTHandler), implementing the setup and process generator methods, registering the handler in the selector functions within s2s_pipeline.py, and exposing CLI arguments via a dedicated dataclass in src/speech_to_speech/arguments_classes/.
The huggingface/speech-to-speech repository implements a modular handler architecture where each pipeline stage—VAD → STT → LM → LM-output-processor → TTS—operates as an independent class. To integrate a custom TTS or STT handler, you extend the base abstractions and register your implementation in the orchestration layer without touching the core queue management logic.
Understanding the Handler Architecture
Every pipeline component inherits from BaseHandler defined in src/speech_to_speech/baseHandler.py. STT handlers specifically inherit from BaseSTTHandler, which extends BaseHandler with speech-to-text-specific input/output types.
The main orchestration in src/speech_to_speech/s2s_pipeline.py uses two key factory functions:
get_stt_handler– Instantiates the active STT implementation based on the--sttCLI argumentget_tts_handler– Instantiates the active TTS implementation based on the--ttsCLI argument
Each handler must implement:
setup(**kwargs)– One-time initialization (model loading, warmup)process(input_item)– A generator yielding output items
Step 1: Create a Custom Handler Class
Subclass the appropriate base class and implement the required methods. The process method must yield type-appropriate objects: Transcription instances for STT, raw audio bytes or AUDIO_RESPONSE_DONE for TTS.
STT Handler Example
Create a file at src/speech_to_speech/STT/echo_stt_handler.py:
from __future__ import annotations
import logging
from typing import Iterator
import numpy as np
from speech_to_speech.pipeline.handler_types import STTIn, STTOut
from speech_to_speech.pipeline.messages import Transcription
from speech_to_speech.STT.base_stt_handler import BaseSTTHandler
logger = logging.getLogger(__name__)
class EchoSTTHandler(BaseSTTHandler):
"""A minimal STT that echoes the duration of the input audio."""
def setup(self, **kwargs) -> None:
# No heavy model loading needed
logger.info("EchoSTTHandler initialised")
def process(self, vad_audio: STTIn) -> Iterator[STTOut]:
# `vad_audio.audio` is a NumPy array (float32) sampled at 16 kHz
duration_s = len(vad_audio.audio) / 16000.0
text = f"[Echo] Detected {duration_s:.2f}s of speech"
yield Transcription(
text=text,
language_code="en",
turn_id=vad_audio.turn_id,
turn_revision=vad_audio.turn_revision,
speech_stopped_at_s=vad_audio.created_at_s,
)
TTS Handler Example
Create a file at src/speech_to_speech/TTS/reverse_tts_handler.py:
from __future__ import annotations
import logging
from typing import Iterator
import numpy as np
from speech_to_speech.pipeline.handler_types import TTSIn, TTSOut
from speech_to_speech.pipeline.messages import AUDIO_RESPONSE_DONE, TTSInput
from speech_to_speech.TTS.base import BaseHandler
logger = logging.getLogger(__name__)
class ReverseTTSHandler(BaseHandler[TTSIn, TTSOut]):
"""A demo TTS that reverses the incoming text and returns a silent audio chunk."""
def setup(self, **kwargs) -> None:
logger.info("ReverseTTSHandler initialised")
def process(self, tts_input: TTSIn) -> Iterator[TTSOut]:
# Reverse the text for demonstration
if isinstance(tts_input, TTSInput):
reversed_text = tts_input.text[::-1]
logger.info("ReverseTTSHandler received: %s → %s", tts_input.text, reversed_text)
# Produce a single silent 1-second chunk (16 kHz int16) just to keep the pipeline happy
silence = np.zeros(16000, dtype=np.int16)
yield silence.tobytes()
# Signal that the assistant's response is finished
yield AUDIO_RESPONSE_DONE
Step 2: Define Handler Arguments
Create a dataclass extending ModuleArguments to expose CLI flags for your handler.
Create src/speech_to_speech/arguments_classes/echo_stt_arguments.py:
from __future__ import annotations
from dataclasses import dataclass
from speech_to_speech.arguments_classes.module_arguments import ModuleArguments
@dataclass
class EchoSTTHandlerArguments(ModuleArguments):
"""Arguments for the EchoSTTHandler."""
pass
Export the class in src/speech_to_speech/arguments_classes/__init__.py:
from .echo_stt_arguments import EchoSTTHandlerArguments
Update the HfArgumentParser tuple in s2s_pipeline.py to include your new arguments class alongside existing ones like WhisperSTTHandlerArguments.
Step 3: Register the Handler in the Pipeline
Modify the selector functions in src/speech_to_speech/s2s_pipeline.py to instantiate your handler when the corresponding CLI argument is provided.
Registering an STT Handler
Update the get_stt_handler function:
def get_stt_handler(..., echo_stt_handler_kwargs, ...) -> BaseHandler[STTIn, STTOut]:
# ... existing branches ...
elif module_kwargs.stt == "echo":
from speech_to_speech.STT.echo_stt_handler import EchoSTTHandler
return EchoSTTHandler(
stop_event,
queue_in=spoken_prompt_queue,
queue_out=text_prompt_queue,
setup_kwargs=vars(echo_stt_handler_kwargs),
)
Registering a TTS Handler
Update the get_tts_handler function:
def get_tts_handler(..., reverse_tts_handler_kwargs, ...) -> BaseHandler[TTSIn, TTSOut]:
# ... existing branches ...
elif module_kwargs.tts == "reverse":
from speech_to_speech.TTS.reverse_tts_handler import ReverseTTSHandler
return ReverseTTSHandler(
stop_event,
queue_in=lm_processed_queue,
queue_out=send_audio_chunks_queue,
setup_kwargs=vars(reverse_tts_handler_kwargs),
)
Update ModuleArguments
Add your handler keys to the allowed choices in src/speech_to_speech/arguments_classes/module_arguments.py:
from dataclasses import dataclass, field
@dataclass
class ModuleArguments:
stt: str = field(default="whisper", metadata={"choices": ["whisper", "faster-whisper", "paraformer", "echo"]})
tts: str = field(default="qwen3", metadata={"choices": ["qwen3", "pocket", "reverse"]})
Step 4: Launch the Pipeline with Your Custom Handler
After reinstalling the package in editable mode:
pip install -e .
Run the pipeline with your custom handlers:
python -m speech_to_speech.s2s_pipeline \
--stt echo \
--tts reverse \
--mode local
The logs will confirm instantiation of EchoSTTHandler and ReverseTTSHandler, and the pipeline will process audio through your custom implementations.
Summary
- Extend
BaseSTTHandlerfor speech-to-text orBaseHandlerfor text-to-speech to ensure compatibility with the queue-based pipeline architecture. - Implement
setup()for initialization andprocess()as a generator yieldingTranscriptionobjects (STT) or audio bytes/AUDIO_RESPONSE_DONE(TTS). - Create an arguments dataclass in
src/speech_to_speech/arguments_classes/and export it in__init__.pyto enable CLI configuration. - Register the handler in
get_stt_handlerorget_tts_handlerinsides2s_pipeline.pyby mapping the CLI value to your class import and instantiation. - Update
ModuleArgumentsto include your handler name in thechoicesmetadata for the--sttor--ttsflags.
Frequently Asked Questions
What base class should I use for a custom STT handler?
Use BaseSTTHandler defined in src/speech_to_speech/STT/base_stt_handler.py. It inherits from BaseHandler and provides the correct type signatures for STTIn and STTOut, ensuring your handler receives VADAudio inputs and produces Transcription outputs that the downstream Language Model stage expects.
How do I pass configuration parameters to my custom handler?
Add fields to your arguments dataclass (e.g., EchoSTTHandlerArguments), then access them via setup_kwargs in the selector function. The pipeline passes these kwargs to your handler's setup() method. For example, if your dataclass has a model_path field, you can access it inside setup() as kwargs.get("model_path").
Can I integrate external model servers or cloud APIs instead of local models?
Yes. Implement the process() method to make HTTP requests or gRPC calls to external services. Since process() is a generator, you can yield results as they arrive from the API. Ensure your setup() method handles any required authentication or connection pooling, and yield properly formatted Transcription or audio byte objects to maintain pipeline flow.
Where is the queue management handled so I don't need to implement it myself?
The BaseHandler class in src/speech_to_speech/baseHandler.py implements the run() method, which manages the input/output queues, handles cancellation via stop_event, and calls your process() implementation. You only need to focus on the business logic inside setup() and process(); the threading and queue orchestration is abstracted away by the base class.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →