How to Add a Custom STT, LLM, or TTS Handler to the Speech-to-Speech Pipeline Architecture
To add a custom handler to the Hugging Face speech-to-speech pipeline, subclass BaseHandler (or BaseSTTHandler for speech-to-text), implement the setup() and process() methods, register the handler in get_stt_handler or get_tts_handler within s2s_pipeline.py, and expose configuration options via dataclasses in arguments_classes/.
The huggingface/speech-to-speech repository implements a modular pipeline architecture where each stage—VAD, STT, Language Model, LM-output-processor, and TTS—operates as an independent handler. This design allows you to inject custom implementations without modifying the core orchestration logic in s2s_pipeline.py. Whether you need to integrate a proprietary STT model, a local LLM, or a specialized TTS engine, the handler interface provides a standardized contract for queue-based communication between pipeline stages.
Understanding the Handler Architecture
Every pipeline component inherits from BaseHandler (defined in src/speech_to_speech/baseHandler.py), which manages queue handling, cancellation signals, and logging. STT handlers specifically inherit from BaseSTTHandler, which extends BaseHandler with speech-to-text specific typing.
The pipeline orchestration relies on two key selector functions in src/speech_to_speech/s2s_pipeline.py:
get_stt_handler()– Instantiates the speech-to-text handler based on the--sttargumentget_tts_handler()– Instantiates the text-to-speech handler based on the--ttsargument
Each handler must implement:
setup(**kwargs): Initialize models, load weights, or configure runtime parametersprocess(input_item): A generator method that yields output items (e.g.,Transcriptionobjects for STT, audio bytes for TTS)
Step 1: Implement the Custom Handler Class
Create a new file in the appropriate handler directory (src/speech_to_speech/STT/ or src/speech_to_speech/TTS/) and subclass the base handler.
STT Handler Example
The following EchoSTTHandler demonstrates the minimal implementation for a speech-to-text handler that returns the audio duration as text:
# src/speech_to_speech/STT/echo_stt_handler.py
from __future__ import annotations
import logging
from typing import Iterator
import numpy as np
from speech_to_speech.pipeline.handler_types import STTIn, STTOut
from speech_to_speech.pipeline.messages import Transcription
from speech_to_speech.STT.base_stt_handler import BaseSTTHandler
logger = logging.getLogger(__name__)
class EchoSTTHandler(BaseSTTHandler):
"""A minimal STT that echoes the duration of the input audio."""
def setup(self, **kwargs) -> None:
# No heavy model loading needed
logger.info("EchoSTTHandler initialised")
def process(self, vad_audio: STTIn) -> Iterator[STTOut]:
# `vad_audio.audio` is a NumPy array (float32) sampled at 16 kHz
duration_s = len(vad_audio.audio) / 16000.0
text = f"[Echo] Detected {duration_s:.2f}s of speech"
yield Transcription(
text=text,
language_code="en",
turn_id=vad_audio.turn_id,
turn_revision=vad_audio.turn_revision,
speech_stopped_at_s=vad_audio.created_at_s,
)
TTS Handler Example
The following ReverseTTSHandler demonstrates a text-to-speech handler that reverses input text and returns silent audio:
# src/speech_to_speech/TTS/reverse_tts_handler.py
from __future__ import annotations
import logging
from typing import Iterator
import numpy as np
from speech_to_speech.pipeline.handler_types import TTSIn, TTSOut
from speech_to_speech.pipeline.messages import AUDIO_RESPONSE_DONE, TTSInput
from speech_to_speech.TTS.base import BaseHandler
logger = logging.getLogger(__name__)
class ReverseTTSHandler(BaseHandler[TTSIn, TTSOut]):
"""A demo TTS that reverses the incoming text and returns a silent audio chunk."""
def setup(self, **kwargs) -> None:
logger.info("ReverseTTSHandler initialised")
def process(self, tts_input: TTSIn) -> Iterator[TTSOut]:
# Reverse the text for demonstration
if isinstance(tts_input, TTSInput):
reversed_text = tts_input.text[::-1]
logger.info("ReverseTTSHandler received: %s → %s", tts_input.text, reversed_text)
# Produce a single silent 1-second chunk (16 kHz int16) just to keep the pipeline happy
silence = np.zeros(16000, dtype=np.int16)
yield silence.tobytes()
# Signal that the assistant's response is finished
yield AUDIO_RESPONSE_DONE
Step 2: Define Arguments and CLI Integration
To expose your handler via command-line arguments, create a dataclass extending ModuleArguments.
# src/speech_to_speech/arguments_classes/echo_stt_arguments.py
from __future__ import annotations
from dataclasses import dataclass
from speech_to_speech.arguments_classes.module_arguments import ModuleArguments
@dataclass
class EchoSTTHandlerArguments(ModuleArguments):
"""Arguments for the EchoSTTHandler."""
# Add any custom parameters here (model paths, temperature, etc.)
pass
Register the new argument class in the package initialization:
# src/speech_to_speech/arguments_classes/__init__.py
from .echo_stt_arguments import EchoSTTHandlerArguments
# from .reverse_tts_arguments import ReverseTTSHandlerArguments # For TTS
Update src/speech_to_speech/arguments_classes/module_arguments.py to include your handler in the allowed choices:
@dataclass
class ModuleArguments:
stt: str = field(default="whisper", metadata={"choices": ["whisper", "faster-whisper", "paraformer", "echo"]})
tts: str = field(default="qwen3", metadata={"choices": ["qwen3", "pocket", "reverse"]})
Step 3: Register the Handler in the Pipeline Selectors
Edit the selector functions in src/speech_to_speech/s2s_pipeline.py to instantiate your handler when the corresponding argument is provided.
For STT handlers, extend get_stt_handler():
def get_stt_handler(..., echo_stt_handler_kwargs, ...) -> BaseHandler[STTIn, STTOut]:
# Existing branches...
elif module_kwargs.stt == "echo":
from speech_to_speech.STT.echo_stt_handler import EchoSTTHandler
return EchoSTTHandler(
stop_event,
queue_in=spoken_prompt_queue,
queue_out=text_prompt_queue,
setup_kwargs=vars(echo_stt_handler_kwargs),
)
For TTS handlers, extend get_tts_handler():
def get_tts_handler(..., reverse_tts_handler_kwargs, ...) -> BaseHandler[TTSIn, TTSOut]:
# Existing branches...
elif module_kwargs.tts == "reverse":
from speech_to_speech.TTS.reverse_tts_handler import ReverseTTSHandler
return ReverseTTSHandler(
stop_event,
queue_in=lm_processed_queue,
queue_out=send_audio_chunks_queue,
setup_kwargs=vars(reverse_tts_handler_kwargs),
)
Step 4: Run the Pipeline with Your Custom Handler
After installing the package in editable mode (pip install -e .), launch the pipeline with your custom handlers:
python -m speech_to_speech.s2s_pipeline \
--stt echo \
--tts reverse \
--mode local
The pipeline logs will confirm instantiation of EchoSTTHandler and ReverseTTSHandler, and the system will process audio through your custom implementations.
Key Implementation Files
src/speech_to_speech/baseHandler.py– Defines theBaseHandlerclass with queue management and theprocess()generator contractsrc/speech_to_speech/STT/base_stt_handler.py– STT-specific base class withSTTInandSTTOuttype definitionssrc/speech_to_speech/s2s_pipeline.py– Containsget_stt_handler()andget_tts_handler()selector functionssrc/speech_to_speech/arguments_classes/module_arguments.py– Defines CLI argument schemas and allowed handler choicessrc/speech_to_speech/arguments_classes/__init__.py– Registration point for new argument dataclasses
Summary
- Inherit from the correct base class: Use
BaseSTTHandlerfor speech-to-text andBaseHandlerfor LLM or TTS components. - Implement the generator pattern: The
process()method must yieldTranscriptionobjects (STT), audio bytes (TTS), or appropriate message types rather than returning single values. - Register in selectors: Add instantiation logic to
get_stt_handler()orget_tts_handler()ins2s_pipeline.pybased on themodule_kwargsvalue. - Expose via arguments: Create dataclasses in
arguments_classes/and updateModuleArgumentsto include your handler in the CLI choices. - Respect the queue contract: Handlers communicate via thread-safe queues; never block indefinitely in
process()without checkingstop_event.
Frequently Asked Questions
How do I add a custom LLM handler instead of STT or TTS?
The process mirrors the TTS implementation: subclass BaseHandler, implement setup() and process(), and register it in the LLM selector function within s2s_pipeline.py. LLM handlers typically yield AssistantResponse objects rather than audio bytes. Reference the existing Qwen or Llama handlers in src/speech_to_speech/LM/ for the specific input/output message types required.
Can I use multiple custom handlers simultaneously?
Yes. The pipeline supports mixing custom and built-in handlers. You can use a custom STT handler with a standard TTS handler, or vice versa, by specifying different values for --stt and --tts. Each handler operates in its own thread with independent queues, so they execute concurrently without interference.
What is the difference between BaseHandler and BaseSTTHandler?
BaseSTTHandler extends BaseHandler with type hints specific to speech-to-text (using STTIn and STTOut type variables) and may include STT-specific utility methods. For LLM and TTS handlers, use BaseHandler directly. Both enforce the same setup() and process() interface, but BaseSTTHandler provides clearer type safety when working with audio transcription inputs.
How do I pass model paths or hyperparameters to my custom handler?
Add fields to your *HandlerArguments dataclass (e.g., model_path: str = field(default="")). These fields automatically become available as command-line arguments (e.g., --model_path). Access them in your handler's setup() method via the setup_kwargs dictionary passed during instantiation in the selector function.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →