How to Add a Custom STT Backend to the Speech-to-Speech Pipeline: A Complete Guide
You can add a custom Speech-to-Text backend to the huggingface/speech-to-speech pipeline by implementing a handler class that inherits from BaseSTTHandler, registering your identifier in the STTBackend enum, and wiring it into the get_stt_handler() dispatch logic in s2s_pipeline.py.
The huggingface/speech-to-speech repository provides a modular framework for real-time voice-to-voice conversion, supporting several built-in STT engines like Whisper and Paraformer. When you need to integrate a proprietary model or a specialized transcription system, the pipeline's dynamic instantiation architecture allows you to register new backends without modifying core processing logic.
Understanding the STT Handler Architecture
The pipeline builds its processing chain dynamically at runtime. STT backends are instantiated through the get_stt_handler() function located in src/speech_to_speech/s2s_pipeline.py (lines 749–789). This function inspects module_kwargs.stt—populated from command-line flags or JSON configuration files—to determine which concrete handler class to return.
All existing handlers inherit from BaseSTTHandler, defined in src/speech_to_speech/STT/base_stt_handler.py (lines 16–22). This base class provides speculative-turn filtering, stale-input dropping, and standardized queue-based messaging, ensuring that custom implementations only need to focus on transcription logic while inheriting robust conversation state management.
Step-by-Step Implementation Guide
Step 1: Create the Handler Class
Create a new file at src/speech_to_speech/STT/<your_name>_handler.py. Your class must inherit from BaseSTTHandler and implement the constructor signature (__init__(stop_event, queue_in, queue_out, setup_kwargs)). Override the run() method to process STTIn messages (typically VADAudio objects) and emit STTOut messages (PartialTranscription or Transcription).
from speech_to_speech.STT.base_stt_handler import BaseSTTHandler
from speech_to_speech.pipeline.handler_types import STTIn, STTOut
from speech_to_speech.pipeline.messages import Transcription
class MySTTHandler(BaseSTTHandler):
"""Example custom STT backend receiving VADAudio and emitting Transcription objects."""
def __init__(self, stop_event, queue_in, queue_out, setup_kwargs):
super().__init__(stop_event, queue_in, queue_out, setup_kwargs)
self.model_path = setup_kwargs.get("model_path", "default.pt")
# Load your model here
def run(self) -> None:
"""Main processing loop called by BaseHandler.run()."""
while not self.stop_event.is_set():
audio: STTIn = self.queue_in.get()
if not self.should_process_input(audio):
continue
# Replace with actual inference logic
transcription = Transcription(
turn_id=audio.turn_id,
turn_revision=audio.turn_revision,
created_at_s=audio.created_at_s,
text="processed transcription",
mode="final",
)
if self.should_emit_output(transcription):
self.queue_out.put(transcription)
By inheriting from BaseSTTHandler, you automatically receive should_process_input() and should_emit_output() filters that manage speculative turns and stale audio segments.
Step 2: Define Configuration Arguments
Create src/speech_to_speech/arguments_classes/<your_name>_stt_arguments.py to expose configuration options. Extend ArgumentsBase so that HfArgumentParser can auto-generate CLI flags.
from dataclasses import dataclass
from speech_to_speech.arguments_classes.module_arguments import ArgumentsBase
@dataclass
class MySTTHandlerArguments(ArgumentsBase):
"""Configuration for the MySTT backend."""
model_path: str = "models/my_stt.pt"
device: str = "cpu"
language: str = "en"
This creates typed configuration objects accessible via --my_stt_model_path, --my_stt_device, and similar flags.
Step 3: Register the Backend Enum
Modify src/speech_to_speech/arguments_classes/module_arguments.py to add your identifier to the STTBackend enum. This enum currently includes values like whisper, faster-whisper, parakeet-tdt, and paraformer.
from enum import Enum
class STTBackend(str, Enum):
WHISPER = "whisper"
WHISPER_MLX = "whisper-mlx"
MLX_AUDIO_WHISPER = "mlx-audio-whisper"
FASTER_WHISPER = "faster-whisper"
PARAKEET_TDT = "parakeet-tdt"
PARAFORMER = "paraformer"
MY_STT = "my_stt" # Your custom backend
Adding your identifier here enables the --stt my_stt CLI option and prevents configuration typos through enum validation.
Step 4: Wire into the Pipeline Dispatch
Add a dispatch branch in get_stt_handler() within src/speech_to_speech/s2s_pipeline.py (around lines 750–790). Import your handler class and return it wrapped with with_speculative_turns() to maintain consistency with built-in backends.
elif module_kwargs.stt == "my_stt":
from speech_to_speech.STT.my_stt_handler import MySTTHandler
return with_speculative_turns(
MySTTHandler(
stop_event,
queue_in=spoken_prompt_queue,
queue_out=text_prompt_queue,
setup_kwargs=vars(my_stt_handler_kwargs),
)
)
This branch mirrors the pattern used for Whisper and other built-in handlers, ensuring your backend integrates seamlessly with the pipeline's speculative-turn management.
Step 5: Handle Argument Parser Registration (Optional)
If your MySTTHandlerArguments introduces field names that clash with other argument classes, adjust the "pre-parse" logic in parse_arguments() within s2s_pipeline.py. Most implementations can simply add the new arguments class to the parser without modification, as HfArgumentParser resolves collisions by registration order.
Step 6: Write Integration Tests
Create tests/test_my_stt_handler.py to validate that the pipeline builds correctly with your backend and that audio flows through your handler to downstream components.
import json
import pathlib
from speech_to_speech.s2s_pipeline import parse_arguments, build_pipeline
def test_my_stt_integration(tmp_path):
cfg = {
"module_kwargs": {"stt": "my_stt", "tts": "qwen3", "mode": "local"},
"my_stt_handler_kwargs": {"model_path": str(tmp_path / "dummy.pt")},
# Add other required configuration sections...
}
cfg_path = tmp_path / "cfg.json"
cfg_path.write_text(json.dumps(cfg))
# Verify pipeline initializes without errors
args = parse_arguments([f"--config", str(cfg_path)])
# Additional assertions to verify handler instantiation...
Tests ensure future refactoring does not break your integration and document the expected message contract for other developers.
Key Implementation Details
The run() method operates on a polling loop until stop_event.is_set() signals shutdown. Input messages implement the STTIn protocol (typically VADAudio instances containing raw audio data and turn metadata), while output messages must conform to STTOut (PartialTranscription for interim results or Transcription for final results).
The with_speculative_turns() wrapper applied in get_stt_handler() manages conversation turn revisions automatically. Your handler receives audio segments tagged with turn_id and turn_revision fields, and the wrapper handles filtering obsolete revisions so your transcription logic can remain stateless regarding turn management.
Summary
- Inherit from
BaseSTTHandlerinsrc/speech_to_speech/STT/base_stt_handler.pyto leverage built-in speculative-turn filtering and stale-input detection - Create an arguments dataclass extending
ArgumentsBaseinsrc/speech_to_speech/arguments_classes/for type-safe configuration - Register your backend identifier in the
STTBackendenum inmodule_arguments.pyto enable--stt <your_id>CLI usage - Add a dispatch branch in
get_stt_handler()ins2s_pipeline.py(lines 749–789) to wire your handler into the pipeline - Wrap with
with_speculative_turns()to maintain consistency with existing backends and automatic turn revision handling - Place tests in
tests/to verify integration and prevent regressions
Frequently Asked Questions
What methods must a custom STT handler implement?
You must implement __init__ accepting stop_event, queue_in, queue_out, and setup_kwargs, plus a run() method that processes the input queue until the stop event is set. The base class provides should_process_input() and should_emit_output() lifecycle hooks for filtering speculative and stale data.
How does the pipeline handle speculative turns with custom backends?
The with_speculative_turns() wrapper in s2s_pipeline.py automatically wraps your handler instance when registered in get_stt_handler(). This wrapper manages turn revision tracking and filtering without requiring changes to your transcription logic, ensuring that only the latest audio for a given turn reaches downstream components.
Can I use a completely custom base class instead of BaseSTTHandler?
Yes, you can inherit directly from BaseHandler[STTIn, STTOut] if you need to bypass the speculative-turn filtering provided by BaseSTTHandler. However, you will need to reimplement stale-input detection and queue management logic yourself, as the base BaseSTTHandler provides these protections by default.
Where should I place my custom handler files in the repository?
Place your handler implementation in src/speech_to_speech/STT/<your_name>_handler.py and your arguments dataclass in src/speech_to_speech/arguments_classes/<your_name>_stt_arguments.py. This maintains consistency with the existing project structure and ensures import paths resolve correctly when the pipeline loads your module.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →