# How to Add a Custom STT Backend to the Speech-to-Speech Pipeline: A Complete Guide

> Learn how to add a custom STT backend to the huggingface speech-to-speech pipeline. Implement a handler class and register it for seamless integration. Complete guide.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-10

---

**You can add a custom Speech-to-Text backend to the huggingface/speech-to-speech pipeline by implementing a handler class that inherits from `BaseSTTHandler`, registering your identifier in the `STTBackend` enum, and wiring it into the `get_stt_handler()` dispatch logic in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py).**

The huggingface/speech-to-speech repository provides a modular framework for real-time voice-to-voice conversion, supporting several built-in STT engines like Whisper and Paraformer. When you need to integrate a proprietary model or a specialized transcription system, the pipeline's dynamic instantiation architecture allows you to register new backends without modifying core processing logic.

## Understanding the STT Handler Architecture

The pipeline builds its processing chain dynamically at runtime. STT backends are instantiated through the `get_stt_handler()` function located in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) (lines 749–789). This function inspects `module_kwargs.stt`—populated from command-line flags or JSON configuration files—to determine which concrete handler class to return.

All existing handlers inherit from `BaseSTTHandler`, defined in [`src/speech_to_speech/STT/base_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/base_stt_handler.py) (lines 16–22). This base class provides speculative-turn filtering, stale-input dropping, and standardized queue-based messaging, ensuring that custom implementations only need to focus on transcription logic while inheriting robust conversation state management.

## Step-by-Step Implementation Guide

### Step 1: Create the Handler Class

Create a new file at `src/speech_to_speech/STT/<your_name>_handler.py`. Your class must inherit from `BaseSTTHandler` and implement the constructor signature `(__init__(stop_event, queue_in, queue_out, setup_kwargs))`. Override the `run()` method to process `STTIn` messages (typically `VADAudio` objects) and emit `STTOut` messages (`PartialTranscription` or `Transcription`).

```python
from speech_to_speech.STT.base_stt_handler import BaseSTTHandler
from speech_to_speech.pipeline.handler_types import STTIn, STTOut
from speech_to_speech.pipeline.messages import Transcription

class MySTTHandler(BaseSTTHandler):
    """Example custom STT backend receiving VADAudio and emitting Transcription objects."""
    
    def __init__(self, stop_event, queue_in, queue_out, setup_kwargs):
        super().__init__(stop_event, queue_in, queue_out, setup_kwargs)
        self.model_path = setup_kwargs.get("model_path", "default.pt")
        # Load your model here

    
    def run(self) -> None:
        """Main processing loop called by BaseHandler.run()."""
        while not self.stop_event.is_set():
            audio: STTIn = self.queue_in.get()
            if not self.should_process_input(audio):
                continue
            
            # Replace with actual inference logic

            transcription = Transcription(
                turn_id=audio.turn_id,
                turn_revision=audio.turn_revision,
                created_at_s=audio.created_at_s,
                text="processed transcription",
                mode="final",
            )
            
            if self.should_emit_output(transcription):
                self.queue_out.put(transcription)

```

By inheriting from `BaseSTTHandler`, you automatically receive `should_process_input()` and `should_emit_output()` filters that manage speculative turns and stale audio segments.

### Step 2: Define Configuration Arguments

Create `src/speech_to_speech/arguments_classes/<your_name>_stt_arguments.py` to expose configuration options. Extend `ArgumentsBase` so that `HfArgumentParser` can auto-generate CLI flags.

```python
from dataclasses import dataclass
from speech_to_speech.arguments_classes.module_arguments import ArgumentsBase

@dataclass
class MySTTHandlerArguments(ArgumentsBase):
    """Configuration for the MySTT backend."""
    model_path: str = "models/my_stt.pt"
    device: str = "cpu"
    language: str = "en"

```

This creates typed configuration objects accessible via `--my_stt_model_path`, `--my_stt_device`, and similar flags.

### Step 3: Register the Backend Enum

Modify [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py) to add your identifier to the `STTBackend` enum. This enum currently includes values like `whisper`, `faster-whisper`, `parakeet-tdt`, and `paraformer`.

```python
from enum import Enum

class STTBackend(str, Enum):
    WHISPER = "whisper"
    WHISPER_MLX = "whisper-mlx"
    MLX_AUDIO_WHISPER = "mlx-audio-whisper"
    FASTER_WHISPER = "faster-whisper"
    PARAKEET_TDT = "parakeet-tdt"
    PARAFORMER = "paraformer"
    MY_STT = "my_stt"  # Your custom backend

```

Adding your identifier here enables the `--stt my_stt` CLI option and prevents configuration typos through enum validation.

### Step 4: Wire into the Pipeline Dispatch

Add a dispatch branch in `get_stt_handler()` within [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) (around lines 750–790). Import your handler class and return it wrapped with `with_speculative_turns()` to maintain consistency with built-in backends.

```python
elif module_kwargs.stt == "my_stt":
    from speech_to_speech.STT.my_stt_handler import MySTTHandler
    return with_speculative_turns(
        MySTTHandler(
            stop_event,
            queue_in=spoken_prompt_queue,
            queue_out=text_prompt_queue,
            setup_kwargs=vars(my_stt_handler_kwargs),
        )
    )

```

This branch mirrors the pattern used for Whisper and other built-in handlers, ensuring your backend integrates seamlessly with the pipeline's speculative-turn management.

### Step 5: Handle Argument Parser Registration (Optional)

If your `MySTTHandlerArguments` introduces field names that clash with other argument classes, adjust the "pre-parse" logic in `parse_arguments()` within [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py). Most implementations can simply add the new arguments class to the parser without modification, as `HfArgumentParser` resolves collisions by registration order.

### Step 6: Write Integration Tests

Create [`tests/test_my_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/tests/test_my_stt_handler.py) to validate that the pipeline builds correctly with your backend and that audio flows through your handler to downstream components.

```python
import json
import pathlib
from speech_to_speech.s2s_pipeline import parse_arguments, build_pipeline

def test_my_stt_integration(tmp_path):
    cfg = {
        "module_kwargs": {"stt": "my_stt", "tts": "qwen3", "mode": "local"},
        "my_stt_handler_kwargs": {"model_path": str(tmp_path / "dummy.pt")},
        # Add other required configuration sections...

    }
    
    cfg_path = tmp_path / "cfg.json"
    cfg_path.write_text(json.dumps(cfg))
    
    # Verify pipeline initializes without errors

    args = parse_arguments([f"--config", str(cfg_path)])
    # Additional assertions to verify handler instantiation...

```

Tests ensure future refactoring does not break your integration and document the expected message contract for other developers.

## Key Implementation Details

The `run()` method operates on a polling loop until `stop_event.is_set()` signals shutdown. Input messages implement the `STTIn` protocol (typically `VADAudio` instances containing raw audio data and turn metadata), while output messages must conform to `STTOut` (`PartialTranscription` for interim results or `Transcription` for final results).

The `with_speculative_turns()` wrapper applied in `get_stt_handler()` manages conversation turn revisions automatically. Your handler receives audio segments tagged with `turn_id` and `turn_revision` fields, and the wrapper handles filtering obsolete revisions so your transcription logic can remain stateless regarding turn management.

## Summary

- **Inherit from `BaseSTTHandler`** in [`src/speech_to_speech/STT/base_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/base_stt_handler.py) to leverage built-in speculative-turn filtering and stale-input detection
- **Create an arguments dataclass** extending `ArgumentsBase` in `src/speech_to_speech/arguments_classes/` for type-safe configuration
- **Register your backend identifier** in the `STTBackend` enum in [`module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/module_arguments.py) to enable `--stt <your_id>` CLI usage
- **Add a dispatch branch** in `get_stt_handler()` in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) (lines 749–789) to wire your handler into the pipeline
- **Wrap with `with_speculative_turns()`** to maintain consistency with existing backends and automatic turn revision handling
- **Place tests** in `tests/` to verify integration and prevent regressions

## Frequently Asked Questions

### What methods must a custom STT handler implement?

You must implement `__init__` accepting `stop_event`, `queue_in`, `queue_out`, and `setup_kwargs`, plus a `run()` method that processes the input queue until the stop event is set. The base class provides `should_process_input()` and `should_emit_output()` lifecycle hooks for filtering speculative and stale data.

### How does the pipeline handle speculative turns with custom backends?

The `with_speculative_turns()` wrapper in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) automatically wraps your handler instance when registered in `get_stt_handler()`. This wrapper manages turn revision tracking and filtering without requiring changes to your transcription logic, ensuring that only the latest audio for a given turn reaches downstream components.

### Can I use a completely custom base class instead of BaseSTTHandler?

Yes, you can inherit directly from `BaseHandler[STTIn, STTOut]` if you need to bypass the speculative-turn filtering provided by `BaseSTTHandler`. However, you will need to reimplement stale-input detection and queue management logic yourself, as the base `BaseSTTHandler` provides these protections by default.

### Where should I place my custom handler files in the repository?

Place your handler implementation in `src/speech_to_speech/STT/<your_name>_handler.py` and your arguments dataclass in `src/speech_to_speech/arguments_classes/<your_name>_stt_arguments.py`. This maintains consistency with the existing project structure and ensures import paths resolve correctly when the pipeline loads your module.