# How to Contribute to the HuggingFace Speech-to-Speech Project: A Complete Developer Guide

> Learn how to contribute to the HuggingFace Speech-to-Speech project. Fork the repo, set up your environment, implement changes, run tests, and submit a PR. Your guide to impactful contributions.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-08-01

---

**Contribute to the huggingface/speech-to-speech project by forking the repository, setting up the development environment with `uv sync`, implementing changes in the modular handler architecture (VAD → STT → LLM → TTS), running tests with `pytest`, and submitting a focused PR referencing the relevant issue.**

The huggingface/speech-to-speech repository implements a low-latency, fully modular voice-assistant pipeline that processes audio through a four-stage architecture. Whether you want to add a new TTS backend, optimize queue handling, or improve documentation, this guide explains exactly how to contribute to the HuggingFace speech-to-speech project using the actual source code structure and contribution workflow.

## Understand the Modular Architecture

The pipeline processes audio through four sequential stages: **VAD** (Voice Activity Detection) → **STT** (Speech-to-Text) → **LLM** (Language Model) → **TTS** (Text-to-Speech). Each stage runs in its own thread and communicates via typed queues.

The architecture centers on three core concepts:

- **Argument classes**: Typed dataclasses in `src/speech_to_speech/arguments_classes/` that expose every CLI flag and are parsed with `HfArgumentParser`.
- **Handlers**: Subclasses of `BaseHandler` in `src/speech_to_speech/VAD/`, `STT/`, `LLM/`, and `TTS/` directories that encapsulate a single stage, receiving input from a queue and pushing results to the next queue.
- **Pipeline builder**: The `build_pipeline` function in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) orchestrates queue creation and handler chains for each deployment mode.

The system supports four deployment modes: **Realtime WebSocket**, **local microphone**, **raw WebSocket**, and **TCP socket**. The entry point `speech-to-speech` CLI parses arguments and invokes the appropriate pipeline configuration.

## Set Up Your Development Environment

Start by cloning the repository and installing dependencies in editable mode:

```bash
git clone https://github.com/huggingface/speech-to-speech.git
cd speech-to-speech
uv sync

```

Validate your installation by running the test suite and linting:

```bash
pytest
ruff check

```

These steps correspond to lines 44-51 of the README. Ensure all tests pass before making changes.

## Make Your First Contribution

### Pick an Issue

Browse the open issues on GitHub to find tasks matching your expertise. Good starting points include documentation improvements, unit test coverage for edge cases, or implementing lightweight backends like a `gTTS` TTS handler.

### Add a New Backend Handler

To add a custom TTS backend, create a new handler class and register it in the pipeline builder.

**Step 1**: Create your handler in [`src/speech_to_speech/TTS/my_custom_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/my_custom_tts_handler.py):

```python
from speech_to_speech.baseHandler import BaseHandler
from speech_to_speech.pipeline.handler_types import TTSIn, TTSOut

class MyCustomTTSHandler(BaseHandler[TTSIn, TTSOut]):
    def __init__(self, stop_event, queue_in, queue_out, *, voice: str = "default"):
        super().__init__(stop_event, queue_in, queue_out)
        self.voice = voice

    def _run_once(self, item: TTSIn):
        # Convert text to audio using your model (pseudo-code)

        audio_bytes = my_tts_model.synthesize(item.text, voice=self.voice)
        self.queue_out.put(audio_bytes)

```

**Step 2**: Add a corresponding arguments dataclass in `src/speech_to_speech/arguments_classes/`.

**Step 3**: Register the handler in `get_tts_handler` within [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) (lines 49-63):

```python
elif module_kwargs.tts == "my_custom":
    from speech_to_speech.TTS.my_custom_tts_handler import MyCustomTTSHandler
    return MyCustomTTSHandler(
        stop_event,
        queue_in=lm_response_queue,
        queue_out=send_audio_chunks_queue,
        voice="alice",
    )

```

Adding an STT backend follows the same pattern: inherit from `BaseSTTHandler` in `src/speech_to_speech/STT/` and register it in `get_stt_handler` (lines 80-78 of [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py)).

### Modify the Pipeline Logic

To change how handlers connect, examine the `_build_pipeline_handlers` function (lines 78-94) in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py). This method wires the queue flow: VAD → STT → TranscriptionNotifier → LLM → LMOutputProcessor → TTS.

For mode-specific behavior, modify the `build_pipeline` function (lines 106-140), which selects between `local`, `raw-websocket`, `realtime`, and `socket` modes.

## Test Your Changes

Write unit tests for new handlers using the `ThreadManager` utility. Here is a test template for a custom STT handler:

```python
def test_my_stt_handler():
    from speech_to_speech.STT.my_stt_handler import MySTTHandler
    from speech_to_speech.utils.thread_manager import ThreadManager
    from queue import Queue
    from threading import Event

    in_q = Queue()
    out_q = Queue()
    stop_evt = Event()
    handler = MySTTHandler(stop_evt, in_q, out_q, model_name="my-model")
    tm = ThreadManager([handler])
    tm.start()
    in_q.put(b"audio data")
    result = out_q.get(timeout=5)
    assert result.text == "expected transcription"
    tm.stop()

```

Run the examples to verify integration:

```bash

# Realtime server mode

speech-to-speech

# In another terminal

python scripts/listen_and_play_realtime.py --host 127.0.0.1 --port 8765

```

## Submit a Pull Request

Follow the standard GitHub workflow to submit your contribution:

1. **Fork** the repository and create a feature branch (`git checkout -b my-feature`).
2. **Implement** your changes and ensure `pytest` passes.
3. **Lint** your code with `ruff check`.
4. **Update** documentation if you introduce new CLI options or APIs.
5. **Open a PR** targeting the `main` branch, referencing the related issue (e.g., "Closes #123").
6. **Monitor CI** – GitHub Actions automatically runs tests and style checks.

Keep pull requests focused on a single logical change to accelerate review.

## Summary

- The **huggingface/speech-to-speech** architecture uses modular handlers (VAD → STT → LLM → TTS) connected by typed queues and orchestrated by [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py).
- Set up your environment with `uv sync`, then validate changes using `pytest` and `ruff check`.
- Add new backends by subclassing `BaseHandler`, creating argument dataclasses, and registering them in `get_stt_handler` or `get_tts_handler`.
- Submit focused PRs that reference specific issues and pass continuous integration tests.

## Frequently Asked Questions

### What is the architecture of the speech-to-speech pipeline?

The pipeline implements a four-stage processing chain: **VAD** (Voice Activity Detection) → **STT** (Speech-to-Text) → **LLM** (Language Model) → **TTS** (Text-to-Speech). Each stage runs as an independent thread managed by `ThreadManager`, with data flowing through typed queues. The `build_pipeline` function in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) constructs this chain based on the selected deployment mode.

### How do I add a new speech-to-text backend to the project?

Create a new file in `src/speech_to_speech/STT/` that inherits from `BaseSTTHandler`, define a corresponding arguments dataclass in `src/speech_to_speech/arguments_classes/`, and register the backend in the `get_stt_handler` function within [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) (around lines 80-78). Follow the existing pattern used by [`parakeet_tdt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/parakeet_tdt_handler.py).

### What testing tools does the project use?

The project uses **pytest** for unit testing and **ruff** for linting. After installing dependencies with `uv sync`, run `pytest` to execute the test suite and `ruff check` to verify code style. Integration testing can be performed by running the Realtime server and connecting with [`scripts/listen_and_play_realtime.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/listen_and_play_realtime.py).

### Do I need to open an issue before submitting a pull request?

For small fixes and documentation improvements, you can submit a PR directly. For larger architectural changes or new features, open an issue first to discuss the approach. As noted in the README (lines 72-75): "For larger changes, open an issue first to discuss the approach." Always reference related issues in your PR description using keywords like "Closes #123".