How to Contribute to the HuggingFace Speech-to-Speech Project: A Complete Developer Guide

Contribute to the huggingface/speech-to-speech project by forking the repository, setting up the development environment with uv sync, implementing changes in the modular handler architecture (VAD → STT → LLM → TTS), running tests with pytest, and submitting a focused PR referencing the relevant issue.

The huggingface/speech-to-speech repository implements a low-latency, fully modular voice-assistant pipeline that processes audio through a four-stage architecture. Whether you want to add a new TTS backend, optimize queue handling, or improve documentation, this guide explains exactly how to contribute to the HuggingFace speech-to-speech project using the actual source code structure and contribution workflow.

Understand the Modular Architecture

The pipeline processes audio through four sequential stages: VAD (Voice Activity Detection) → STT (Speech-to-Text) → LLM (Language Model) → TTS (Text-to-Speech). Each stage runs in its own thread and communicates via typed queues.

The architecture centers on three core concepts:

  • Argument classes: Typed dataclasses in src/speech_to_speech/arguments_classes/ that expose every CLI flag and are parsed with HfArgumentParser.
  • Handlers: Subclasses of BaseHandler in src/speech_to_speech/VAD/, STT/, LLM/, and TTS/ directories that encapsulate a single stage, receiving input from a queue and pushing results to the next queue.
  • Pipeline builder: The build_pipeline function in src/speech_to_speech/s2s_pipeline.py orchestrates queue creation and handler chains for each deployment mode.

The system supports four deployment modes: Realtime WebSocket, local microphone, raw WebSocket, and TCP socket. The entry point speech-to-speech CLI parses arguments and invokes the appropriate pipeline configuration.

Set Up Your Development Environment

Start by cloning the repository and installing dependencies in editable mode:

git clone https://github.com/huggingface/speech-to-speech.git
cd speech-to-speech
uv sync

Validate your installation by running the test suite and linting:

pytest
ruff check

These steps correspond to lines 44-51 of the README. Ensure all tests pass before making changes.

Make Your First Contribution

Pick an Issue

Browse the open issues on GitHub to find tasks matching your expertise. Good starting points include documentation improvements, unit test coverage for edge cases, or implementing lightweight backends like a gTTS TTS handler.

Add a New Backend Handler

To add a custom TTS backend, create a new handler class and register it in the pipeline builder.

Step 1: Create your handler in src/speech_to_speech/TTS/my_custom_tts_handler.py:

from speech_to_speech.baseHandler import BaseHandler
from speech_to_speech.pipeline.handler_types import TTSIn, TTSOut

class MyCustomTTSHandler(BaseHandler[TTSIn, TTSOut]):
    def __init__(self, stop_event, queue_in, queue_out, *, voice: str = "default"):
        super().__init__(stop_event, queue_in, queue_out)
        self.voice = voice

    def _run_once(self, item: TTSIn):
        # Convert text to audio using your model (pseudo-code)

        audio_bytes = my_tts_model.synthesize(item.text, voice=self.voice)
        self.queue_out.put(audio_bytes)

Step 2: Add a corresponding arguments dataclass in src/speech_to_speech/arguments_classes/.

Step 3: Register the handler in get_tts_handler within src/speech_to_speech/s2s_pipeline.py (lines 49-63):

elif module_kwargs.tts == "my_custom":
    from speech_to_speech.TTS.my_custom_tts_handler import MyCustomTTSHandler
    return MyCustomTTSHandler(
        stop_event,
        queue_in=lm_response_queue,
        queue_out=send_audio_chunks_queue,
        voice="alice",
    )

Adding an STT backend follows the same pattern: inherit from BaseSTTHandler in src/speech_to_speech/STT/ and register it in get_stt_handler (lines 80-78 of s2s_pipeline.py).

Modify the Pipeline Logic

To change how handlers connect, examine the _build_pipeline_handlers function (lines 78-94) in src/speech_to_speech/s2s_pipeline.py. This method wires the queue flow: VAD → STT → TranscriptionNotifier → LLM → LMOutputProcessor → TTS.

For mode-specific behavior, modify the build_pipeline function (lines 106-140), which selects between local, raw-websocket, realtime, and socket modes.

Test Your Changes

Write unit tests for new handlers using the ThreadManager utility. Here is a test template for a custom STT handler:

def test_my_stt_handler():
    from speech_to_speech.STT.my_stt_handler import MySTTHandler
    from speech_to_speech.utils.thread_manager import ThreadManager
    from queue import Queue
    from threading import Event

    in_q = Queue()
    out_q = Queue()
    stop_evt = Event()
    handler = MySTTHandler(stop_evt, in_q, out_q, model_name="my-model")
    tm = ThreadManager([handler])
    tm.start()
    in_q.put(b"audio data")
    result = out_q.get(timeout=5)
    assert result.text == "expected transcription"
    tm.stop()

Run the examples to verify integration:


# Realtime server mode

speech-to-speech

# In another terminal

python scripts/listen_and_play_realtime.py --host 127.0.0.1 --port 8765

Submit a Pull Request

Follow the standard GitHub workflow to submit your contribution:

  1. Fork the repository and create a feature branch (git checkout -b my-feature).
  2. Implement your changes and ensure pytest passes.
  3. Lint your code with ruff check.
  4. Update documentation if you introduce new CLI options or APIs.
  5. Open a PR targeting the main branch, referencing the related issue (e.g., "Closes #123").
  6. Monitor CI – GitHub Actions automatically runs tests and style checks.

Keep pull requests focused on a single logical change to accelerate review.

Summary

  • The huggingface/speech-to-speech architecture uses modular handlers (VAD → STT → LLM → TTS) connected by typed queues and orchestrated by s2s_pipeline.py.
  • Set up your environment with uv sync, then validate changes using pytest and ruff check.
  • Add new backends by subclassing BaseHandler, creating argument dataclasses, and registering them in get_stt_handler or get_tts_handler.
  • Submit focused PRs that reference specific issues and pass continuous integration tests.

Frequently Asked Questions

What is the architecture of the speech-to-speech pipeline?

The pipeline implements a four-stage processing chain: VAD (Voice Activity Detection) → STT (Speech-to-Text) → LLM (Language Model) → TTS (Text-to-Speech). Each stage runs as an independent thread managed by ThreadManager, with data flowing through typed queues. The build_pipeline function in src/speech_to_speech/s2s_pipeline.py constructs this chain based on the selected deployment mode.

How do I add a new speech-to-text backend to the project?

Create a new file in src/speech_to_speech/STT/ that inherits from BaseSTTHandler, define a corresponding arguments dataclass in src/speech_to_speech/arguments_classes/, and register the backend in the get_stt_handler function within src/speech_to_speech/s2s_pipeline.py (around lines 80-78). Follow the existing pattern used by parakeet_tdt_handler.py.

What testing tools does the project use?

The project uses pytest for unit testing and ruff for linting. After installing dependencies with uv sync, run pytest to execute the test suite and ruff check to verify code style. Integration testing can be performed by running the Realtime server and connecting with scripts/listen_and_play_realtime.py.

Do I need to open an issue before submitting a pull request?

For small fixes and documentation improvements, you can submit a PR directly. For larger architectural changes or new features, open an issue first to discuss the approach. As noted in the README (lines 72-75): "For larger changes, open an issue first to discuss the approach." Always reference related issues in your PR description using keywords like "Closes #123".

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →