How to Swap Components in the Speech-to-Speech Pipeline: Complete Backend Customization Guide
Swapping components in the Hugging Face Speech-to-Speech pipeline requires only command-line flags or JSON configuration—no code changes needed—thanks to a modular backend registry system.
The Speech-to-Speech repository (huggingface/speech-to-speech) implements a pluggable pipeline architecture where each major stage (STT, LLM, TTS) is a swappable backend. Understanding how to swap these components lets you optimize for latency, quality, or hardware constraints without touching the core pipeline logic.
Understanding the Pipeline Architecture
The system processes audio through a fixed handler chain:
VAD → STT → (optional) TranscriptionNotifier → LLM → LMOutputProcessor → TTS
Each swappable stage conforms to a common interface and is registered in speech_to_speech/backend_registry.py. The registries STT_BACKENDS, LLM_BACKENDS, and TTS_BACKENDS map string names to factory functions, enabling runtime selection.
How Backend Selection Works
The pipeline uses a three-step resolution process:
- Argument parsing —
speech_to_speech/s2s_pipeline.py::parse_arguments()(lines 70-104) reads--stt,--llm_backend, and--ttsflags - Registry lookup —
backend_registry.py::select_backend()(lines 62-68) converts names toBackendSelectionobjects with normalized configuration viaBackendSpec.normalize - Handler instantiation —
backend_registry.py::create_backend_handler()(lines 82-88) executes the factory to build the concrete handler - Chain assembly —
s2s_pipeline.py::_build_handlers()(lines 48-69) inserts handlers into the runtime queue structure
Because the registries use factory functions, adding or swapping backends requires zero changes to s2s_pipeline.py.
Swapping Components via Command Line
The simplest method uses flags when starting the server:
# Replace STT engine
speech-to-speech serve --stt whisper
# Use different language model
speech-to-speech serve --llm_backend mlx-lm
# Change TTS engine
speech-to-speech serve --tts qwen3
Combine multiple swaps in one command:
speech-to-speech serve \
--stt paraformer \
--llm_backend responses-api \
--tts pocket \
--log-level debug
Swapping Components via JSON Configuration
For reproducible deployments, use a JSON configuration file:
{
"stt": "faster-whisper",
"llm_backend": "chat-completions",
"tts": "chatTTS",
"stt_device": "cuda",
"llm_device": "cuda",
"tts_device": "cuda",
"log_level": "info"
}
Launch with:
speech-to-speech serve pipeline.json
The parse_arguments() function detects JSON paths through its _is_json branch and merges settings with platform defaults.
Swapping Components Programmatically
For dynamic configuration in Python applications:
from speech_to_speech.s2s_pipeline import parse_arguments, run_pipeline_command
# Configure paraformer STT and kokoro TTS
cli_args = ["--stt", "parakeet-tdt", "--tts", "kokoro"]
parsed = parse_arguments(cli_args, command="serve")
# Verify selections
print(parsed.stt_backend.name) # "parakeet-tdt"
print(parsed.llm_backend.name) # default value
print(parsed.tts_backend.name) # "kokoro"
# Execute pipeline
run_pipeline_command("serve", cli_args)
This approach enables runtime backend selection based on user preferences or hardware detection.
Available Backend Options
The registry supports these backend categories (see backend_registry.py for current registrations):
STT Backends:
whisper— OpenAI Whisperwhisper-mlx— Apple Silicon optimizedfaster-whisper— Faster Whisper implementationparaformer— Alibaba Paraformerparakeet-tdt— NVIDIA Parakeet
LLM Backends:
mlx-lm— Apple MLX frameworktransformers— Hugging Face Transformerschat-completions— OpenAI-compatible APIresponses-api— Responses API interface
TTS Backends:
qwen3— Qwen3 TTSchatTTS— ChatTTS modelfacebookMMS— Meta MMSpocket— PocketSynthesiskokoro— Kokoro TTS
Configuration File with Backend-Specific Options
Some backends require additional parameters:
{
"stt": "facebookMMS",
"llm_backend": "transformers",
"tts": "facebookMMS",
"facebook_mms_tts_language": "en",
"stt_device": "cpu",
"llm_device": "cuda:0",
"tts_device": "cpu",
"log_level": "debug"
}
Backend-specific arguments are defined in speech_to_speech/arguments_classes/ (e.g., whisper_stt_arguments.py).
Complete Working Examples
Full CLI swap with debugging:
speech-to-speech serve \
--stt whisper-mlx \
--llm_backend mlx-lm \
--tts qwen3 \
--log-level debug
JSON configuration for multilingual deployment:
cat > pipeline.json <<'EOF'
{
"stt": "paraformer",
"llm_backend": "transformers",
"tts": "facebookMMS",
"facebook_mms_tts_language": "en",
"log_level": "info"
}
EOF
speech-to-speech serve pipeline.json
Python API for conditional backend selection:
from speech_to_speech.s2s_pipeline import parse_arguments, run_pipeline_command
import torch
# Auto-select based on availability
stt_choice = "whisper-mlx" if torch.backends.mps.is_available() else "faster-whisper"
cli_args = ["--stt", stt_choice, "--llm_backend", "mlx-lm"]
run_pipeline_command("serve", cli_args)
Key Source Files
| File | Purpose |
|---|---|
src/speech_to_speech/s2s_pipeline.py |
Pipeline orchestration, argument parsing, handler chain construction |
src/speech_to_speech/backend_registry.py |
Backend declarations, factory registration, configuration normalization |
src/speech_to_speech/arguments_classes/ |
Typed dataclasses for backend-specific options |
src/speech_to_speech/STT/*.py |
Concrete STT implementations (e.g., whisper_stt_handler.py) |
src/speech_to_speech/LLM/*.py |
Concrete LLM implementations (e.g., language_model.py) |
src/speech_to_speech/TTS/*.py |
Concrete TTS implementations (e.g., qwen3_tts_handler.py) |
Summary
- Command-line flags (
--stt,--llm_backend,--tts) provide instant backend swapping without code changes - JSON configuration files enable reproducible, version-controlled pipeline deployments
- Python API (
parse_arguments,run_pipeline_command) supports dynamic, conditional backend selection - Registry-based architecture in
backend_registry.pydecouples backend implementations from pipeline orchestration - Factory pattern allows new backends to be added by registering name-to-constructor mappings
Frequently Asked Questions
What backends are available in the Speech-to-Speech pipeline?
According to the backend_registry.py source code, available backends include STT options (whisper, faster-whisper, paraformer, parakeet-tdt), LLM options (mlx-lm, transformers, chat-completions, responses-api), and TTS options (qwen3, chatTTS, facebookMMS, pocket, kokoro). Run speech-to-speech serve --help to see current registrations for your installation.
Can I use different hardware devices for each pipeline stage?
Yes. The JSON configuration format accepts stt_device, llm_device, and tts_device fields to assign specific GPUs or CPU for each component. The parse_arguments() function passes these values through BackendSpec.normalize to handler constructors.
How do I add a custom backend to the pipeline?
Create a handler class following the existing patterns in src/speech_to_speech/STT/, LLM/, or TTS/, then register it in backend_registry.py by adding an entry to the appropriate *_BACKENDS dictionary with a factory function. No changes to s2s_pipeline.py are required.
Why does my JSON configuration merge with defaults instead of replacing them?
The parse_arguments() function in s2s_pipeline.py applies platform presets (notably macOS optimizations) before merging your JSON values. To override completely, explicitly set all relevant fields in your configuration file rather than relying on selective specification.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →