Common Speech-to-Speech Errors and Troubleshooting Tips for the Hugging Face Pipeline
Most speech-to-speech pipeline failures stem from platform-device mismatches, missing optional dependencies, or invalid backend configurations that can be resolved by validating arguments against the source code in s2s_pipeline.py.
The Hugging Face speech-to-speech repository implements a modular pipeline connecting voice activity detection (VAD), speech-to-text (STT), language models (LM), and text-to-speech (TTS). Because this architecture supports multiple interchangeable backends, users frequently encounter configuration errors when device settings, quantization formats, or optional dependencies are mismatched.
Platform-Specific Configuration Errors
Invalid CUDA Selection on macOS
The check_mac_settings function in src/speech_to_speech/s2s_pipeline.py explicitly forbids --device cuda on macOS. If you attempt to run with GPU acceleration on Apple hardware, the pipeline raises: "Cannot use CUDA on macOS. Please set the device to 'cpu' or 'mps'."
Fix: Run with --device mps for Apple Silicon or --device cpu for compatibility mode. Verify your module_kwargs.device value in the argument configuration file.
MLX Contention with Multiple Pipelines
On Apple Silicon, the global MLX lock in src/speech_to_speech/utils/mlx_lock.py serializes inference. Running with --num_pipelines greater than 1 triggers the warning: "MLX contention: --num_pipelines=… > 1 on Apple Silicon → disabling live transcription."
Fix: Either reduce --num_pipelines to 1, disable live transcription with --enable_live_transcription false, or migrate to a non-Apple-Silicon platform.
Backend and Quantization Validation Errors
Unsupported Qwen-3 TTS Configurations
The Qwen3TTSHandlerArguments class in src/speech_to_speech/TTS/qwen3_tts_handler.py validates quantization formats and backends strictly. Attempting to use unsupported values raises "Unsupported qwen3_tts_mlx_quantization" or "Unsupported qwen3_tts_backend."
Fix: Select supported combinations such as quantization bf16 and backend mlx. When using the GGUF backend, you must provide both gguf_talker_path and gguf_codec_path together; omitting one triggers a validation error.
TTS Backend Incompatibility on macOS
The default TTS backend is incompatible with Apple Silicon. Users see: "For macOS users, it is recommended to use qwen3 for TTS (pocket and kokoro are also valid options)."
Fix: Explicitly set --tts qwen3, --tts pocket, or --tts kokoro when running on macOS. Alternatively, use --local_mac_optimal_settings to automatically configure device=mps, llm_backend=mlx-lm, and tts=qwen3.
Dependency and Installation Errors
Missing Optional TTS Handlers
The pipeline treats ChatTTS, Pocket, and Kokoro as optional extras. Importing these without installation raises specific import errors:
- "Error importing ChatTTSHandler. Install it with
pip install "speech-to-speech[chattts]"" - "Pocket TTS is optional. Install it with
pip install "speech-to-speech[pocket]"" - "Kokoro is optional. Install it with
pip install "speech-to-speech[kokoro]""
Fix: Install the specific extra:
pip install "speech-to-speech[chattts]"
pip install "speech-to-speech[pocket]"
pip install "speech-to-speech[kokoro]"
Missing Reference Assets
Benchmark scripts expect packaged reference audio files for Qwen-3 TTS that are not included in the base installation. Running benchmarks without these files causes "Packaged Qwen3-TTS reference audio is missing."
Fix: Run scripts/benchmark_tts.py with the --download flag first, or manually add the missing audio files to the assets directory.
Pipeline Validation and Runtime Errors
Invalid Pipeline Count Configuration
The parse_arguments function in src/speech_to_speech/s2s_pipeline.py validates that --num_pipelines must be >= 1, and values > 1 are only supported with --mode realtime. Violating either constraint produces explicit error messages.
Fix: Use --num_pipelines 1 for non-realtime modes, or switch to --mode realtime if parallel processing is required.
Missing Required Arguments
The argument parser raises "Missing required …" when mandatory fields are absent from JSON/YAML configs or command-line flags.
Fix: Validate your configuration against all required fields in parse_arguments. Run python -m speech_to_speech.s2s_pipeline --help to inspect mandatory parameters.
API Provider Unavailability
For the responses-api and chat-completions LM backends, invalid API keys or network issues produce RuntimeError("provider unavailable") as defined in the respective handler arguments classes.
Fix: Verify your API key, upstream URL, and model name in ResponsesApiLanguageModelHandlerArguments. Ensure network connectivity to the provider endpoints.
General Troubleshooting Workflow
When errors persist, follow this diagnostic sequence:
-
Enable debug logging – Add
--log_level debugto see detailed pipeline actions and exception tracebacks handled bysrc/speech_to_speech/pipeline/log_context.py. -
Validate arguments – Use the built-in help flag to verify all required fields are present, especially for optional backends.
-
Check platform settings – On macOS, use
--local_mac_optimal_settingsto automatically configure compatible device, LLM, and TTS backends. -
Inspect queue connections – For suspected deadlocks, insert diagnostic prints in
speech_to_speech/pipeline/_build_pipeline_handlers.pyto verify queue creation and handler connections.
Code Examples
Running with recommended macOS settings:
python -m speech_to_speech.s2s_pipeline \
--mode realtime \
--device mps \
--llm_backend mlx-lm \
--tts qwen3 \
--log_level info
Validating configuration programmatically:
from speech_to_speech.s2s_pipeline import parse_arguments
args = parse_arguments()
print(args.module_kwargs)
print(args.qwen3_tts_handler_kwargs)
Installing all optional TTS backends:
pip install "speech-to-speech[chattts,pocket,kokoro]"
Summary
- Platform mismatches cause the most common startup failures; use
--device mpson macOS and avoid CUDA on Apple Silicon. - Optional dependencies for ChatTTS, Pocket, and Kokoro require bracketed pip installs (
speech-to-speech[handler]). - Configuration validation occurs in
parse_argumentsand handler-specific argument classes—always provide bothgguf_talker_pathandgguf_codec_pathwhen using GGUF backends. - MLX contention on Apple Silicon limits parallel pipelines; use
--num_pipelines 1or disable live transcription. - Debug logging (
--log_level debug) and the help flag are the primary diagnostic tools for resolving argument and runtime errors.
Frequently Asked Questions
Why do I get a CUDA error on my Mac when running speech-to-speech?
The check_mac_settings function in s2s_pipeline.py explicitly blocks CUDA on macOS because Apple hardware uses Metal Performance Shaders (MPS) instead. Set --device mps for Apple Silicon or --device cpu for compatibility. You can also use --local_mac_optimal_settings to automatically configure the correct device and compatible backends.
How do I fix "Missing required argument" errors when using JSON configuration files?
The parse_arguments function validates that all mandatory fields are present. Run python -m speech_to_speech.s2s_pipeline --help to identify required parameters, then verify your JSON or YAML config includes fields like module_kwargs and backend-specific arguments (e.g., qwen3_tts_handler_kwargs). Required fields vary by selected backend.
Can I run multiple parallel pipelines on Apple Silicon?
No, values of --num_pipelines greater than 1 create MLX contention on Apple Silicon due to the global lock in src/speech_to_speech/utils/mlx_lock.py. Either set --num_pipelines 1, disable live transcription with --enable_live_transcription false, or run on Linux/Windows with CUDA support. Multiple pipelines require --mode realtime on any platform.
Why is my TTS backend failing with an import error?
ChatTTS, Pocket, and Kokoro are optional dependencies not installed by default. The handlers in src/speech_to_speech/TTS/ raise import errors with specific install commands. Install the required handler using pip install "speech-to-speech[chattts]" (or [pocket], [kokoro]), or switch to the default qwen3 TTS which requires no extra dependencies.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →