Common Speech-to-Speech Errors and Troubleshooting Tips for the Hugging Face Pipeline

Most speech-to-speech pipeline failures stem from platform-device mismatches, missing optional dependencies, or invalid backend configurations that can be resolved by validating arguments against the source code in s2s_pipeline.py.

The Hugging Face speech-to-speech repository implements a modular pipeline connecting voice activity detection (VAD), speech-to-text (STT), language models (LM), and text-to-speech (TTS). Because this architecture supports multiple interchangeable backends, users frequently encounter configuration errors when device settings, quantization formats, or optional dependencies are mismatched.

Platform-Specific Configuration Errors

Invalid CUDA Selection on macOS

The check_mac_settings function in src/speech_to_speech/s2s_pipeline.py explicitly forbids --device cuda on macOS. If you attempt to run with GPU acceleration on Apple hardware, the pipeline raises: "Cannot use CUDA on macOS. Please set the device to 'cpu' or 'mps'."

Fix: Run with --device mps for Apple Silicon or --device cpu for compatibility mode. Verify your module_kwargs.device value in the argument configuration file.

MLX Contention with Multiple Pipelines

On Apple Silicon, the global MLX lock in src/speech_to_speech/utils/mlx_lock.py serializes inference. Running with --num_pipelines greater than 1 triggers the warning: "MLX contention: --num_pipelines=… > 1 on Apple Silicon → disabling live transcription."

Fix: Either reduce --num_pipelines to 1, disable live transcription with --enable_live_transcription false, or migrate to a non-Apple-Silicon platform.

Backend and Quantization Validation Errors

Unsupported Qwen-3 TTS Configurations

The Qwen3TTSHandlerArguments class in src/speech_to_speech/TTS/qwen3_tts_handler.py validates quantization formats and backends strictly. Attempting to use unsupported values raises "Unsupported qwen3_tts_mlx_quantization" or "Unsupported qwen3_tts_backend."

Fix: Select supported combinations such as quantization bf16 and backend mlx. When using the GGUF backend, you must provide both gguf_talker_path and gguf_codec_path together; omitting one triggers a validation error.

TTS Backend Incompatibility on macOS

The default TTS backend is incompatible with Apple Silicon. Users see: "For macOS users, it is recommended to use qwen3 for TTS (pocket and kokoro are also valid options)."

Fix: Explicitly set --tts qwen3, --tts pocket, or --tts kokoro when running on macOS. Alternatively, use --local_mac_optimal_settings to automatically configure device=mps, llm_backend=mlx-lm, and tts=qwen3.

Dependency and Installation Errors

Missing Optional TTS Handlers

The pipeline treats ChatTTS, Pocket, and Kokoro as optional extras. Importing these without installation raises specific import errors:

  • "Error importing ChatTTSHandler. Install it with pip install "speech-to-speech[chattts]""
  • "Pocket TTS is optional. Install it with pip install "speech-to-speech[pocket]""
  • "Kokoro is optional. Install it with pip install "speech-to-speech[kokoro]""

Fix: Install the specific extra:

pip install "speech-to-speech[chattts]"
pip install "speech-to-speech[pocket]"
pip install "speech-to-speech[kokoro]"

Missing Reference Assets

Benchmark scripts expect packaged reference audio files for Qwen-3 TTS that are not included in the base installation. Running benchmarks without these files causes "Packaged Qwen3-TTS reference audio is missing."

Fix: Run scripts/benchmark_tts.py with the --download flag first, or manually add the missing audio files to the assets directory.

Pipeline Validation and Runtime Errors

Invalid Pipeline Count Configuration

The parse_arguments function in src/speech_to_speech/s2s_pipeline.py validates that --num_pipelines must be >= 1, and values > 1 are only supported with --mode realtime. Violating either constraint produces explicit error messages.

Fix: Use --num_pipelines 1 for non-realtime modes, or switch to --mode realtime if parallel processing is required.

Missing Required Arguments

The argument parser raises "Missing required …" when mandatory fields are absent from JSON/YAML configs or command-line flags.

Fix: Validate your configuration against all required fields in parse_arguments. Run python -m speech_to_speech.s2s_pipeline --help to inspect mandatory parameters.

API Provider Unavailability

For the responses-api and chat-completions LM backends, invalid API keys or network issues produce RuntimeError("provider unavailable") as defined in the respective handler arguments classes.

Fix: Verify your API key, upstream URL, and model name in ResponsesApiLanguageModelHandlerArguments. Ensure network connectivity to the provider endpoints.

General Troubleshooting Workflow

When errors persist, follow this diagnostic sequence:

  1. Enable debug logging – Add --log_level debug to see detailed pipeline actions and exception tracebacks handled by src/speech_to_speech/pipeline/log_context.py.

  2. Validate arguments – Use the built-in help flag to verify all required fields are present, especially for optional backends.

  3. Check platform settings – On macOS, use --local_mac_optimal_settings to automatically configure compatible device, LLM, and TTS backends.

  4. Inspect queue connections – For suspected deadlocks, insert diagnostic prints in speech_to_speech/pipeline/_build_pipeline_handlers.py to verify queue creation and handler connections.

Code Examples

Running with recommended macOS settings:

python -m speech_to_speech.s2s_pipeline \
    --mode realtime \
    --device mps \
    --llm_backend mlx-lm \
    --tts qwen3 \
    --log_level info

Validating configuration programmatically:

from speech_to_speech.s2s_pipeline import parse_arguments

args = parse_arguments()
print(args.module_kwargs)
print(args.qwen3_tts_handler_kwargs)

Installing all optional TTS backends:

pip install "speech-to-speech[chattts,pocket,kokoro]"

Summary

  • Platform mismatches cause the most common startup failures; use --device mps on macOS and avoid CUDA on Apple Silicon.
  • Optional dependencies for ChatTTS, Pocket, and Kokoro require bracketed pip installs (speech-to-speech[handler]).
  • Configuration validation occurs in parse_arguments and handler-specific argument classes—always provide both gguf_talker_path and gguf_codec_path when using GGUF backends.
  • MLX contention on Apple Silicon limits parallel pipelines; use --num_pipelines 1 or disable live transcription.
  • Debug logging (--log_level debug) and the help flag are the primary diagnostic tools for resolving argument and runtime errors.

Frequently Asked Questions

Why do I get a CUDA error on my Mac when running speech-to-speech?

The check_mac_settings function in s2s_pipeline.py explicitly blocks CUDA on macOS because Apple hardware uses Metal Performance Shaders (MPS) instead. Set --device mps for Apple Silicon or --device cpu for compatibility. You can also use --local_mac_optimal_settings to automatically configure the correct device and compatible backends.

How do I fix "Missing required argument" errors when using JSON configuration files?

The parse_arguments function validates that all mandatory fields are present. Run python -m speech_to_speech.s2s_pipeline --help to identify required parameters, then verify your JSON or YAML config includes fields like module_kwargs and backend-specific arguments (e.g., qwen3_tts_handler_kwargs). Required fields vary by selected backend.

Can I run multiple parallel pipelines on Apple Silicon?

No, values of --num_pipelines greater than 1 create MLX contention on Apple Silicon due to the global lock in src/speech_to_speech/utils/mlx_lock.py. Either set --num_pipelines 1, disable live transcription with --enable_live_transcription false, or run on Linux/Windows with CUDA support. Multiple pipelines require --mode realtime on any platform.

Why is my TTS backend failing with an import error?

ChatTTS, Pocket, and Kokoro are optional dependencies not installed by default. The handlers in src/speech_to_speech/TTS/ raise import errors with specific install commands. Install the required handler using pip install "speech-to-speech[chattts]" (or [pocket], [kokoro]), or switch to the default qwen3 TTS which requires no extra dependencies.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →