# Common Speech-to-Speech Errors and Troubleshooting Tips for the Hugging Face Pipeline

> Resolve common speech-to-speech errors with our expert troubleshooting guide. Fix platform-device mismatches, dependency issues, and backend configurations for seamless Hugging Face pipeline performance.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-08-01

---

**Most speech-to-speech pipeline failures stem from platform-device mismatches, missing optional dependencies, or invalid backend configurations that can be resolved by validating arguments against the source code in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py).**

The Hugging Face `speech-to-speech` repository implements a modular pipeline connecting voice activity detection (VAD), speech-to-text (STT), language models (LM), and text-to-speech (TTS). Because this architecture supports multiple interchangeable backends, users frequently encounter configuration errors when device settings, quantization formats, or optional dependencies are mismatched.

## Platform-Specific Configuration Errors

### Invalid CUDA Selection on macOS

The `check_mac_settings` function in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) explicitly forbids `--device cuda` on macOS. If you attempt to run with GPU acceleration on Apple hardware, the pipeline raises: "Cannot use CUDA on macOS. Please set the device to 'cpu' or 'mps'."

**Fix:** Run with `--device mps` for Apple Silicon or `--device cpu` for compatibility mode. Verify your `module_kwargs.device` value in the argument configuration file.

### MLX Contention with Multiple Pipelines

On Apple Silicon, the global MLX lock in [`src/speech_to_speech/utils/mlx_lock.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/mlx_lock.py) serializes inference. Running with `--num_pipelines` greater than 1 triggers the warning: "MLX contention: --num_pipelines=… > 1 on Apple Silicon → disabling live transcription."

**Fix:** Either reduce `--num_pipelines` to 1, disable live transcription with `--enable_live_transcription false`, or migrate to a non-Apple-Silicon platform.

## Backend and Quantization Validation Errors

### Unsupported Qwen-3 TTS Configurations

The `Qwen3TTSHandlerArguments` class in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) validates quantization formats and backends strictly. Attempting to use unsupported values raises "Unsupported qwen3_tts_mlx_quantization" or "Unsupported qwen3_tts_backend."

**Fix:** Select supported combinations such as quantization `bf16` and backend `mlx`. When using the GGUF backend, you must provide both `gguf_talker_path` and `gguf_codec_path` together; omitting one triggers a validation error.

### TTS Backend Incompatibility on macOS

The default TTS backend is incompatible with Apple Silicon. Users see: "For macOS users, it is recommended to use qwen3 for TTS (pocket and kokoro are also valid options)."

**Fix:** Explicitly set `--tts qwen3`, `--tts pocket`, or `--tts kokoro` when running on macOS. Alternatively, use `--local_mac_optimal_settings` to automatically configure `device=mps`, `llm_backend=mlx-lm`, and `tts=qwen3`.

## Dependency and Installation Errors

### Missing Optional TTS Handlers

The pipeline treats ChatTTS, Pocket, and Kokoro as optional extras. Importing these without installation raises specific import errors:

- "Error importing ChatTTSHandler. Install it with `pip install "speech-to-speech[chattts]"`"
- "Pocket TTS is optional. Install it with `pip install "speech-to-speech[pocket]"`"
- "Kokoro is optional. Install it with `pip install "speech-to-speech[kokoro]"`"

**Fix:** Install the specific extra:

```bash
pip install "speech-to-speech[chattts]"
pip install "speech-to-speech[pocket]"
pip install "speech-to-speech[kokoro]"

```

### Missing Reference Assets

Benchmark scripts expect packaged reference audio files for Qwen-3 TTS that are not included in the base installation. Running benchmarks without these files causes "Packaged Qwen3-TTS reference audio is missing."

**Fix:** Run [`scripts/benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_tts.py) with the `--download` flag first, or manually add the missing audio files to the assets directory.

## Pipeline Validation and Runtime Errors

### Invalid Pipeline Count Configuration

The `parse_arguments` function in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) validates that `--num_pipelines` must be >= 1, and values > 1 are only supported with `--mode realtime`. Violating either constraint produces explicit error messages.

**Fix:** Use `--num_pipelines 1` for non-realtime modes, or switch to `--mode realtime` if parallel processing is required.

### Missing Required Arguments

The argument parser raises "Missing required …" when mandatory fields are absent from JSON/YAML configs or command-line flags.

**Fix:** Validate your configuration against all required fields in `parse_arguments`. Run `python -m speech_to_speech.s2s_pipeline --help` to inspect mandatory parameters.

### API Provider Unavailability

For the `responses-api` and `chat-completions` LM backends, invalid API keys or network issues produce `RuntimeError("provider unavailable")` as defined in the respective handler arguments classes.

**Fix:** Verify your API key, upstream URL, and model name in `ResponsesApiLanguageModelHandlerArguments`. Ensure network connectivity to the provider endpoints.

## General Troubleshooting Workflow

When errors persist, follow this diagnostic sequence:

1. **Enable debug logging** – Add `--log_level debug` to see detailed pipeline actions and exception tracebacks handled by [`src/speech_to_speech/pipeline/log_context.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/log_context.py).

2. **Validate arguments** – Use the built-in help flag to verify all required fields are present, especially for optional backends.

3. **Check platform settings** – On macOS, use `--local_mac_optimal_settings` to automatically configure compatible device, LLM, and TTS backends.

4. **Inspect queue connections** – For suspected deadlocks, insert diagnostic prints in [`speech_to_speech/pipeline/_build_pipeline_handlers.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/pipeline/_build_pipeline_handlers.py) to verify queue creation and handler connections.

## Code Examples

**Running with recommended macOS settings:**

```bash
python -m speech_to_speech.s2s_pipeline \
    --mode realtime \
    --device mps \
    --llm_backend mlx-lm \
    --tts qwen3 \
    --log_level info

```

**Validating configuration programmatically:**

```python
from speech_to_speech.s2s_pipeline import parse_arguments

args = parse_arguments()
print(args.module_kwargs)
print(args.qwen3_tts_handler_kwargs)

```

**Installing all optional TTS backends:**

```bash
pip install "speech-to-speech[chattts,pocket,kokoro]"

```

## Summary

- **Platform mismatches** cause the most common startup failures; use `--device mps` on macOS and avoid CUDA on Apple Silicon.
- **Optional dependencies** for ChatTTS, Pocket, and Kokoro require bracketed pip installs (`speech-to-speech[handler]`).
- **Configuration validation** occurs in `parse_arguments` and handler-specific argument classes—always provide both `gguf_talker_path` and `gguf_codec_path` when using GGUF backends.
- **MLX contention** on Apple Silicon limits parallel pipelines; use `--num_pipelines 1` or disable live transcription.
- **Debug logging** (`--log_level debug`) and the help flag are the primary diagnostic tools for resolving argument and runtime errors.

## Frequently Asked Questions

### Why do I get a CUDA error on my Mac when running speech-to-speech?

The `check_mac_settings` function in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) explicitly blocks CUDA on macOS because Apple hardware uses Metal Performance Shaders (MPS) instead. Set `--device mps` for Apple Silicon or `--device cpu` for compatibility. You can also use `--local_mac_optimal_settings` to automatically configure the correct device and compatible backends.

### How do I fix "Missing required argument" errors when using JSON configuration files?

The `parse_arguments` function validates that all mandatory fields are present. Run `python -m speech_to_speech.s2s_pipeline --help` to identify required parameters, then verify your JSON or YAML config includes fields like `module_kwargs` and backend-specific arguments (e.g., `qwen3_tts_handler_kwargs`). Required fields vary by selected backend.

### Can I run multiple parallel pipelines on Apple Silicon?

No, values of `--num_pipelines` greater than 1 create MLX contention on Apple Silicon due to the global lock in [`src/speech_to_speech/utils/mlx_lock.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/mlx_lock.py). Either set `--num_pipelines 1`, disable live transcription with `--enable_live_transcription false`, or run on Linux/Windows with CUDA support. Multiple pipelines require `--mode realtime` on any platform.

### Why is my TTS backend failing with an import error?

ChatTTS, Pocket, and Kokoro are optional dependencies not installed by default. The handlers in `src/speech_to_speech/TTS/` raise import errors with specific install commands. Install the required handler using `pip install "speech-to-speech[chattts]"` (or `[pocket]`, `[kokoro]`), or switch to the default `qwen3` TTS which requires no extra dependencies.