# How to Configure the Qwen3-TTS Backend: GGML vs Torch on Linux, Windows, and macOS

> Master the Qwen3-TTS backend! Learn to configure GGML vs Torch on Linux, Windows, and macOS. Optimize your speech-to-speech setup easily.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-30

---

**On Linux and Windows, the Qwen3-TTS handler defaults to the GGML backend but allows switching to Torch via the `--qwen3_tts_backend` flag, while macOS (Apple Silicon) automatically uses the MLX-Audio stack and ignores backend selection entirely.**

The `huggingface/speech-to-speech` repository provides a flexible text-to-speech pipeline supporting multiple inference runtimes for the Qwen3-TTS model. Configuring the correct backend—whether GGML for low-latency CPU inference or Torch for CUDA-optimized GPU performance—requires understanding the platform-specific defaults implemented in the handler source code. The system automatically selects compatible runtimes based on the host operating system, with explicit overrides available only for non-macOS platforms.

## Platform-Specific Backend Defaults

The `Qwen3TTSHandler` class enforces distinct backend strategies depending on the host platform, hardcoded in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) (lines 84-86).

### macOS (Apple Silicon)

On macOS systems (`sys.platform == "darwin"`), the handler **forces the MLX-Audio stack** regardless of user input. The constructor explicitly sets `self.backend = "mlx"`, causing the `qwen3_tts_backend` argument to be silently ignored. You cannot configure GGML or Torch backends on Apple Silicon; instead, control quantization levels via the `qwen3_tts_mlx_quantization` parameter (supporting `bf16`, `4bit`, `6bit`, or `8bit`).

### Linux and Windows (Non-macOS)

On Linux and Windows systems, the handler defaults to **faster-qwen3-tts with the GGML backend**, setting `self.backend = "faster_qwen3_tts"` and `self.faster_backend = "ggml"`. Users can explicitly override this selection to use the Torch (CUDA-graphs) implementation by passing `torch` as the backend value. The validator allows only `("ggml", "torch")` as permitted values, raising a `ValueError` for invalid inputs (lines 42-48).

## How Backend Selection Works Internally

The handler implements a three-stage initialization process that determines which native runtime libraries load into memory.

### Platform Detection

During `__init__`, the handler checks `sys.platform` to determine the execution path. If the platform equals `"darwin"`, it routes to the MLX audio pipeline; otherwise, it initializes the faster-qwen3-tts framework (lines 84-86).

### Backend Normalization

For non-macOS platforms, the `_normalize_faster_backend` method processes the user-provided string. This validation occurs against the `VALID_FASTER_BACKENDS` tuple defined at module level, ensuring only `"ggml"` or `"torch"` propagate to the model loader (lines 42-48). Invalid values trigger an immediate exception at line 246.

### Model Loading Path Divergence

The `setup()` method branches based on `self.backend`:

- **MLX Path**: Invokes `_setup_mlx()`, loading the model via `mlx-audio` with quantization specified by `qwen3_tts_mlx_quantization`.
- **Faster-Qwen3-TTS Path**: Invokes `_setup_faster()`, which instantiates `FasterQwen3TTS.from_pretrained(..., backend=self.faster_backend)` (lines 98-112). The `backend` argument passed here determines whether the underlying C++ GGML bindings or the PyTorch CUDA-graphs implementation execute inference.

## Configuration Options: CLI and Python API

Configuration interfaces reside in [`src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py) (lines 31-35), defining `qwen3_tts_backend` as a `Literal["ggml", "torch"]` with a default value of `"ggml"`.

### Command-Line Interface

Pass the `--qwen3_tts_backend` flag when launching the speech-to-speech pipeline:

```bash

# Linux / Windows: Default GGML backend (optimal for CPU/low-latency inference)

python s2s_pipeline.py \
  --tts qwen3 \
  --qwen3_tts_model_name Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice \
  --qwen3_tts_device cuda \
  --qwen3_tts_backend ggml \
  --qwen3_tts_speaker Aiden

# Linux / Windows: Switch to Torch CUDA-graphs backend

python s2s_pipeline.py \
  --tts qwen3 \
  --qwen3_tts_model_name Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice \
  --qwen3_tts_device cuda \
  --qwen3_tts_backend torch \
  --qwen3_tts_speaker Aiden

```

*Note:* On macOS, the `--qwen3_tts_backend` flag is ignored; use `--qwen3_tts_mlx_quantization` instead to control performance.

### Python API Usage

Instantiate the handler directly and call `setup()` with the `backend` parameter:

```python
from speech_to_speech.TTS.qwen3_tts_handler import Qwen3TTSHandler
from threading import Event

# GGML backend example (Linux/Windows)

handler = Qwen3TTSHandler()
handler.setup(
    should_listen=Event(),
    model_name="Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice",
    device="cuda",
    backend="ggml",
    speaker="Aiden",
)

# Torch backend example (Linux/Windows)

handler_torch = Qwen3TTSHandler()
handler_torch.setup(
    should_listen=Event(),
    model_name="Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice",
    device="cuda",
    backend="torch",
    speaker="Aiden",
)

```

### Installing Required Dependencies

Install the appropriate extras for your selected backend:

```bash

# For GGML backend (default, CPU/GPU compatible)

pip install "faster-qwen3-tts[ggml]"

# For Torch backend (CUDA-graphs optimized)

pip install "faster-qwen3-tts[torch]"

```

The `faster-qwen3-tts` package provides CUDA-specific wheels; consult [`src/speech_to_speech/TTS/README.md`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/README.md) (lines 96-104) for selecting specific CUDA versions if the default wheel conflicts with your driver version.

## Summary

- **macOS (Apple Silicon)**: Automatically uses MLX-Audio; `qwen3_tts_backend` parameter is ignored. Configure via `qwen3_tts_mlx_quantization` instead.
- **Linux/Windows**: Defaults to GGML backend for broad hardware compatibility.
- **Override capability**: Pass `--qwen3_tts_backend torch` on Linux/Windows to enable CUDA-graphs acceleration.
- **Validation**: Both CLI and Python API enforce backend values against `("ggml", "torch")`, throwing `ValueError` for invalid entries.
- **Source locations**: Backend logic resides in [`qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/qwen3_tts_handler.py) (lines 42-48, 84-86, 98-112), with arguments defined in [`qwen3_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/qwen3_tts_arguments.py) (lines 31-35).

## Frequently Asked Questions

### Can I use the Torch backend on macOS?

No. The `Qwen3TTSHandler` forces `self.backend = "mlx"` on Darwin systems (macOS) during initialization (lines 84-86), bypassing the faster-qwen3-tts framework entirely. The Torch and GGML backends are unavailable on Apple Silicon; you must use the MLX-Audio stack with quantization controls via `qwen3_tts_mlx_quantization`.

### What is the performance difference between GGML and Torch backends?

The **GGML** backend utilizes optimized C++ bindings for low-latency inference suitable for both CPU and GPU deployments, while the **Torch** backend leverages CUDA-graphs for reduced kernel launch overhead on NVIDIA GPUs. Choose GGML for broader hardware compatibility or when memory constraints are tight; select Torch for maximum throughput on CUDA-capable systems.

### Why does my backend argument raise a ValueError?

The handler validates the `backend` parameter against `VALID_FASTER_BACKENDS = ("ggml", "torch")` in [`qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/qwen3_tts_handler.py) (lines 42-48). Passing any value other than these literals—such as `"cuda"` or `"cpu"`—triggers a validation error at line 246. Ensure you use the exact strings `"ggml"` or `"torch"` when configuring the Qwen3-TTS backend.

### How do I verify which backend is actually running?

Inspect the handler's internal state after initialization. On Linux/Windows, check `handler.faster_backend` (which stores `"ggml"` or `"torch"`) or `handler.backend` (which stores `"faster_qwen3_tts"`). On macOS, `handler.backend` will equal `"mlx"`, indicating the MLX-Audio stack is active regardless of the input arguments.