# How `--qwen3_tts_backend` GGML vs Torch Affects Speech-to-Speech Performance

> Explore how --qwen3_tts_backend GGML vs Torch impacts speech-to-speech performance. Discover Torch's lower latency benefits and GGML's wider compatibility for your projects.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: performance
- Published: 2026-07-11

---

**The `--qwen3_tts_backend` flag selects between GGML (default, broad compatibility) and Torch (CUDA-optimized) execution engines, where Torch delivers lower latency through CUDA-graph capture while GGML offers wider hardware support at the cost of higher per-inference overhead.**

The `huggingface/speech-to-speech` repository provides a `--qwen3_tts_backend` option for the Qwen3-TTS model that lets you choose between two distinct inference engines. Understanding the **GGML** versus **Torch** backend distinction is crucial for optimizing real-time speech synthesis pipelines, as the choice directly determines latency, throughput, and hardware compatibility.

## Backend Architecture and Implementation

The system validates backend choices through `VALID_FASTER_BACKENDS = ("ggml", "torch")` at line 47 of the arguments class, with normalization handled by `_normalize_faster_backend` (lines 42-48). The actual implementation diverges significantly based on your selection.

### GGML Backend (Default)

The **GGML** backend utilizes the inference engine provided by the `faster-qwen3-tts[ggml]` package. It executes on GPU or CPU using standard GPU kernels without CUDA-graph optimization. This approach avoids the warm-up requirements of graph-based execution and maintains a lower memory footprint, making it suitable for diverse hardware configurations including older CUDA versions and CPU-only deployments.

### Torch Backend with CUDA-Graph Acceleration

The **Torch** backend leverages PyTorch with **CUDA-graph** acceleration. When initialized, the handler calls `warmup()` which executes `self.model._warmup(prefill_len=100)` on the first forward pass to capture the computation graph. This eliminates kernel launch overhead during subsequent inference steps. However, this backend requires a CUDA runtime that matches the compiled wheel (for example, CUDA 12.8) and utilizes Tensor cores for fast matrix operations.

## Performance Characteristics and Trade-offs

The performance delta between these backends stems from how they schedule GPU work.

**GGML** routes each inference step through regular GPU kernels. This results in **higher latency** compared to the Torch implementation, though it still maintains real-time streaming capabilities. The throughput (measured in audio seconds per wall-clock second) remains modest due to per-step kernel launch overhead.

**Torch** delivers **lower latency** and a higher real-time factor (RTF) because the captured CUDA graph removes kernel launch bottlenecks. This comes with slightly higher GPU memory usage due to the stored graph representation, and it requires compatible NVIDIA hardware with sufficient VRAM to hold the captured execution graph.

## Code Implementation Details

In [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py), the handler instantiates the model via `FasterQwen3TTS.from_pretrained(..., backend=backend)` (lines 100-112). For non-MLX backends, a warm-up path executes unless `parity_mode` is enabled (lines 58-66). This warm-up is **only** effective for the Torch backend, giving it a performance head-start that GGML cannot utilize.

The arguments definition in [`src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py) documents the default value (`ggml`) and explains the compatibility requirements for the Torch option.

## Configuration and Usage Examples

Select the **GGML** backend for universal compatibility across GPU generations and CPU environments:

```bash
python s2s_pipeline.py \
  --tts qwen3 \
  --qwen3_tts_model_name Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice \
  --qwen3_tts_device cuda \
  --qwen3_tts_backend ggml \
  --qwen3_tts_speaker Aiden

```

Select the **Torch** backend when running on compatible CUDA 12.x systems where latency minimization is critical:

```bash
python s2s_pipeline.py \
  --tts qwen3 \
  --qwen3_tts_model_name Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice \
  --qwen3_tts_device cuda \
  --qwen3_tts_backend torch \
  --qwen3_tts_speaker Aiden

```

## Summary

- **GGML** provides broad hardware compatibility (including older CUDA versions and CPUs) with lower memory usage but higher inference latency.
- **Torch** requires specific CUDA runtime versions (e.g., CUDA 12.8) but delivers superior performance through CUDA-graph capture and Tensor core utilization.
- The Torch backend executes a mandatory `warmup()` step that GGML skips, contributing to its lower per-step latency.
- Choose GGML for deployment flexibility; choose Torch for maximum real-time performance on modern NVIDIA GPUs.

## Frequently Asked Questions

### Which backend should I use for real-time speech-to-speech applications?

Choose **Torch** if your hardware supports the required CUDA runtime (e.g., CUDA 12.8) and you have sufficient GPU memory, as the CUDA-graph optimization eliminates kernel launch overhead and reduces latency. Select **GGML** if you need to support diverse GPU generations, CPU-only inference, or cannot match the specific CUDA wheel requirements.

### Why does the Torch backend require a warm-up step?

The Torch backend calls `self.model._warmup(prefill_len=100)` during initialization to capture a CUDA graph of the execution flow. This warm-up step pre-compiles the GPU kernel sequence, allowing subsequent inference passes to launch the entire computation graph as a single unit rather than dispatching individual kernels, which dramatically reduces latency.

### Can I use the Torch backend on CPU or older GPUs?

No, the Torch backend requires a CUDA-capable GPU with a runtime version matching the compiled `faster-qwen3-tts` wheel (such as CUDA 12.8). The GGML backend is the appropriate choice for CPU inference or older GPU architectures that lack the specific CUDA version support or sufficient memory for graph capture.

### How do I verify which backend is currently active?

The `speech-to-speech` pipeline logs the selected backend during initialization. You can also inspect the command-line arguments in [`src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py) where the `--qwen3_tts_backend` default is set to `ggml`, or check the handler logs in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) to confirm whether the CUDA-graph warm-up path executed (indicating Torch mode).