# How to Configure CUDA Settings for Qwen3-TTS on Linux

> Optimize Qwen3-TTS on Linux by configuring CUDA settings. Set device to cuda, backend to torch and dtype to float16 or auto for efficient GPU acceleration.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-30

---

**To configure CUDA settings for Qwen3-TTS on Linux, set `qwen3_tts_device` to `"cuda"`, ensure `qwen3_tts_backend` is `"torch"`, and select `"float16"` or `"auto"` dtype to leverage GPU acceleration with optimal memory efficiency.**

The **Qwen3-TTS** handler in the **huggingface/speech-to-speech** repository provides dedicated CUDA support for Linux systems, enabling low-latency text-to-speech synthesis via the **Faster-Qwen3-TTS** backend. Proper configuration involves mapping CLI arguments through the `Qwen3TTSHandlerArguments` dataclass to the runtime initialization in `Qwen3TTSHandler.setup`, which validates GPU capabilities and selects appropriate precision formats automatically.

## Understanding the CUDA Configuration Architecture

The CUDA configuration flow follows a strict path from user input to GPU execution:

```

CLI / Python args → Qwen3TTSHandlerArguments → Qwen3TTSHandler.setup → Faster-Qwen3-TTS (CUDA backend)

```

### Qwen3TTSHandlerArguments Dataclass

The argument definitions reside in [`src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py). This dataclass exposes four critical CUDA-related fields:

- **`qwen3_tts_device`**: Defaults to `"cuda"`. Accepts `"cpu"`, `"mps"`, or `"auto"`.
- **`qwen3_tts_dtype`**: Controls tensor precision (`"float16"`, `"bfloat16"`, `"float32"`, or `"auto"`). When set to `"auto"`, the handler queries `torch.cuda.is_bf16_supported()` to select `torch.bfloat16` on compatible hardware, falling back to `torch.float16` otherwise.
- **`qwen3_tts_attn_implementation`**: Selects the attention kernel (`"eager"`, `"flash_attention_2"`, or `"sdpa"`). The default `"eager"` works universally, while `"flash_attention_2"` requires Ampere-generation GPUs or newer.
- **`qwen3_tts_backend`**: Must be `"torch"` for CUDA support. The `"ggml"` backend does not utilize CUDA acceleration.

### Qwen3TTSHandler.setup Runtime Initialization

The [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) file contains the `setup` method (lines 91-112) that translates arguments into concrete PyTorch objects. This method:

1. Assigns the device string directly to `self.device`
2. Resolves `"auto"` dtype using `torch.cuda.is_bf16_supported()` (lines 91-97)
3. Imports `faster_qwen3_tts` and instantiates the model with the specified `device`, `dtype`, and `attn_implementation` (lines 98-112)

## Essential CUDA Configuration Parameters

Configure these arguments to optimize throughput and compatibility on Linux:

| Argument | Purpose | Recommended Value |
|----------|---------|-----------------|
| `qwen3_tts_device` | CUDA device identifier | `"cuda"` (uses default GPU) |
| `qwen3_tts_backend` | Inference backend | `"torch"` (required for CUDA) |
| `qwen3_tts_dtype` | Numerical precision | `"auto"` or `"float16"` |
| `qwen3_tts_attn_implementation` | Attention algorithm | `"flash_attention_2"` (if supported) or `"sdpa"` |
| `qwen3_tts_parity_mode` | CUDA graph stability | `False` (set `True` only if encountering graph-related crashes) |

## Configuration Methods

### Command-Line Interface

Pass flags directly when launching the speech-to-speech pipeline:

```bash
python -m speech_to_speech \
  --qwen3_tts_device cuda \
  --qwen3_tts_dtype float16 \
  --qwen3_tts_attn_implementation flash_attention_2 \
  --qwen3_tts_backend torch \
  --qwen3_tts_model_name Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice

```

All flags map directly to fields in `Qwen3TTSHandlerArguments`, which the pipeline parses and passes to the handler.

### Programmatic Handler Setup

Instantiate and configure the handler directly for embedded applications:

```python
from speech_to_speech.TTS.qwen3_tts_handler import Qwen3TTSHandler
from threading import Event

handler = Qwen3TTSHandler()
handler.setup(
    should_listen=Event(),
    model_name="Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice",
    device="cuda",
    dtype="float16",
    attn_implementation="flash_attention_2",
    backend="torch",
)

```

The `setup` method validates GPU availability and raises clear import errors if CUDA initialization fails.

### Pipeline Integration

Modify handler arguments before creating an `S2SPipeline` instance:

```python
from speech_to_speech.s2s_pipeline import S2SPipeline
from speech_to_speech.arguments_classes.qwen3_tts_arguments import Qwen3TTSHandlerArguments

qwen_args = Qwen3TTSHandlerArguments()
qwen_args.qwen3_tts_device = "cuda"
qwen_args.qwen3_tts_dtype = "auto"
qwen_args.qwen3_tts_attn_implementation = "flash_attention_2"

pipeline = S2SPipeline(
    # ... other component arguments ...

    qwen3_tts_handler_kwargs=qwen_args,
)

```

The pipeline passes `vars(qwen3_tts_handler_kwargs)` directly to the handler's `setup` method according to the implementation in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) (lines 97-108).

## Verifying CUDA Configuration

Confirm your settings are active by running the backend validation test:

```bash
pytest tests/test_qwen3_tts_handler_backend.py -k cuda

```

The test `test_qwen3_tts_handler_backend` asserts that `handler.backend == "faster_qwen3_tts"` when running on Linux with CUDA available (lines 106-120). You can also verify GPU utilization by inspecting the handler's device attribute after setup:

```python
print(f"Device: {handler.device}")  # Should output: cuda

```

## Summary

- **Set `qwen3_tts_device`** to `"cuda"` and **`qwen3_tts_backend`** to `"torch"` to enable GPU acceleration on Linux.
- **Use `"auto"` dtype** to let the handler automatically select BF16 on compatible GPUs (RTX 30-series, A100, etc.) or fall back to FP16.
- **Enable Flash Attention 2** via `qwen3_tts_attn_implementation="flash_attention_2"` for significant speedups on modern architectures.
- **Reference the source files** [`qwen3_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/qwen3_tts_arguments.py) and [`qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/qwen3_tts_handler.py) to understand how CLI arguments translate to PyTorch CUDA tensors.

## Frequently Asked Questions

### What is the difference between "auto" and explicit dtype settings?

When `qwen3_tts_dtype` is set to `"auto"`, the handler calls `torch.cuda.is_bf16_supported()` to detect BF16 capability. If your GPU supports BF16 (Compute Capability 8.0+), it selects `torch.bfloat16` for better numerical stability; otherwise, it uses `torch.float16`. Explicitly setting `"float16"` forces lower precision regardless of hardware support, while `"float32"` disables mixed precision entirely.

### How do I troubleshoot CUDA out-of-memory errors?

Reduce memory pressure by switching `qwen3_tts_dtype` to `"float16"` instead of `"float32"`, or decrease the streaming chunk size if configured. Ensure `qwen3_tts_attn_implementation` is not set to `"eager"` on large models, as standard attention consumes significantly more VRAM than SDPA or Flash Attention 2.

### Can I use the "ggml" backend with CUDA?

No. According to the source in [`qwen3_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/qwen3_tts_arguments.py), the `"ggml"` backend does not support CUDA acceleration. You must set `qwen3_tts_backend="torch"` to utilize NVIDIA GPUs on Linux.

### Why does Qwen3-TTS default to "eager" attention instead of Flash Attention 2?

The default `"eager"` implementation provides maximum compatibility across all GPU generations. Flash Attention 2 requires specific CUDA compute capabilities and additional kernel support. Set `qwen3_tts_attn_implementation="flash_attention_2"` manually if your hardware supports it (NVIDIA Ampere, Ada Lovelace, or Hopper architectures).