How to Configure CUDA Settings for Qwen3-TTS on Linux

To configure CUDA settings for Qwen3-TTS on Linux, set qwen3_tts_device to "cuda", ensure qwen3_tts_backend is "torch", and select "float16" or "auto" dtype to leverage GPU acceleration with optimal memory efficiency.

The Qwen3-TTS handler in the huggingface/speech-to-speech repository provides dedicated CUDA support for Linux systems, enabling low-latency text-to-speech synthesis via the Faster-Qwen3-TTS backend. Proper configuration involves mapping CLI arguments through the Qwen3TTSHandlerArguments dataclass to the runtime initialization in Qwen3TTSHandler.setup, which validates GPU capabilities and selects appropriate precision formats automatically.

Understanding the CUDA Configuration Architecture

The CUDA configuration flow follows a strict path from user input to GPU execution:


CLI / Python args → Qwen3TTSHandlerArguments → Qwen3TTSHandler.setup → Faster-Qwen3-TTS (CUDA backend)

Qwen3TTSHandlerArguments Dataclass

The argument definitions reside in src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py. This dataclass exposes four critical CUDA-related fields:

  • qwen3_tts_device: Defaults to "cuda". Accepts "cpu", "mps", or "auto".
  • qwen3_tts_dtype: Controls tensor precision ("float16", "bfloat16", "float32", or "auto"). When set to "auto", the handler queries torch.cuda.is_bf16_supported() to select torch.bfloat16 on compatible hardware, falling back to torch.float16 otherwise.
  • qwen3_tts_attn_implementation: Selects the attention kernel ("eager", "flash_attention_2", or "sdpa"). The default "eager" works universally, while "flash_attention_2" requires Ampere-generation GPUs or newer.
  • qwen3_tts_backend: Must be "torch" for CUDA support. The "ggml" backend does not utilize CUDA acceleration.

Qwen3TTSHandler.setup Runtime Initialization

The src/speech_to_speech/TTS/qwen3_tts_handler.py file contains the setup method (lines 91-112) that translates arguments into concrete PyTorch objects. This method:

  1. Assigns the device string directly to self.device
  2. Resolves "auto" dtype using torch.cuda.is_bf16_supported() (lines 91-97)
  3. Imports faster_qwen3_tts and instantiates the model with the specified device, dtype, and attn_implementation (lines 98-112)

Essential CUDA Configuration Parameters

Configure these arguments to optimize throughput and compatibility on Linux:

Argument Purpose Recommended Value
qwen3_tts_device CUDA device identifier "cuda" (uses default GPU)
qwen3_tts_backend Inference backend "torch" (required for CUDA)
qwen3_tts_dtype Numerical precision "auto" or "float16"
qwen3_tts_attn_implementation Attention algorithm "flash_attention_2" (if supported) or "sdpa"
qwen3_tts_parity_mode CUDA graph stability False (set True only if encountering graph-related crashes)

Configuration Methods

Command-Line Interface

Pass flags directly when launching the speech-to-speech pipeline:

python -m speech_to_speech \
  --qwen3_tts_device cuda \
  --qwen3_tts_dtype float16 \
  --qwen3_tts_attn_implementation flash_attention_2 \
  --qwen3_tts_backend torch \
  --qwen3_tts_model_name Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice

All flags map directly to fields in Qwen3TTSHandlerArguments, which the pipeline parses and passes to the handler.

Programmatic Handler Setup

Instantiate and configure the handler directly for embedded applications:

from speech_to_speech.TTS.qwen3_tts_handler import Qwen3TTSHandler
from threading import Event

handler = Qwen3TTSHandler()
handler.setup(
    should_listen=Event(),
    model_name="Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice",
    device="cuda",
    dtype="float16",
    attn_implementation="flash_attention_2",
    backend="torch",
)

The setup method validates GPU availability and raises clear import errors if CUDA initialization fails.

Pipeline Integration

Modify handler arguments before creating an S2SPipeline instance:

from speech_to_speech.s2s_pipeline import S2SPipeline
from speech_to_speech.arguments_classes.qwen3_tts_arguments import Qwen3TTSHandlerArguments

qwen_args = Qwen3TTSHandlerArguments()
qwen_args.qwen3_tts_device = "cuda"
qwen_args.qwen3_tts_dtype = "auto"
qwen_args.qwen3_tts_attn_implementation = "flash_attention_2"

pipeline = S2SPipeline(
    # ... other component arguments ...

    qwen3_tts_handler_kwargs=qwen_args,
)

The pipeline passes vars(qwen3_tts_handler_kwargs) directly to the handler's setup method according to the implementation in src/speech_to_speech/s2s_pipeline.py (lines 97-108).

Verifying CUDA Configuration

Confirm your settings are active by running the backend validation test:

pytest tests/test_qwen3_tts_handler_backend.py -k cuda

The test test_qwen3_tts_handler_backend asserts that handler.backend == "faster_qwen3_tts" when running on Linux with CUDA available (lines 106-120). You can also verify GPU utilization by inspecting the handler's device attribute after setup:

print(f"Device: {handler.device}")  # Should output: cuda

Summary

  • Set qwen3_tts_device to "cuda" and qwen3_tts_backend to "torch" to enable GPU acceleration on Linux.
  • Use "auto" dtype to let the handler automatically select BF16 on compatible GPUs (RTX 30-series, A100, etc.) or fall back to FP16.
  • Enable Flash Attention 2 via qwen3_tts_attn_implementation="flash_attention_2" for significant speedups on modern architectures.
  • Reference the source files qwen3_tts_arguments.py and qwen3_tts_handler.py to understand how CLI arguments translate to PyTorch CUDA tensors.

Frequently Asked Questions

What is the difference between "auto" and explicit dtype settings?

When qwen3_tts_dtype is set to "auto", the handler calls torch.cuda.is_bf16_supported() to detect BF16 capability. If your GPU supports BF16 (Compute Capability 8.0+), it selects torch.bfloat16 for better numerical stability; otherwise, it uses torch.float16. Explicitly setting "float16" forces lower precision regardless of hardware support, while "float32" disables mixed precision entirely.

How do I troubleshoot CUDA out-of-memory errors?

Reduce memory pressure by switching qwen3_tts_dtype to "float16" instead of "float32", or decrease the streaming chunk size if configured. Ensure qwen3_tts_attn_implementation is not set to "eager" on large models, as standard attention consumes significantly more VRAM than SDPA or Flash Attention 2.

Can I use the "ggml" backend with CUDA?

No. According to the source in qwen3_tts_arguments.py, the "ggml" backend does not support CUDA acceleration. You must set qwen3_tts_backend="torch" to utilize NVIDIA GPUs on Linux.

Why does Qwen3-TTS default to "eager" attention instead of Flash Attention 2?

The default "eager" implementation provides maximum compatibility across all GPU generations. Flash Attention 2 requires specific CUDA compute capabilities and additional kernel support. Set qwen3_tts_attn_implementation="flash_attention_2" manually if your hardware supports it (NVIDIA Ampere, Ada Lovelace, or Hopper architectures).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →