How to Run Speech-to-Speech on a GPU: Complete Setup Guide

Install CUDA-compatible Qwen3-TTS wheels, ensure PyTorch with CUDA support is present, and launch the pipeline with --device cuda to execute all heavy inference stages on the GPU.

The huggingface/speech-to-speech repository provides a modular voice-to-voice pipeline that supports GPU acceleration for every compute-intensive stage. By default, the argument classes in the source code set device="cuda" for compatible models, but proper installation of CUDA-enabled dependencies is required to leverage GPU hardware.

Prerequisites for CUDA Support

NVIDIA Driver Requirements

Before installing Python dependencies, verify that your host system has NVIDIA drivers compatible with CUDA 12.4 or newer. The Qwen3-TTS backend specifically requires a CUDA runtime that matches the compiled wheel version you install.

Installing GPU-Compatible Dependencies

The most critical step for GPU acceleration is installing the correct qwentts-cpp-python wheel for your CUDA version. According to the repository README, Linux users must manually specify the wheel matching their CUDA runtime (e.g., cu124 or cu130).

Install the CUDA-specific wheel before the main package:


# Replace cu124 with your CUDA runtime version (cu124 or cu130)

pip install "qwentts-cpp-python==0.3.1+cu124" \
    -f https://huggingface.co/datasets/andito/qwentts-cpp-python-wheels/tree/main/whl/cu124

Then install the speech-to-speech package, which will detect available GPU backends:

pip install speech-to-speech

Configuring Device Settings for GPU Execution

The pipeline automatically defaults to CUDA when available. In src/speech_to_speech/arguments_classes/whisper_stt_arguments.py, the device parameter defaults to "cuda", as does the configuration in src/speech_to_speech/TTS/qwen3_tts_handler.py.

You can explicitly enforce GPU usage across all components using the global device flag:

speech-to-speech \
    --stt whisper \
    --llm_backend transformers \
    --tts qwen3 \
    --device cuda

When device="cuda" is specified, the pipeline checks torch.cuda.is_available() before loading models onto the GPU.

Running the Complete Pipeline on GPU

Starting the GPU-Accelerated Server

Launch the server with CUDA-enabled backends for real-time speech processing:

export OPENAI_API_KEY=your-api-key  # Required for LLM backend

export CUDA_VISIBLE_DEVICES=0        # Optional: specify GPU ID

speech-to-speech \
    --stt whisper \
    --llm_backend transformers \
    --tts qwen3 \
    --device cuda \
    --host 0.0.0.0 \
    --port 8765

Connecting a Local Client

With the server running on GPU, connect a local microphone client:

python scripts/listen_and_play_realtime.py --host 127.0.0.1 --port 8765

The server handles VAD, STT, LLM inference, and TTS generation on the GPU, while the client only manages audio I/O.

Component-Specific GPU Configuration

Each pipeline stage supports specific GPU backends:

  • STT (Whisper): Defaults to "cuda" in src/speech_to_speech/arguments_classes/whisper_stt_arguments.py. The Faster-Whisper backend also supports GPU execution via CUDA.
  • TTS (Qwen3): Uses GGML CUDA when the correct wheel is installed, as implemented in src/speech_to_speech/TTS/qwen3_tts_handler.py. Alternatively, use --qwen3_tts_backend torch for PyTorch CUDA execution.
  • LLM: Transformers backend automatically utilizes CUDA when device="cuda" is passed.
  • VAD: The Silero VAD runs on CPU by default, which is intentional as the load is negligible compared to other stages.

Verifying GPU Utilization

To confirm that TTS inference is utilizing the GPU, run the benchmark script provided in the repository:

python scripts/benchmark_tts.py --device cuda --backend qwen3

This script explicitly sets device="cuda" and measures throughput when models are loaded on the GPU versus CPU.

Summary

  • Install the matching qwentts-cpp-python CUDA wheel (e.g., +cu124) before installing the main package to enable Qwen3-TTS GPU acceleration.
  • The pipeline defaults to device="cuda" in whisper_stt_arguments.py and qwen3_tts_handler.py, but explicitly pass --device cuda to ensure all components use GPU.
  • Use --llm_backend transformers for CUDA-accelerated language model inference.
  • The VAD component runs on CPU, which does not impact overall pipeline latency.
  • Verify GPU usage with scripts/benchmark_tts.py before deploying production workloads.

Frequently Asked Questions

Do I need a specific CUDA version for Qwen3-TTS?

Yes, you must install the qwentts-cpp-python wheel that matches your system's CUDA runtime version (either cu124 or cu130). Mismatched versions will cause the TTS backend to fail or fall back to CPU execution. Check your CUDA version with nvcc --version before installing.

Can I mix CPU and GPU across different pipeline stages?

While possible, it is not recommended. The s2s_pipeline.py orchestrator initializes all heavy models (STT, LLM, TTS) with the same device parameter. Running some stages on CPU and others on GPU creates bottlenecks. If GPU memory is limited, prioritize placing the LLM and TTS on GPU while keeping STT on CPU using backend-specific flags.

Why is my GPU not being detected by the speech-to-speech server?

First, verify that PyTorch was installed with CUDA support by running python -c "import torch; print(torch.cuda.is_available())". If this returns False, reinstall PyTorch with the correct CUDA wheel from pytorch.org. Also ensure that CUDA_VISIBLE_DEVICES is not masking your GPU and that the qwentts-cpp-python wheel matches your CUDA runtime.

Is the VAD component GPU-accelerated?

No, the Silero VAD implementation in this repository runs exclusively on CPU. This design choice is intentional because VAD computation is lightweight compared to STT, LLM, and TTS inference. Keeping VAD on CPU preserves GPU memory for the heavy generative models without impacting end-to-end latency.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →