# How to Run Speech-to-Speech on a GPU: Complete Setup Guide

> Accelerate speech-to-speech with GPU acceleration. Follow our complete setup guide to install Qwen3-TTS, configure PyTorch for CUDA, and run inference on your GPU for faster results.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-08-01

---

**Install CUDA-compatible Qwen3-TTS wheels, ensure PyTorch with CUDA support is present, and launch the pipeline with `--device cuda` to execute all heavy inference stages on the GPU.**

The `huggingface/speech-to-speech` repository provides a modular voice-to-voice pipeline that supports GPU acceleration for every compute-intensive stage. By default, the argument classes in the source code set `device="cuda"` for compatible models, but proper installation of CUDA-enabled dependencies is required to leverage GPU hardware.

## Prerequisites for CUDA Support

### NVIDIA Driver Requirements

Before installing Python dependencies, verify that your host system has NVIDIA drivers compatible with CUDA 12.4 or newer. The Qwen3-TTS backend specifically requires a CUDA runtime that matches the compiled wheel version you install.

### Installing GPU-Compatible Dependencies

The most critical step for GPU acceleration is installing the correct `qwentts-cpp-python` wheel for your CUDA version. According to the repository README, Linux users must manually specify the wheel matching their CUDA runtime (e.g., `cu124` or `cu130`).

Install the CUDA-specific wheel before the main package:

```bash

# Replace cu124 with your CUDA runtime version (cu124 or cu130)

pip install "qwentts-cpp-python==0.3.1+cu124" \
    -f https://huggingface.co/datasets/andito/qwentts-cpp-python-wheels/tree/main/whl/cu124

```

Then install the speech-to-speech package, which will detect available GPU backends:

```bash
pip install speech-to-speech

```

## Configuring Device Settings for GPU Execution

The pipeline automatically defaults to CUDA when available. In [`src/speech_to_speech/arguments_classes/whisper_stt_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/whisper_stt_arguments.py), the device parameter defaults to `"cuda"`, as does the configuration in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py).

You can explicitly enforce GPU usage across all components using the global device flag:

```bash
speech-to-speech \
    --stt whisper \
    --llm_backend transformers \
    --tts qwen3 \
    --device cuda

```

When `device="cuda"` is specified, the pipeline checks `torch.cuda.is_available()` before loading models onto the GPU.

## Running the Complete Pipeline on GPU

### Starting the GPU-Accelerated Server

Launch the server with CUDA-enabled backends for real-time speech processing:

```bash
export OPENAI_API_KEY=your-api-key  # Required for LLM backend

export CUDA_VISIBLE_DEVICES=0        # Optional: specify GPU ID

speech-to-speech \
    --stt whisper \
    --llm_backend transformers \
    --tts qwen3 \
    --device cuda \
    --host 0.0.0.0 \
    --port 8765

```

### Connecting a Local Client

With the server running on GPU, connect a local microphone client:

```bash
python scripts/listen_and_play_realtime.py --host 127.0.0.1 --port 8765

```

The server handles VAD, STT, LLM inference, and TTS generation on the GPU, while the client only manages audio I/O.

## Component-Specific GPU Configuration

Each pipeline stage supports specific GPU backends:

- **STT (Whisper)**: Defaults to `"cuda"` in [`src/speech_to_speech/arguments_classes/whisper_stt_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/whisper_stt_arguments.py). The Faster-Whisper backend also supports GPU execution via CUDA.
- **TTS (Qwen3)**: Uses GGML CUDA when the correct wheel is installed, as implemented in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py). Alternatively, use `--qwen3_tts_backend torch` for PyTorch CUDA execution.
- **LLM**: Transformers backend automatically utilizes CUDA when `device="cuda"` is passed.
- **VAD**: The Silero VAD runs on CPU by default, which is intentional as the load is negligible compared to other stages.

## Verifying GPU Utilization

To confirm that TTS inference is utilizing the GPU, run the benchmark script provided in the repository:

```bash
python scripts/benchmark_tts.py --device cuda --backend qwen3

```

This script explicitly sets `device="cuda"` and measures throughput when models are loaded on the GPU versus CPU.

## Summary

- Install the matching `qwentts-cpp-python` CUDA wheel (e.g., `+cu124`) before installing the main package to enable Qwen3-TTS GPU acceleration.
- The pipeline defaults to `device="cuda"` in [`whisper_stt_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/whisper_stt_arguments.py) and [`qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/qwen3_tts_handler.py), but explicitly pass `--device cuda` to ensure all components use GPU.
- Use `--llm_backend transformers` for CUDA-accelerated language model inference.
- The VAD component runs on CPU, which does not impact overall pipeline latency.
- Verify GPU usage with [`scripts/benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_tts.py) before deploying production workloads.

## Frequently Asked Questions

### Do I need a specific CUDA version for Qwen3-TTS?

Yes, you must install the `qwentts-cpp-python` wheel that matches your system's CUDA runtime version (either `cu124` or `cu130`). Mismatched versions will cause the TTS backend to fail or fall back to CPU execution. Check your CUDA version with `nvcc --version` before installing.

### Can I mix CPU and GPU across different pipeline stages?

While possible, it is not recommended. The [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) orchestrator initializes all heavy models (STT, LLM, TTS) with the same device parameter. Running some stages on CPU and others on GPU creates bottlenecks. If GPU memory is limited, prioritize placing the LLM and TTS on GPU while keeping STT on CPU using backend-specific flags.

### Why is my GPU not being detected by the speech-to-speech server?

First, verify that PyTorch was installed with CUDA support by running `python -c "import torch; print(torch.cuda.is_available())"`. If this returns `False`, reinstall PyTorch with the correct CUDA wheel from pytorch.org. Also ensure that `CUDA_VISIBLE_DEVICES` is not masking your GPU and that the `qwentts-cpp-python` wheel matches your CUDA runtime.

### Is the VAD component GPU-accelerated?

No, the Silero VAD implementation in this repository runs exclusively on CPU. This design choice is intentional because VAD computation is lightweight compared to STT, LLM, and TTS inference. Keeping VAD on CPU preserves GPU memory for the heavy generative models without impacting end-to-end latency.