# Recommended Hardware for Training and Inference in Speech-to-Speech Pipelines

> Discover recommended hardware for speech-to-speech inference. Optimize performance with NVIDIA CUDA GPUs for low latency or choose Apple Silicon M1/M2/M3 for native speed.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: performance
- Published: 2026-08-02

---

**The huggingface/speech-to-speech repository does not include training scripts; it is designed for inference only, with NVIDIA CUDA GPUs recommended for lowest latency and Apple Silicon (M1/M2/M3) as a fully native alternative.**

This modular voice-agent pipeline chains Voice Activity Detection (VAD) → Speech-to-Text (STT) → Large Language Model (LLM) → Text-to-Speech (TTS). While you cannot train models directly in this codebase, selecting the right hardware for inference dramatically impacts real-time performance. Below is the complete hardware matrix drawn from the official repository documentation and source files.

---

## Training vs. Inference Scope

The `speech-to-speech` project **does not contain training implementations** for any of its component models. According to the repository structure and [`README.md`](https://github.com/huggingface/speech-to-speech/blob/main/README.md), this is a runtime inference engine only. Training guidance lives with the upstream model repositories (e.g., NVIDIA NeMo for Parakeet, OpenAI for Whisper, Alibaba for Qwen3-TTS).

For inference, hardware recommendations vary by backend selection. The pipeline automatically selects optimal backends based on your platform, or you can override them via CLI flags defined in [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py).

---

## Inference Hardware by Pipeline Component

### Voice Activity Detection (VAD)

**Silero VAD v5** runs on **CPU with negligible load**. No GPU acceleration is needed or implemented.

```bash

# Default behavior — no flags required

speech-to-speech --mode realtime

```

---

### Speech-to-Text (STT) Hardware Options

| Backend | Recommended Hardware | CLI Flag |
|---------|---------------------|----------|
| **Parakeet TDT** (default) | CUDA GPU or Apple Silicon (MLX) | `--stt parakeet-tdt` |
| **Whisper / Faster-Whisper** | CUDA GPU for speed; CPU fallback | `--stt faster-whisper` |
| **Lightning Whisper MLX** | Apple Silicon (MPS) | `--stt whisper-mlx` |
| **Paraformer (FunASR)** | CUDA GPU or CPU | `--stt paraformer` |

The **Parakeet TDT** backend provides the best CUDA performance. On macOS, `whisper-mlx` leverages Apple's Neural Engine through the MLX framework.

```bash

# Optimal STT on NVIDIA GPU

speech-to-speech --stt parakeet-tdt

# Optimal STT on Apple Silicon

speech-to-speech --stt whisper-mlx

```

---

### Large Language Model (LLM) Hardware Options

| Backend | Recommended Hardware | CLI Flag |
|---------|---------------------|----------|
| **OpenAI-compatible API** | Any platform (remote service) | `--llm_backend responses-api` |
| **Transformers** (local) | CUDA GPU for throughput; CPU fallback | `--llm_backend transformers` |
| **mlx-lm** | Apple Silicon (MPS) — built into macOS | `--llm_backend mlx-lm` |

Local LLM inference in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) automatically detects macOS and recommends `mlx-lm`. For CUDA systems, the Transformers backend provides maximum throughput.

**Memory requirements vary dramatically:**

- **Gemma 4 31B**: >30 GB VRAM required
- **Qwen3 4B / Gemma 2B**: 12–16 GB GPUs sufficient

```bash

# Local LLM on CUDA with Transformers

speech-to-speech --llm_backend transformers

# Local LLM optimized for Apple Silicon

speech-to-speech --llm_backend mlx-lm

```

---

### Text-to-Speech (TTS) Hardware Options

| Backend | Recommended Hardware | CLI Flag |
|---------|---------------------|----------|
| **Qwen3-TTS** (default) | CUDA GPU (GGML) **or** Apple Silicon (MLX) | `--tts qwen3` |
| **Kokoro-82M** | CUDA/CPU **or** Apple Silicon | `--tts kokoro` |
| **Pocket TTS / ChatTTS / MMS-TTS** | CUDA/CPU (no GPU-specific optimization) | `--tts pocket`, `chattts`, `facebook-mms` |

---

## Critical Platform-Specific Requirements

### NVIDIA CUDA GPUs: CUDA 12.x Compatibility

The **Qwen3-TTS GGML wheels target CUDA 12.8**. Running on mismatched CUDA versions causes backend initialization failures. As documented in the README at lines 102–118, you must install the wheel matching your system's CUDA runtime from the Hugging Face wheelhouse before installing the package.

```bash

# Example: Install CUDA 12.4 wheel explicitly

pip install https://huggingface.co/spaces/huggingface/speech-to-speech/resolve/main/qwen3-tts-ggml-cu124.whl
pip install speech-to-speech

```

Check your CUDA version:

```bash
nvidia-smi | grep -i "cuda version"

```

---

### Apple Silicon: Native MLX Integration

On macOS, the pipeline automatically uses:

- **MLX-audio** for STT (when `whisper-mlx` selected)
- **mlx-lm** for LLM inference
- **MLX-audio backend** for Qwen3-TTS

This provides GPU-class performance without discrete hardware. The runtime logic in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) detects `sys.platform == "darwin"` and surfaces macOS-specific recommendations.

```bash

# Complete Apple Silicon optimized pipeline

speech-to-speech \
    --stt whisper-mlx \
    --llm_backend mlx-lm \
    --tts qwen3 \
    --mode realtime

```

---

### CPU-Only Fallback

All components support CPU execution. Install the Qwen3-TTS `cpu` wheel for TTS without GPU:

```bash
pip install https://huggingface.co/spaces/huggingface/speech-to-speech/resolve/main/qwen3-tts-ggml-cpu.whl

```

Expect **significantly higher latency** — suitable for prototyping only. Lines 115–118 of the README note this as a fallback, not a production configuration.

---

## Profiling Hardware Performance

The repository includes benchmark scripts to validate your hardware setup:

```bash

# Benchmark TTS backends with MLX quantization variants

python scripts/benchmark_tts.py \
    --handlers qwen3 kokoro pocket \
    --iterations 5 \
    --qwen3_mlx_quantizations bf16 4bit 6bit 8bit

```

```bash

# Benchmark STT backend latency

python scripts/benchmark_stt.py \
    --handlers parakeet-tdt faster-whisper whisper-mlx \
    --iterations 10

```

These scripts report per-utterance latency and throughput, allowing data-driven backend selection. The [`benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/benchmark_tts.py) script at line 240 specifically tests quantization-aware performance for memory-constrained deployments.

---

## Complete Hardware Configuration Examples

### Production NVIDIA Setup (CUDA 12.8, ≥16 GB VRAM)

```bash
export OPENAI_API_KEY=your_key_here

speech-to-speech \
    --stt parakeet-tdt \
    --llm_backend transformers \
    --tts qwen3 \
    --mode realtime

```

### MacBook Pro M3 (On-Device Inference, No API Calls)

```bash
speech-to-speech \
    --stt whisper-mlx \
    --llm_backend mlx-lm \
    --tts qwen3 \
    --mode realtime

```

### Headless Server (CPU-Only, API-Based LLM)

```bash
export OPENAI_API_KEY=your_key_here

speech-to-speech \
    --stt faster-whisper \
    --llm_backend responses-api \
    --tts qwen3 \
    --qwen3_tts_backend ggml \
    --mode realtime

```

---

## Memory and Quantization Strategies

For **VRAM-constrained systems**, use MLX quantization on Apple Silicon. The [`benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/benchmark_tts.py) script validates these configurations:

| Quantization | Memory Reduction | Quality Impact |
|-------------|------------------|----------------|
| `bf16` | 2× vs FP32 | Minimal |
| `8bit` | 4× | Slight |
| `6bit` | 5.3× | Moderate |
| `4bit` | 8× | Noticeable |

On NVIDIA GPUs, select smaller models (2B–4B parameters) rather than quantizing, as the GGML backend does not expose runtime quantization in this release.

---

## Summary

- **No training hardware guidance exists** in the `speech-to-speech` repository; use upstream model documentation for training requirements.
- **NVIDIA CUDA GPUs with CUDA 12.x** provide lowest latency for parallel STT/LLM/TTS inference on Linux/Windows.
- **Apple Silicon (M1/M2/M3)** offers a fully native pipeline via MLX without discrete GPU.
- **CPU-only execution works for prototyping** with substantially higher latency.
- **Benchmark scripts** in [`scripts/benchmark_tts.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_tts.py) and [`scripts/benchmark_stt.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/benchmark_stt.py) validate hardware-specific performance.
- **Memory requirements scale with model size**: 31B LLMs need >30 GB VRAM; 2B–4B models fit 12–16 GB.

---

## Frequently Asked Questions

### Does the speech-to-speech repository support model training?

No. The repository contains **inference-only code** for running pretrained VAD, STT, LLM, and TTS models. Training scripts reside in upstream repositories (NVIDIA NeMo, OpenAI Whisper, Qwen3-TTS). For training hardware, consult those projects' documentation.

### What is the minimum VRAM for local LLM inference?

**12 GB VRAM** suffices for 2B–4B parameter models (Qwen3-4B, Gemma-2B). **30+ GB VRAM** is required for Gemma-4 31B. The `mlx-lm` backend on Apple Silicon uses unified memory, allowing larger models through quantization.

### Can I run the full pipeline without any GPU?

Yes, but with trade-offs. Set `--qwen3_tts_backend ggml` with the CPU wheel, use `--stt faster-whisper`, and either `--llm_backend responses-api` (remote) or `--llm_backend transformers` (local CPU). Latency increases 5–10× versus GPU.

### How do I verify my CUDA version matches Qwen3-TTS requirements?

Run `nvidia-smi` and check the "CUDA Version" line. Install the matching wheel from the Hugging Face wheelhouse before `pip install speech-to-speech`. Mismatched CUDA runtimes cause silent backend failures in TTS initialization.