Recommended Hardware for Training and Inference in Speech-to-Speech Pipelines

The huggingface/speech-to-speech repository does not include training scripts; it is designed for inference only, with NVIDIA CUDA GPUs recommended for lowest latency and Apple Silicon (M1/M2/M3) as a fully native alternative.

This modular voice-agent pipeline chains Voice Activity Detection (VAD) → Speech-to-Text (STT) → Large Language Model (LLM) → Text-to-Speech (TTS). While you cannot train models directly in this codebase, selecting the right hardware for inference dramatically impacts real-time performance. Below is the complete hardware matrix drawn from the official repository documentation and source files.


Training vs. Inference Scope

The speech-to-speech project does not contain training implementations for any of its component models. According to the repository structure and README.md, this is a runtime inference engine only. Training guidance lives with the upstream model repositories (e.g., NVIDIA NeMo for Parakeet, OpenAI for Whisper, Alibaba for Qwen3-TTS).

For inference, hardware recommendations vary by backend selection. The pipeline automatically selects optimal backends based on your platform, or you can override them via CLI flags defined in src/speech_to_speech/arguments_classes/module_arguments.py.


Inference Hardware by Pipeline Component

Voice Activity Detection (VAD)

Silero VAD v5 runs on CPU with negligible load. No GPU acceleration is needed or implemented.


# Default behavior — no flags required

speech-to-speech --mode realtime

Speech-to-Text (STT) Hardware Options

Backend Recommended Hardware CLI Flag
Parakeet TDT (default) CUDA GPU or Apple Silicon (MLX) --stt parakeet-tdt
Whisper / Faster-Whisper CUDA GPU for speed; CPU fallback --stt faster-whisper
Lightning Whisper MLX Apple Silicon (MPS) --stt whisper-mlx
Paraformer (FunASR) CUDA GPU or CPU --stt paraformer

The Parakeet TDT backend provides the best CUDA performance. On macOS, whisper-mlx leverages Apple's Neural Engine through the MLX framework.


# Optimal STT on NVIDIA GPU

speech-to-speech --stt parakeet-tdt

# Optimal STT on Apple Silicon

speech-to-speech --stt whisper-mlx

Large Language Model (LLM) Hardware Options

Backend Recommended Hardware CLI Flag
OpenAI-compatible API Any platform (remote service) --llm_backend responses-api
Transformers (local) CUDA GPU for throughput; CPU fallback --llm_backend transformers
mlx-lm Apple Silicon (MPS) — built into macOS --llm_backend mlx-lm

Local LLM inference in src/speech_to_speech/s2s_pipeline.py automatically detects macOS and recommends mlx-lm. For CUDA systems, the Transformers backend provides maximum throughput.

Memory requirements vary dramatically:

  • Gemma 4 31B: >30 GB VRAM required
  • Qwen3 4B / Gemma 2B: 12–16 GB GPUs sufficient

# Local LLM on CUDA with Transformers

speech-to-speech --llm_backend transformers

# Local LLM optimized for Apple Silicon

speech-to-speech --llm_backend mlx-lm

Text-to-Speech (TTS) Hardware Options

Backend Recommended Hardware CLI Flag
Qwen3-TTS (default) CUDA GPU (GGML) or Apple Silicon (MLX) --tts qwen3
Kokoro-82M CUDA/CPU or Apple Silicon --tts kokoro
Pocket TTS / ChatTTS / MMS-TTS CUDA/CPU (no GPU-specific optimization) --tts pocket, chattts, facebook-mms

Critical Platform-Specific Requirements

NVIDIA CUDA GPUs: CUDA 12.x Compatibility

The Qwen3-TTS GGML wheels target CUDA 12.8. Running on mismatched CUDA versions causes backend initialization failures. As documented in the README at lines 102–118, you must install the wheel matching your system's CUDA runtime from the Hugging Face wheelhouse before installing the package.


# Example: Install CUDA 12.4 wheel explicitly

pip install https://huggingface.co/spaces/huggingface/speech-to-speech/resolve/main/qwen3-tts-ggml-cu124.whl
pip install speech-to-speech

Check your CUDA version:

nvidia-smi | grep -i "cuda version"

Apple Silicon: Native MLX Integration

On macOS, the pipeline automatically uses:

  • MLX-audio for STT (when whisper-mlx selected)
  • mlx-lm for LLM inference
  • MLX-audio backend for Qwen3-TTS

This provides GPU-class performance without discrete hardware. The runtime logic in src/speech_to_speech/s2s_pipeline.py detects sys.platform == "darwin" and surfaces macOS-specific recommendations.


# Complete Apple Silicon optimized pipeline

speech-to-speech \
    --stt whisper-mlx \
    --llm_backend mlx-lm \
    --tts qwen3 \
    --mode realtime

CPU-Only Fallback

All components support CPU execution. Install the Qwen3-TTS cpu wheel for TTS without GPU:

pip install https://huggingface.co/spaces/huggingface/speech-to-speech/resolve/main/qwen3-tts-ggml-cpu.whl

Expect significantly higher latency — suitable for prototyping only. Lines 115–118 of the README note this as a fallback, not a production configuration.


Profiling Hardware Performance

The repository includes benchmark scripts to validate your hardware setup:


# Benchmark TTS backends with MLX quantization variants

python scripts/benchmark_tts.py \
    --handlers qwen3 kokoro pocket \
    --iterations 5 \
    --qwen3_mlx_quantizations bf16 4bit 6bit 8bit

# Benchmark STT backend latency

python scripts/benchmark_stt.py \
    --handlers parakeet-tdt faster-whisper whisper-mlx \
    --iterations 10

These scripts report per-utterance latency and throughput, allowing data-driven backend selection. The benchmark_tts.py script at line 240 specifically tests quantization-aware performance for memory-constrained deployments.


Complete Hardware Configuration Examples

Production NVIDIA Setup (CUDA 12.8, ≥16 GB VRAM)

export OPENAI_API_KEY=your_key_here

speech-to-speech \
    --stt parakeet-tdt \
    --llm_backend transformers \
    --tts qwen3 \
    --mode realtime

MacBook Pro M3 (On-Device Inference, No API Calls)

speech-to-speech \
    --stt whisper-mlx \
    --llm_backend mlx-lm \
    --tts qwen3 \
    --mode realtime

Headless Server (CPU-Only, API-Based LLM)

export OPENAI_API_KEY=your_key_here

speech-to-speech \
    --stt faster-whisper \
    --llm_backend responses-api \
    --tts qwen3 \
    --qwen3_tts_backend ggml \
    --mode realtime

Memory and Quantization Strategies

For VRAM-constrained systems, use MLX quantization on Apple Silicon. The benchmark_tts.py script validates these configurations:

Quantization Memory Reduction Quality Impact
bf16 2× vs FP32 Minimal
8bit 4× Slight
6bit 5.3× Moderate
4bit 8× Noticeable

On NVIDIA GPUs, select smaller models (2B–4B parameters) rather than quantizing, as the GGML backend does not expose runtime quantization in this release.


Summary

  • No training hardware guidance exists in the speech-to-speech repository; use upstream model documentation for training requirements.
  • NVIDIA CUDA GPUs with CUDA 12.x provide lowest latency for parallel STT/LLM/TTS inference on Linux/Windows.
  • Apple Silicon (M1/M2/M3) offers a fully native pipeline via MLX without discrete GPU.
  • CPU-only execution works for prototyping with substantially higher latency.
  • Benchmark scripts in scripts/benchmark_tts.py and scripts/benchmark_stt.py validate hardware-specific performance.
  • Memory requirements scale with model size: 31B LLMs need >30 GB VRAM; 2B–4B models fit 12–16 GB.

Frequently Asked Questions

Does the speech-to-speech repository support model training?

No. The repository contains inference-only code for running pretrained VAD, STT, LLM, and TTS models. Training scripts reside in upstream repositories (NVIDIA NeMo, OpenAI Whisper, Qwen3-TTS). For training hardware, consult those projects' documentation.

What is the minimum VRAM for local LLM inference?

12 GB VRAM suffices for 2B–4B parameter models (Qwen3-4B, Gemma-2B). 30+ GB VRAM is required for Gemma-4 31B. The mlx-lm backend on Apple Silicon uses unified memory, allowing larger models through quantization.

Can I run the full pipeline without any GPU?

Yes, but with trade-offs. Set --qwen3_tts_backend ggml with the CPU wheel, use --stt faster-whisper, and either --llm_backend responses-api (remote) or --llm_backend transformers (local CPU). Latency increases 5–10× versus GPU.

How do I verify my CUDA version matches Qwen3-TTS requirements?

Run nvidia-smi and check the "CUDA Version" line. Install the matching wheel from the Hugging Face wheelhouse before pip install speech-to-speech. Mismatched CUDA runtimes cause silent backend failures in TTS initialization.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →