GPT-SoVITS Inference Speed (RTF) Benchmark and Real-Time Optimization Guide

GPT-SoVITS v2 ProPlus achieves an inference RTF of 0.014 on an NVIDIA RTX 4090 (71× faster than real-time) and can be optimized further through half-precision inference, batch processing, and parallel generation to meet strict sub-100 ms latency requirements.

The RVC-Boss/GPT-SoVITS repository delivers state-of-the-art zero-shot text-to-speech synthesis with industry-leading speed metrics. Understanding the real-time factor (RTF)—the ratio of audio duration to processing time—is essential for deploying this model in live applications such as voice assistants, real-time dubbing, and interactive AI agents.

Official GPT-SoVITS RTF Benchmarks

The project’s README.md reports concrete RTF measurements for the latest v2 ProPlus checkpoint across different hardware configurations. An RTF value below 1.0 indicates generation faster than playback speed.

Hardware RTF Real-Time Performance
NVIDIA RTX 4060 Ti 0.028 ~35× faster than real-time
NVIDIA RTX 4090 0.014 ~71× faster (1,400 words / 4 min synthesizes in 3.36 s)
Apple M4 CPU 0.526 ~2× slower than real-time (usable for short prompts)

These figures demonstrate that modern GPUs easily achieve the real-time threshold (RTF < 1), while CPU-only deployments require optimization for latency-sensitive use cases.

Where Inference Speed Is Controlled in the Code

Performance tuning in GPT-SoVITS centers on four key parameters located in specific source files:

  • inference_webui_fast.py (lines 50‑73) – The inference function constructs the TTS pipeline and yields audio chunks.
  • speed_factor – Passed via the inputs dictionary (lines 86‑88) and applied in GPT_SoVITS/TTS_infer_pack/TTS.py (lines 1384‑1389) to scale the internal upsampling rate.
  • parallel_infer – A Gradio checkbox defined at lines 418‑420 that toggles batch-wise concurrent generation.
  • is_half – An environment variable read at line 52 that selects torch.float16 precision when available.
  • tts_infer.yaml – Configuration file storing defaults for batch_size, speed_factor, parallel_infer, and is_half.

Optimization Strategies for Real-Time Applications

Enable Half-Precision (FP16) Inference

Set is_half=True (auto-detected by default) to halve memory bandwidth and increase tensor throughput. In inference_webui_fast.py, this flag triggers torch.float16 dtype for all GPU tensors. Launch the Web UI with the --no_half flag only if your GPU lacks FP16 support.

Increase Batch Size

Adjust the batch_size parameter in GPT_SoVITS/configs/tts_infer.yaml or via the UI slider. Larger batches amortize the fixed overhead of the transformer encoder/decoder across multiple frames, improving throughput for bulk generation tasks.

Activate Parallel Inference

Enable the 并行推理 (parallel inference) checkbox in the Gradio interface or set "parallel_infer": true in the YAML config. This mode processes multiple audio fragments concurrently inside the TTS pipeline, exploiting GPU parallelism to reduce per-chunk latency.

Adjust the Speed Factor

Reduce speed_factor below 1.0 (e.g., 0.5 for 2× speed) in the configuration or UI slider. According to the implementation in TTS.py lines 1384‑1389, this parameter directly modifies the upsampling denominator, forcing the BigVGAN vocoder to generate fewer samples for the same semantic content. This trade-off lowers latency at the potential cost of audio fidelity.

Model Selection and Export

  • Use lighter checkpoints: Switch from v2 ProPlus to standard v2 models via weight.json to reduce parameter count.
  • Export to ONNX: Run python GPT_SoVITS/onnx_export.py --weights v2_proplus.ckpt --output gpt_sovits.onnx --fp16. The exported model supports onnxruntime execution, yielding 10‑15 % additional speedup on both GPU and CPU targets.
  • TorchScript: Use export_torch_script.py for JIT-optimized graphs via torch.jit.optimize_for_inference.

Hardware and Environment Tuning

  • GPU Selection: Target compute capability ≥ 8.0 (Ampere/Ada Lovelace) for optimal FP16 performance.
  • cuDNN Benchmarking: Add torch.backends.cudnn.benchmark = True to webui.py to enable automatic algorithm selection for convolution operations.
  • Sampling Rate: Lower the output sr in tts_infer.yaml from 48 kHz to 24 kHz to halve vocoder workload, though this reduces high-frequency fidelity.

Practical Implementation Examples

Launching the Web UI with Maximum Performance

python webui.py --device cuda \
    --batch_size 4 \
    --speed_factor 1.0 \
    --parallel_infer

Programmatic Inference with Custom Speed

from GPT_SoVITS.inference_webui_fast import inference, tts_config

# Configure for low-latency generation

tts_config.is_half = True
tts_config.update({
    "batch_size": 2,
    "speed_factor": 0.5,
    "parallel_infer": False
})

# Stream audio chunks

for wav_chunk, seed in inference(
    text="Low latency real-time synthesis.",
    text_lang="en",
    ref_audio_path="reference.wav",
    batch_size=2,
    speed_factor=0.5,
    parallel_infer=False
):
    # wav_chunk is a numpy float32 array ready for playback

    process_audio(wav_chunk)

Exporting to ONNX for Edge Deployment

python GPT_SoVITS/onnx_export.py \
    --weights path/to/v2_proplus.ckpt \
    --output gpt_sovits.onnx \
    --fp16

Load the resulting gpt_sovits.onnx with onnxruntime to achieve deterministic, low-overhead inference suitable for containerized or ARM-based edge devices.

Summary

  • Baseline Performance: GPT-SoVITS v2 ProPlus achieves RTF 0.014 on an RTX 4090, enabling 71× real-time synthesis.
  • Precision: Enable is_half=True to reduce memory bandwidth by 50 %.
  • Throughput: Increase batch_size and activate parallel_infer to maximize GPU utilization.
  • Latency: Reduce speed_factor to trade audio length for generation speed.
  • Deployment: Export to ONNX/TorchScript for cross-platform, ultra-low-latency serving.

Frequently Asked Questions

What is considered a good RTF value for real-time TTS applications?

An RTF below 0.1 is generally considered excellent for streaming, ensuring that 10 seconds of audio generates in under 1 second. GPT-SoVITS achieves RTF 0.014 on high-end GPUs, providing ample headroom for network transmission and audio buffering in live scenarios.

Does reducing the speed_factor degrade voice quality?

Reducing speed_factor to 0.5‑0.8 produces intelligible speech but may compress prosody and alter speaking cadence. For critical applications, benchmark perceived quality with human evaluators when operating below 1.0.

Can GPT-SoVITS achieve real-time inference without a GPU?

On Apple M4 CPUs, the model achieves RTF 0.526 (roughly half real-time speed), suitable for short asynchronous clips but insufficient for streaming. Exporting to ONNX and using optimized runtimes like onnxruntime can improve CPU performance, though a discrete GPU remains recommended for live applications.

Where exactly is the speed_factor parameter applied in the source code?

The parameter is consumed in GPT_SoVITS/TTS_infer_pack/TTS.py at lines 1384‑1389, where it scales the upsample_rate variable inside the TTS pipeline’s forward pass, directly influencing the temporal resolution of the generated waveform.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →