GPT-SoVITS Inference Speed (RTF) Benchmark and Real-Time Optimization Guide
GPT-SoVITS v2 ProPlus achieves an inference RTF of 0.014 on an NVIDIA RTX 4090 (71× faster than real-time) and can be optimized further through half-precision inference, batch processing, and parallel generation to meet strict sub-100 ms latency requirements.
The RVC-Boss/GPT-SoVITS repository delivers state-of-the-art zero-shot text-to-speech synthesis with industry-leading speed metrics. Understanding the real-time factor (RTF)—the ratio of audio duration to processing time—is essential for deploying this model in live applications such as voice assistants, real-time dubbing, and interactive AI agents.
Official GPT-SoVITS RTF Benchmarks
The project’s README.md reports concrete RTF measurements for the latest v2 ProPlus checkpoint across different hardware configurations. An RTF value below 1.0 indicates generation faster than playback speed.
| Hardware | RTF | Real-Time Performance |
|---|---|---|
| NVIDIA RTX 4060 Ti | 0.028 | ~35× faster than real-time |
| NVIDIA RTX 4090 | 0.014 | ~71× faster (1,400 words / 4 min synthesizes in 3.36 s) |
| Apple M4 CPU | 0.526 | ~2× slower than real-time (usable for short prompts) |
These figures demonstrate that modern GPUs easily achieve the real-time threshold (RTF < 1), while CPU-only deployments require optimization for latency-sensitive use cases.
Where Inference Speed Is Controlled in the Code
Performance tuning in GPT-SoVITS centers on four key parameters located in specific source files:
inference_webui_fast.py(lines 50‑73) – Theinferencefunction constructs theTTSpipeline and yields audio chunks.speed_factor– Passed via theinputsdictionary (lines 86‑88) and applied inGPT_SoVITS/TTS_infer_pack/TTS.py(lines 1384‑1389) to scale the internal upsampling rate.parallel_infer– A Gradio checkbox defined at lines 418‑420 that toggles batch-wise concurrent generation.is_half– An environment variable read at line 52 that selectstorch.float16precision when available.tts_infer.yaml– Configuration file storing defaults forbatch_size,speed_factor,parallel_infer, andis_half.
Optimization Strategies for Real-Time Applications
Enable Half-Precision (FP16) Inference
Set is_half=True (auto-detected by default) to halve memory bandwidth and increase tensor throughput. In inference_webui_fast.py, this flag triggers torch.float16 dtype for all GPU tensors. Launch the Web UI with the --no_half flag only if your GPU lacks FP16 support.
Increase Batch Size
Adjust the batch_size parameter in GPT_SoVITS/configs/tts_infer.yaml or via the UI slider. Larger batches amortize the fixed overhead of the transformer encoder/decoder across multiple frames, improving throughput for bulk generation tasks.
Activate Parallel Inference
Enable the 并行推理 (parallel inference) checkbox in the Gradio interface or set "parallel_infer": true in the YAML config. This mode processes multiple audio fragments concurrently inside the TTS pipeline, exploiting GPU parallelism to reduce per-chunk latency.
Adjust the Speed Factor
Reduce speed_factor below 1.0 (e.g., 0.5 for 2× speed) in the configuration or UI slider. According to the implementation in TTS.py lines 1384‑1389, this parameter directly modifies the upsampling denominator, forcing the BigVGAN vocoder to generate fewer samples for the same semantic content. This trade-off lowers latency at the potential cost of audio fidelity.
Model Selection and Export
- Use lighter checkpoints: Switch from v2 ProPlus to standard v2 models via
weight.jsonto reduce parameter count. - Export to ONNX: Run
python GPT_SoVITS/onnx_export.py --weights v2_proplus.ckpt --output gpt_sovits.onnx --fp16. The exported model supportsonnxruntimeexecution, yielding 10‑15 % additional speedup on both GPU and CPU targets. - TorchScript: Use
export_torch_script.pyfor JIT-optimized graphs viatorch.jit.optimize_for_inference.
Hardware and Environment Tuning
- GPU Selection: Target compute capability ≥ 8.0 (Ampere/Ada Lovelace) for optimal FP16 performance.
- cuDNN Benchmarking: Add
torch.backends.cudnn.benchmark = Truetowebui.pyto enable automatic algorithm selection for convolution operations. - Sampling Rate: Lower the output
srintts_infer.yamlfrom 48 kHz to 24 kHz to halve vocoder workload, though this reduces high-frequency fidelity.
Practical Implementation Examples
Launching the Web UI with Maximum Performance
python webui.py --device cuda \
--batch_size 4 \
--speed_factor 1.0 \
--parallel_infer
Programmatic Inference with Custom Speed
from GPT_SoVITS.inference_webui_fast import inference, tts_config
# Configure for low-latency generation
tts_config.is_half = True
tts_config.update({
"batch_size": 2,
"speed_factor": 0.5,
"parallel_infer": False
})
# Stream audio chunks
for wav_chunk, seed in inference(
text="Low latency real-time synthesis.",
text_lang="en",
ref_audio_path="reference.wav",
batch_size=2,
speed_factor=0.5,
parallel_infer=False
):
# wav_chunk is a numpy float32 array ready for playback
process_audio(wav_chunk)
Exporting to ONNX for Edge Deployment
python GPT_SoVITS/onnx_export.py \
--weights path/to/v2_proplus.ckpt \
--output gpt_sovits.onnx \
--fp16
Load the resulting gpt_sovits.onnx with onnxruntime to achieve deterministic, low-overhead inference suitable for containerized or ARM-based edge devices.
Summary
- Baseline Performance: GPT-SoVITS v2 ProPlus achieves RTF 0.014 on an RTX 4090, enabling 71× real-time synthesis.
- Precision: Enable
is_half=Trueto reduce memory bandwidth by 50 %. - Throughput: Increase
batch_sizeand activateparallel_inferto maximize GPU utilization. - Latency: Reduce
speed_factorto trade audio length for generation speed. - Deployment: Export to ONNX/TorchScript for cross-platform, ultra-low-latency serving.
Frequently Asked Questions
What is considered a good RTF value for real-time TTS applications?
An RTF below 0.1 is generally considered excellent for streaming, ensuring that 10 seconds of audio generates in under 1 second. GPT-SoVITS achieves RTF 0.014 on high-end GPUs, providing ample headroom for network transmission and audio buffering in live scenarios.
Does reducing the speed_factor degrade voice quality?
Reducing speed_factor to 0.5‑0.8 produces intelligible speech but may compress prosody and alter speaking cadence. For critical applications, benchmark perceived quality with human evaluators when operating below 1.0.
Can GPT-SoVITS achieve real-time inference without a GPU?
On Apple M4 CPUs, the model achieves RTF 0.526 (roughly half real-time speed), suitable for short asynchronous clips but insufficient for streaming. Exporting to ONNX and using optimized runtimes like onnxruntime can improve CPU performance, though a discrete GPU remains recommended for live applications.
Where exactly is the speed_factor parameter applied in the source code?
The parameter is consumed in GPT_SoVITS/TTS_infer_pack/TTS.py at lines 1384‑1389, where it scales the upsample_rate variable inside the TTS pipeline’s forward pass, directly influencing the temporal resolution of the generated waveform.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →