# GPT-SoVITS Inference Speed (RTF) Benchmark and Real-Time Optimization Guide

> Discover GPT-SoVITS inference speed benchmarks and learn how to optimize for real-time applications. Achieve sub-100ms latency with half-precision inference batching and parallel generation.

- Repository: [RVC-Boss/GPT-SoVITS](https://github.com/RVC-Boss/GPT-SoVITS)
- Tags: performance
- Published: 2026-03-07

---

**GPT-SoVITS v2 ProPlus achieves an inference RTF of 0.014 on an NVIDIA RTX 4090 (71× faster than real-time) and can be optimized further through half-precision inference, batch processing, and parallel generation to meet strict sub-100 ms latency requirements.**

The RVC-Boss/GPT-SoVITS repository delivers state-of-the-art zero-shot text-to-speech synthesis with industry-leading speed metrics. Understanding the **real-time factor (RTF)**—the ratio of audio duration to processing time—is essential for deploying this model in live applications such as voice assistants, real-time dubbing, and interactive AI agents.

## Official GPT-SoVITS RTF Benchmarks

The project’s README.md reports concrete RTF measurements for the latest **v2 ProPlus** checkpoint across different hardware configurations. An RTF value below 1.0 indicates generation faster than playback speed.

| Hardware | RTF | Real-Time Performance |
|----------|-----|----------------------|
| NVIDIA RTX 4060 Ti | **0.028** | ~35× faster than real-time |
| NVIDIA RTX 4090 | **0.014** | ~71× faster (1,400 words / 4 min synthesizes in 3.36 s) |
| Apple M4 CPU | **0.526** | ~2× slower than real-time (usable for short prompts) |

These figures demonstrate that modern GPUs easily achieve the **real-time threshold** (RTF < 1), while CPU-only deployments require optimization for latency-sensitive use cases.

## Where Inference Speed Is Controlled in the Code

Performance tuning in GPT-SoVITS centers on four key parameters located in specific source files:

- **[`inference_webui_fast.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/inference_webui_fast.py)** (lines 50‑73) – The `inference` function constructs the `TTS` pipeline and yields audio chunks.
- **`speed_factor`** – Passed via the `inputs` dictionary (lines 86‑88) and applied in [`GPT_SoVITS/TTS_infer_pack/TTS.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/GPT_SoVITS/TTS_infer_pack/TTS.py) (lines 1384‑1389) to scale the internal upsampling rate.
- **`parallel_infer`** – A Gradio checkbox defined at lines 418‑420 that toggles batch-wise concurrent generation.
- **`is_half`** – An environment variable read at line 52 that selects `torch.float16` precision when available.
- **[`tts_infer.yaml`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/tts_infer.yaml)** – Configuration file storing defaults for `batch_size`, `speed_factor`, `parallel_infer`, and `is_half`.

## Optimization Strategies for Real-Time Applications

### Enable Half-Precision (FP16) Inference

Set `is_half=True` (auto-detected by default) to halve memory bandwidth and increase tensor throughput. In [`inference_webui_fast.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/inference_webui_fast.py), this flag triggers `torch.float16` dtype for all GPU tensors. Launch the Web UI with the `--no_half` flag only if your GPU lacks FP16 support.

### Increase Batch Size

Adjust the `batch_size` parameter in [`GPT_SoVITS/configs/tts_infer.yaml`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/GPT_SoVITS/configs/tts_infer.yaml) or via the UI slider. Larger batches amortize the fixed overhead of the transformer encoder/decoder across multiple frames, improving throughput for bulk generation tasks.

### Activate Parallel Inference

Enable the **并行推理** (parallel inference) checkbox in the Gradio interface or set `"parallel_infer": true` in the YAML config. This mode processes multiple audio fragments concurrently inside the `TTS` pipeline, exploiting GPU parallelism to reduce per-chunk latency.

### Adjust the Speed Factor

Reduce `speed_factor` below 1.0 (e.g., 0.5 for 2× speed) in the configuration or UI slider. According to the implementation in [`TTS.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/TTS.py) lines 1384‑1389, this parameter directly modifies the upsampling denominator, forcing the BigVGAN vocoder to generate fewer samples for the same semantic content. This trade-off lowers latency at the potential cost of audio fidelity.

### Model Selection and Export

- **Use lighter checkpoints**: Switch from v2 ProPlus to standard v2 models via [`weight.json`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/weight.json) to reduce parameter count.
- **Export to ONNX**: Run `python GPT_SoVITS/onnx_export.py --weights v2_proplus.ckpt --output gpt_sovits.onnx --fp16`. The exported model supports `onnxruntime` execution, yielding **10‑15 %** additional speedup on both GPU and CPU targets.
- **TorchScript**: Use [`export_torch_script.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/export_torch_script.py) for JIT-optimized graphs via `torch.jit.optimize_for_inference`.

### Hardware and Environment Tuning

- **GPU Selection**: Target compute capability ≥ 8.0 (Ampere/Ada Lovelace) for optimal FP16 performance.
- **cuDNN Benchmarking**: Add `torch.backends.cudnn.benchmark = True` to [`webui.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/webui.py) to enable automatic algorithm selection for convolution operations.
- **Sampling Rate**: Lower the output `sr` in [`tts_infer.yaml`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/tts_infer.yaml) from 48 kHz to 24 kHz to halve vocoder workload, though this reduces high-frequency fidelity.

## Practical Implementation Examples

### Launching the Web UI with Maximum Performance

```bash
python webui.py --device cuda \
    --batch_size 4 \
    --speed_factor 1.0 \
    --parallel_infer

```

### Programmatic Inference with Custom Speed

```python
from GPT_SoVITS.inference_webui_fast import inference, tts_config

# Configure for low-latency generation

tts_config.is_half = True
tts_config.update({
    "batch_size": 2,
    "speed_factor": 0.5,
    "parallel_infer": False
})

# Stream audio chunks

for wav_chunk, seed in inference(
    text="Low latency real-time synthesis.",
    text_lang="en",
    ref_audio_path="reference.wav",
    batch_size=2,
    speed_factor=0.5,
    parallel_infer=False
):
    # wav_chunk is a numpy float32 array ready for playback

    process_audio(wav_chunk)

```

### Exporting to ONNX for Edge Deployment

```bash
python GPT_SoVITS/onnx_export.py \
    --weights path/to/v2_proplus.ckpt \
    --output gpt_sovits.onnx \
    --fp16

```

Load the resulting `gpt_sovits.onnx` with `onnxruntime` to achieve deterministic, low-overhead inference suitable for containerized or ARM-based edge devices.

## Summary

- **Baseline Performance**: GPT-SoVITS v2 ProPlus achieves RTF 0.014 on an RTX 4090, enabling 71× real-time synthesis.
- **Precision**: Enable `is_half=True` to reduce memory bandwidth by 50 %.
- **Throughput**: Increase `batch_size` and activate `parallel_infer` to maximize GPU utilization.
- **Latency**: Reduce `speed_factor` to trade audio length for generation speed.
- **Deployment**: Export to ONNX/TorchScript for cross-platform, ultra-low-latency serving.

## Frequently Asked Questions

### What is considered a good RTF value for real-time TTS applications?

An RTF below 0.1 is generally considered excellent for streaming, ensuring that 10 seconds of audio generates in under 1 second. GPT-SoVITS achieves RTF 0.014 on high-end GPUs, providing ample headroom for network transmission and audio buffering in live scenarios.

### Does reducing the speed_factor degrade voice quality?

Reducing `speed_factor` to 0.5‑0.8 produces intelligible speech but may compress prosody and alter speaking cadence. For critical applications, benchmark perceived quality with human evaluators when operating below 1.0.

### Can GPT-SoVITS achieve real-time inference without a GPU?

On Apple M4 CPUs, the model achieves RTF 0.526 (roughly half real-time speed), suitable for short asynchronous clips but insufficient for streaming. Exporting to ONNX and using optimized runtimes like `onnxruntime` can improve CPU performance, though a discrete GPU remains recommended for live applications.

### Where exactly is the speed_factor parameter applied in the source code?

The parameter is consumed in [`GPT_SoVITS/TTS_infer_pack/TTS.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/GPT_SoVITS/TTS_infer_pack/TTS.py) at lines 1384‑1389, where it scales the `upsample_rate` variable inside the TTS pipeline’s forward pass, directly influencing the temporal resolution of the generated waveform.