# GPT-SoVITS Inference Performance: FP16 (`is_half=True`) vs Full Precision Explained

> Boost GPT-SoVITS inference speed up to 2x with FP16 half precision. Reduce memory by 50% on Tensor-Core GPUs with minimal quality loss. Learn the performance differences.

- Repository: [RVC-Boss/GPT-SoVITS](https://github.com/RVC-Boss/GPT-SoVITS)
- Tags: performance
- Published: 2026-03-07

---

**Running GPT-SoVITS inference with `is_half=True` enables half-precision (FP16) mode, delivering 1.5–2× speed-ups and 50% memory reduction on Tensor-Core GPUs with minimal quality trade-offs.**

The `is_half` flag in the [RVC-Boss/GPT-SoVITS](https://github.com/RVC-Boss/GPT-SoVITS) repository controls whether the text-to-speech pipeline operates in **float16 (FP16)** or **float32 (FP32)** precision. This single configuration parameter propagates through every component—from BERT embedding extraction to waveform generation—directly impacting latency, VRAM usage, and audio fidelity.

## How `is_half=True` Works Internally

When enabled, the flag triggers systematic dtype conversion across the inference stack. Understanding where these conversions occur helps diagnose compatibility issues and performance bottlenecks.

### Environment Detection in [`config.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/config.py)

The pipeline first validates GPU capabilities in [`config.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/config.py). The code checks compute capability (SM ≥ 6.1) or 16-series GPU architecture to ensure Tensor-Core support before allowing FP16 initialization.

```python

# config.py

is_half = any(dtype == torch.float16 for _, dtype, _, _ in tmp)

```

This detection feeds into `config.get_device_dtype_sm`, which returns `torch.float16` for compatible GPUs and `torch.float32` for older hardware or CPU-only environments.

### Model Weight Conversion

During model loading in [`inference_webui.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/inference_webui.py), all transformer and acoustic model weights are explicitly cast to FP16 when the flag is active.

```python

# inference_webui.py excerpt

if is_half:
    bert_model = bert_model.half().to(device)   # FP16 BERT embeddings

else:
    bert_model = bert_model.to(device)          # FP32 BERT embeddings

```

The SV (speaker verification) model in [`sv.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/sv.py) stores `self.is_half` as an instance variable, applying `.half()` to every intermediate tensor during the speech synthesis forward pass.

### Audio Buffer Allocation

The precision setting extends to output buffers. In [`api.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/api.py), the zero-padding waveform buffer uses conditional dtype selection to match the model's operating mode.

```python

# api.py

zero_wav = np.zeros(int(hps.data.sampling_rate * 0.3),
                    dtype=np.float16 if is_half else np.float32)

```

## Performance Benefits of FP16 Inference

**Reduced GPU Memory Consumption** — FP16 tensors occupy half the storage of FP32 equivalents. On a 12 GB GPU, this allows doubling batch sizes or processing longer audio segments without out-of-memory errors.

**Higher Throughput via Tensor Cores** — Modern NVIDIA GPUs (RTX 30xx/40xx series) leverage dedicated Tensor Cores optimized for FP16 matrix multiplication. This yields **1.5×–2× speed-ups** during the autoregressive sampling and vocoder phases.

**Faster Data Transfer** — Smaller tensor footprints reduce PCIe bandwidth pressure between host and device, providing marginal latency improvements in CPU-bound preprocessing pipelines.

## Trade-offs and Compatibility Considerations

While FP16 accelerates inference, several constraints govern its safe deployment.

### Numerical Precision Loss

FP16 provides approximately 3× less mantissa precision than FP32. Although GPT-SoVITS trains in full precision, aggressive quantization can cause subtle acoustic artifacts—slightly noisier harmonics or less stable pitch contours in edge cases.

### Hardware Requirements

FP16 inference requires GPUs with Tensor-Core support (SM ≥ 6.1) or 16-series cards (e.g., GTX 1660). On older architectures or CPU-only systems, `config.get_device_dtype_sm` silently falls back to FP32 regardless of the flag setting.

### Operator Support Overhead

Certain PyTorch operations lack native FP16 kernels. When encountering these, the framework automatically promotes tensors to FP32, computes the result, then casts back—introducing small synchronization penalties. The codebase guards against this by explicitly converting tensors to FP32 only when required for specific mathematical operations.

### Deployment Constraints

When exporting models via ONNX for production deployment, ensure the target runtime supports FP16 execution. If the inference engine expects FP32 inputs, you must export a full-precision version to avoid runtime dtype mismatches.

## Configuring Precision in GPT-SoVITS

Control the precision mode via environment variables or CLI arguments in [`webui.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/webui.py).

**Enable half-precision (default on compatible GPUs):**

```bash
python webui.py

```

**Force full-precision mode:**

```bash
python webui.py -fp

```

Override the environment variable directly for API usage:

```python
import os
os.environ["is_half"] = "False"  # Forces FP32

```

## Summary

- **`is_half=True`** activates FP16 inference throughout [`inference_webui.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/inference_webui.py), [`sv.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/sv.py), and [`api.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/api.py), reducing VRAM usage by 50%.
- **Performance gains** of 1.5×–2× are typical on RTX 30/40 series GPUs due to Tensor-Core acceleration.
- **Compatibility** requires SM ≥ 6.1; older GPUs automatically fall back to FP32 via [`config.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/config.py) detection logic.
- **Quality impact** is generally marginal, though extreme dynamic ranges may exhibit minor quantization noise.
- **Configuration** is controlled via the `-fp` CLI flag or the `is_half` environment variable parsed in [`webui.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/webui.py).

## Frequently Asked Questions

### Does `is_half=True` affect audio quality in GPT-SoVITS?

The impact is typically minimal for most use cases. Since the model trains in FP32, FP16 inference introduces slight quantization noise that may manifest as subtle harmonic instability in quiet passages. However, the perceptual difference is negligible for standard TTS applications, making FP16 the recommended default for real-time inference.

### Which GPUs support FP16 inference in GPT-SoVITS?

GPUs with NVIDIA Tensor Cores (compute capability SM ≥ 6.1) or 16-series cards like the GTX 1660 support FP16. The [`config.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/config.py) detection logic automatically disables FP16 on incompatible hardware (e.g., GTX 10-series or older) to prevent runtime errors, falling back to FP32 silently.

### How do I force full precision if `is_half` is causing issues?

Pass the `-fp` flag when launching the web UI: `python webui.py -fp`. This overrides the environment variable and forces FP32 mode across all model loading routines in [`inference_webui.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/inference_webui.py) and [`sv.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/sv.py), eliminating any FP16-specific numerical instability at the cost of reduced performance.

### Can I export GPT-SoVITS models in FP16 format?

Yes, but verify your deployment target supports FP16 execution natively. When exporting through `tools/` scripts, ensure the ONNX runtime or TensorRT engine is configured for mixed precision. If deploying to hardware without FP16 support, maintain FP32 weights to avoid automatic casting overhead during inference.