# GPT-SoVITS V4 vs V3 Audio Quality: 24 kHz vs 48 kHz Native Output Differences

> Discover GPT-SoVITS V4's superior audio quality. Learn how native 48 kHz output eliminates metallic artifacts and preserves high frequencies, outperforming V3's 24 kHz generation.

- Repository: [RVC-Boss/GPT-SoVITS](https://github.com/RVC-Boss/GPT-SoVITS)
- Tags: performance
- Published: 2026-03-07

---

**GPT-SoVITS V4 eliminates the metallic artifacts present in V3 by natively generating 48 kHz audio instead of 24 kHz, preserving higher-frequency details without destructive upsampling interpolation.**

The **RVC-Boss/GPT-SoVITS** repository introduced a fundamental architectural change in V4 that directly impacts output fidelity. While V3 was trained to synthesize speech at a **native sampling rate of 24 kHz**, V4 upgrades the vocoder pipeline to output **48 kHz natively**, resolving muffling issues caused by non-integer upsampling factors in the previous version.

## Why V4 Delivers Higher-Fidelity Audio

The audio quality difference stems from two critical fixes implemented in V4:

- **Elimination of metallic artifacts**: V3 suffered from "muffled" metallic artifacts caused by non-integer upsampling factors during audio generation.
- **Doubled native sampling rate**: V4 moves the target sample rate from 24 kHz to 48 kHz, allowing the model to synthesize high-frequency content directly rather than interpolating it afterward.

According to the project documentation in [`README.md`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/README.md) (lines 337-341), *"V4 fixes the issue of metallic artifacts in Version 3 … and natively outputs 48 kHz audio … (V3 only natively outputs 24 kHz audio)"*. This claim is corroborated across multilingual documentation in [`docs/ko/README.md`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/docs/ko/README.md), [`docs/ja/README.md`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/docs/ja/README.md), and [`docs/tr/README.md`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/docs/tr/README.md), which confirm that **V4 기본적으로 48 kHz 오디오를 출력합니다** (V4 basically outputs 48 kHz audio).

## Technical Architecture: Configs and Sample Rates

### V3 Configuration: 24 kHz Target

In V3, the mel-spectrogram and vocoder pipeline hardcodes a **target sample rate of 24 kHz**. In [`GPT_SoVITS/f5_tts/model/modules.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/GPT_SoVITS/f5_tts/model/modules.py) (lines 34-45), the model modules explicitly set `target_sample_rate=24000`, constraining the output bandwidth to telephone-quality frequencies.

The configuration file [`GPT_SoVITS/configs/s2v3.json`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/GPT_SoVITS/configs/s2v3.json) maintains this constraint, ensuring backward compatibility with legacy models trained on 24 kHz datasets.

### V4 Configuration: 48 kHz Native Output

V4 shifts the vocoder's `upsample_rate` to **480** (representing 48 kHz) in [`GPT_SoVITS/configs/s2v4.json`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/GPT_SoVITS/configs/s2v4.json). This change allows the neural vocoder to generate waveforms with full broadband fidelity, capturing subtle harmonics and sibilance that 24 kHz down-sampling destroys.

The inference logic in [`GPT_SoVITS/TTS_infer_pack/TTS.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/GPT_SoVITS/TTS_infer_pack/TTS.py) (lines 1305-1409) reads this configuration value and applies it directly to the generated waveform, ensuring the output buffer matches the native 48 kHz specification without post-processing interpolation.

### Version Selection Mechanism

The WebUI and programmatic API control the output quality through an environment variable that selects the appropriate JSON configuration:

```python

# webui.py (excerpt)

import os
os.environ["version"] = version = "v4"   # set to v3, v4, v2Pro, etc.

...

# The chosen version determines which config JSON is loaded:

config_path = f"GPT_SoVITS/configs/s2{version}.json"

```

When `version` is set to `"v4"`, the loader imports [`s2v4.json`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/s2v4.json) with its 48 kHz vocoder settings. Setting it to `"v3"` loads the legacy 24 kHz configuration.

## Practical Implementation Examples

### Checking Your Current Output Rate

```python
from GPT_SoVITS.TTS_infer_pack.TTS import TTS

# Initialize with V4 for 48 kHz native output

tts_v4 = TTS(version="v4")
print(f"V4 upsample rate: {tts_v4.vocoder_configs['upsample_rate']} -> 48 kHz output")

# Initialize with V3 for comparison

tts_v3 = TTS(version="v3")
print(f"V3 upsample rate: {tts_v3.vocoder_configs['upsample_rate']} -> 24 kHz output")

```

The `TTS` class automatically populates `vocoder_configs` from the version-specific JSON file, exposing the `upsample_rate` parameter that dictates the final PCM sample rate.

### Generating Comparison Audio Files

```python
from GPT_SoVITS.TTS_infer_pack.TTS import TTS
import soundfile as sf

text = "This is a quality comparison between V3 and V4 sampling rates."

for version in ["v3", "v4"]:
    tts = TTS(version=version)
    wav = tts.infer(text)
    
    # Extract native sample rate from config

    sr = tts.vocoder_configs["upsample_rate"] * 100  # 240 or 480 -> 24000 or 48000

    sf.write(f"output_{version}.wav", wav, samplerate=sr)
    print(f"{version}: Generated {sr} Hz audio")

```

Running this script produces `output_v3.wav` at **24 kHz** and `output_v4.wav` at **48 kHz**, demonstrating the native-rate difference without resampling artifacts.

### Verifying Configuration Files Directly

```json
// GPT_SoVITS/configs/s2v4.json (excerpt)
{
  "vocoder": {
    "upsample_rate": 480,   // 48 kHz native output
    "type": "nsf_hifigan"
  }
}

```

```json
// GPT_SoVITS/configs/s2v3.json (excerpt)
{
  "vocoder": {
    "upsample_rate": 240,   // 24 kHz native output
    "type": "nsf_hifigan"
  }
}

```

The `upsample_rate` value represents the target sample rate divided by 100, serving as the authoritative source for the synthesis pipeline's output frequency.

## Summary

- **GPT-SoVITS V4** outputs **48 kHz audio natively**, while **V3** is limited to **24 kHz**.
- V4 eliminates **metallic artifacts** caused by non-integer upsampling in V3's post-processing pipeline.
- The `version` environment variable in [`webui.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/webui.py) controls which configuration ([`s2v3.json`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/s2v3.json) or [`s2v4.json`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/s2v4.json)) the `TTS` class loads.
- **Higher-frequency preservation** in V4 results from the vocoder's `upsample_rate` of **480** (48 kHz) versus V3's **240** (24 kHz).
- Both versions are accessible programmatically via `TTS(version="v3")` or `TTS(version="v4")`.

## Frequently Asked Questions

### Can I upsample V3's 24 kHz output to 48 kHz without quality loss?

No. According to the release notes in [`docs/tr/README.md`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/docs/tr/README.md), V3's 24 kHz audio lacks the high-frequency content above 12 kHz (Nyquist limit). While V4 includes a **24 kHz → 48 kHz audio super-resolution model** to enhance legacy V3-generated files, native V4 synthesis produces superior results by generating those frequencies during inference rather than reconstructing them afterward.

### Does V4 require more GPU memory than V3?

Yes. Processing 48 kHz audio doubles the temporal resolution of the waveform buffers, increasing VRAM usage during vocoder inference. The [`TTS.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/TTS.py) implementation (lines 1305-1409) handles larger tensor dimensions for the 48 kHz path compared to the 24 kHz V3 pipeline.

### Is V4 backward compatible with V3 models?

Partially. While the `TTS` class supports version selection via the `version` parameter, V3-trained checkpoints are optimized for 24 kHz mel-spectrograms. For best results with V4's 48 kHz native output, use models fine-tuned on the V4 configuration, though the codebase maintains compatibility layers for mixed-version workflows.

### Why does V3 have metallic artifacts while V4 doesn't?

V3's artifacts originate from **non-integer upsampling factors** in the audio generation pipeline. When the 24 kHz native output is resampled to higher rates for playback, interpolation errors create metallic, muffled timbres. V4 avoids this by synthesizing at 48 kHz directly, bypassing the destructive upsampling step entirely as documented in [`README.md`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/README.md) (lines 337-341).