GPT-SoVITS V4 vs V3 Audio Quality: 24 kHz vs 48 kHz Native Output Differences

GPT-SoVITS V4 eliminates the metallic artifacts present in V3 by natively generating 48 kHz audio instead of 24 kHz, preserving higher-frequency details without destructive upsampling interpolation.

The RVC-Boss/GPT-SoVITS repository introduced a fundamental architectural change in V4 that directly impacts output fidelity. While V3 was trained to synthesize speech at a native sampling rate of 24 kHz, V4 upgrades the vocoder pipeline to output 48 kHz natively, resolving muffling issues caused by non-integer upsampling factors in the previous version.

Why V4 Delivers Higher-Fidelity Audio

The audio quality difference stems from two critical fixes implemented in V4:

  • Elimination of metallic artifacts: V3 suffered from "muffled" metallic artifacts caused by non-integer upsampling factors during audio generation.
  • Doubled native sampling rate: V4 moves the target sample rate from 24 kHz to 48 kHz, allowing the model to synthesize high-frequency content directly rather than interpolating it afterward.

According to the project documentation in README.md (lines 337-341), "V4 fixes the issue of metallic artifacts in Version 3 … and natively outputs 48 kHz audio … (V3 only natively outputs 24 kHz audio)". This claim is corroborated across multilingual documentation in docs/ko/README.md, docs/ja/README.md, and docs/tr/README.md, which confirm that V4 기본적으로 48 kHz 오디오를 출력합니다 (V4 basically outputs 48 kHz audio).

Technical Architecture: Configs and Sample Rates

V3 Configuration: 24 kHz Target

In V3, the mel-spectrogram and vocoder pipeline hardcodes a target sample rate of 24 kHz. In GPT_SoVITS/f5_tts/model/modules.py (lines 34-45), the model modules explicitly set target_sample_rate=24000, constraining the output bandwidth to telephone-quality frequencies.

The configuration file GPT_SoVITS/configs/s2v3.json maintains this constraint, ensuring backward compatibility with legacy models trained on 24 kHz datasets.

V4 Configuration: 48 kHz Native Output

V4 shifts the vocoder's upsample_rate to 480 (representing 48 kHz) in GPT_SoVITS/configs/s2v4.json. This change allows the neural vocoder to generate waveforms with full broadband fidelity, capturing subtle harmonics and sibilance that 24 kHz down-sampling destroys.

The inference logic in GPT_SoVITS/TTS_infer_pack/TTS.py (lines 1305-1409) reads this configuration value and applies it directly to the generated waveform, ensuring the output buffer matches the native 48 kHz specification without post-processing interpolation.

Version Selection Mechanism

The WebUI and programmatic API control the output quality through an environment variable that selects the appropriate JSON configuration:


# webui.py (excerpt)

import os
os.environ["version"] = version = "v4"   # set to v3, v4, v2Pro, etc.

...

# The chosen version determines which config JSON is loaded:

config_path = f"GPT_SoVITS/configs/s2{version}.json"

When version is set to "v4", the loader imports s2v4.json with its 48 kHz vocoder settings. Setting it to "v3" loads the legacy 24 kHz configuration.

Practical Implementation Examples

Checking Your Current Output Rate

from GPT_SoVITS.TTS_infer_pack.TTS import TTS

# Initialize with V4 for 48 kHz native output

tts_v4 = TTS(version="v4")
print(f"V4 upsample rate: {tts_v4.vocoder_configs['upsample_rate']} -> 48 kHz output")

# Initialize with V3 for comparison

tts_v3 = TTS(version="v3")
print(f"V3 upsample rate: {tts_v3.vocoder_configs['upsample_rate']} -> 24 kHz output")

The TTS class automatically populates vocoder_configs from the version-specific JSON file, exposing the upsample_rate parameter that dictates the final PCM sample rate.

Generating Comparison Audio Files

from GPT_SoVITS.TTS_infer_pack.TTS import TTS
import soundfile as sf

text = "This is a quality comparison between V3 and V4 sampling rates."

for version in ["v3", "v4"]:
    tts = TTS(version=version)
    wav = tts.infer(text)
    
    # Extract native sample rate from config

    sr = tts.vocoder_configs["upsample_rate"] * 100  # 240 or 480 -> 24000 or 48000

    sf.write(f"output_{version}.wav", wav, samplerate=sr)
    print(f"{version}: Generated {sr} Hz audio")

Running this script produces output_v3.wav at 24 kHz and output_v4.wav at 48 kHz, demonstrating the native-rate difference without resampling artifacts.

Verifying Configuration Files Directly

// GPT_SoVITS/configs/s2v4.json (excerpt)
{
  "vocoder": {
    "upsample_rate": 480,   // 48 kHz native output
    "type": "nsf_hifigan"
  }
}
// GPT_SoVITS/configs/s2v3.json (excerpt)
{
  "vocoder": {
    "upsample_rate": 240,   // 24 kHz native output
    "type": "nsf_hifigan"
  }
}

The upsample_rate value represents the target sample rate divided by 100, serving as the authoritative source for the synthesis pipeline's output frequency.

Summary

  • GPT-SoVITS V4 outputs 48 kHz audio natively, while V3 is limited to 24 kHz.
  • V4 eliminates metallic artifacts caused by non-integer upsampling in V3's post-processing pipeline.
  • The version environment variable in webui.py controls which configuration (s2v3.json or s2v4.json) the TTS class loads.
  • Higher-frequency preservation in V4 results from the vocoder's upsample_rate of 480 (48 kHz) versus V3's 240 (24 kHz).
  • Both versions are accessible programmatically via TTS(version="v3") or TTS(version="v4").

Frequently Asked Questions

Can I upsample V3's 24 kHz output to 48 kHz without quality loss?

No. According to the release notes in docs/tr/README.md, V3's 24 kHz audio lacks the high-frequency content above 12 kHz (Nyquist limit). While V4 includes a 24 kHz → 48 kHz audio super-resolution model to enhance legacy V3-generated files, native V4 synthesis produces superior results by generating those frequencies during inference rather than reconstructing them afterward.

Does V4 require more GPU memory than V3?

Yes. Processing 48 kHz audio doubles the temporal resolution of the waveform buffers, increasing VRAM usage during vocoder inference. The TTS.py implementation (lines 1305-1409) handles larger tensor dimensions for the 48 kHz path compared to the 24 kHz V3 pipeline.

Is V4 backward compatible with V3 models?

Partially. While the TTS class supports version selection via the version parameter, V3-trained checkpoints are optimized for 24 kHz mel-spectrograms. For best results with V4's 48 kHz native output, use models fine-tuned on the V4 configuration, though the codebase maintains compatibility layers for mixed-version workflows.

Why does V3 have metallic artifacts while V4 doesn't?

V3's artifacts originate from non-integer upsampling factors in the audio generation pipeline. When the 24 kHz native output is resampled to higher rates for playback, interpolation errors create metallic, muffled timbres. V4 avoids this by synthesizing at 48 kHz directly, bypassing the destructive upsampling step entirely as documented in README.md (lines 337-341).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →