How GPT-SoVITS Handles Version Compatibility Between v1, v2, v3, v4, v2Pro, and v2ProPlus Models

GPT-SoVITS achieves seamless version compatibility through six coordinated mechanisms including centralized weight mapping in config.py, conditional configuration loading in webui.py, runtime backend selection in inference_webui.py, and strict API validation in TTS_infer_pack/TTS.py.

The RVC-Boss/GPT-SoVITS repository supports six distinct model architectures—v1, v2, v3, v4, v2Pro, and v2ProPlus—each with unique hyperparameters, vocoder requirements, and sampling rates. Understanding how the framework manages GPT-SoVITS version compatibility ensures you select the correct training configurations and inference pipelines for your specific checkpoint files.

Centralized Version Configuration and Weight Mapping

The foundation of compatibility lies in explicit version-to-resource mapping that prevents cross-contamination of model artifacts.

Version Registry in config.py

In config.py, the framework maintains version-specific weight maps that bind each version identifier to its corresponding pretrained SoVITS and GPT checkpoint files. This central registry also defines the directory structures where new checkpoints are saved during training, ensuring that v1, v2, v3, v4, v2Pro, and v2ProPlus artifacts remain organized and accessible.

Conditional Configuration Loading in webui.py

When initiating training via webui.py, the open1Ba function selects JSON configuration files based on the supplied version string. Standard versions (v1, v2, v3, v4) load from s2.json, while specialized Pro variants require s2v2Pro.json or s2v2ProPlus.json. This conditional loading guarantees that hyperparameters such as sample rate and vocoder type match the architectural requirements of the target version.

Runtime Backend Selection and Model Initialization

Inference pipelines adapt dynamically to version-specific synthesis requirements through runtime flag inspection.

Backend Selection in inference_webui.py

At inference time, inference_webui.py inspects the model_version parameter to instantiate the correct synthesis backend. The framework supports three distinct vocoder implementations: BigVGAN for standard versions, Hifigan for specific legacy configurations, and the custom SV-CN model exclusively for Pro families (v2Pro and v2ProPlus).

The v2pro_set Identifier in module/models.py

To streamline Pro-version detection throughout the codebase, module/models.py defines a helper set:

v2pro_set = {"v2Pro", "v2ProPlus"}

This constant is reused across multiple modules to test version membership without hardcoding strings, reducing error risk when implementing version-specific logic branches.

Data Pipeline and Sampling Rate Adaptation

Training data preprocessing adjusts automatically based on version capabilities to prevent spectral mismatches.

Dynamic Processing in module/data_utils.py

The DataCollator class in module/data_utils.py sets a boolean flag during initialization:

self.is_v2Pro = version in {"v2Pro","v2ProPlus"}

Downstream components query this flag to apply 16 kHz-only processing required by Pro models, while standard versions may utilize different spectrogram resampling strategies. This encapsulation prevents data pipeline mismatches that could degrade model performance.

Checkpoint Integrity and API Validation

Robust safeguards prevent accidental mixing of incompatible model artifacts across different versions.

Binary ID Mapping in process_ckpt.py

Legacy checkpoint compatibility is maintained through binary tags in process_ckpt.py. The system maps specific byte sequences—b"05" for v2Pro and b"06" for v2ProPlus—to version strings within stored .pth files. This metadata allows the framework to identify checkpoint origins even when files are mixed in shared logs/ directories.

Strict Version Validation in TTS_infer_pack/TTS.py

The public inference API exposed in TTS_infer_pack/TTS.py strictly validates the version parameter against the allowed list:

["v1","v2","v3","v4","v2Pro","v2ProPlus"]

The constructor raises assertions if an unsupported version is supplied, preventing runtime errors from mismatched SoVITS and GPT checkpoint pairs.

Practical Implementation Examples

Configuring Training for Pro Versions

To initiate training with v2Pro or v2ProPlus architectures, specify the version parameter in the UI backend call:


# webui.py entry point for training (open1Ba function)

open1Ba(
    version="v2Pro",          # Selects s2v2Pro.json automatically

    batch_size=4,
    total_epoch=2000,
    exp_name="my_pro_exp",
    text_low_lr_rate=0.1,
    if_save_latest=True,
    if_save_every_weights=False,
    save_every_epoch=10,
    gpu_numbers1Ba="0-1",
    pretrained_s2G="",
    pretrained_s2D="",
    if_grad_ckpt=False,
    lora_rank=0,
)

The function automatically resolves the correct configuration path based on the version string.

Initializing Inference with Version Specification

When loading checkpoints for synthesis, explicitly declare the version to ensure proper backend initialization:

from GPT_SoVITS.TTS_infer_pack.TTS import TTS

tts = TTS(
    sovits_path="GPT_SoVITS/pretrained_models/v2Pro/s2Gv2Pro.pth",
    gpt_path="GPT_SoVITS/pretrained_models/s1v3.ckpt",
    version="v2Pro",          # Mandatory for correct pipeline selection

    language="zh",
)

audio = tts.infer(
    text="你好,世界!",
    speed=1.0,
    top_k=5,
    top_p=0.7,
)

The TTS class validates the version parameter and wires the appropriate model paths and vocoder settings.

Inspecting Data Pipeline Flags

Verify Pro-version detection within the data utilities:

from GPT_SoVITS.module.data_utils import DataCollator

collator = DataCollator(version="v2ProPlus")
print(collator.is_v2Pro)   # Output: True

This flag drives conditional logic for spectrogram generation and audio resampling throughout the training loop.

Summary

  • Centralized mapping in config.py maintains version-specific weight directories and pretrained model paths for all six supported architectures.
  • Conditional loading in webui.py selects appropriate JSON configurations (s2.json vs. s2v2Pro.json) based on the version parameter passed to training functions.
  • Runtime backend selection in inference_webui.py instantiates BigVGAN, Hifigan, or SV-CN vocoders according to version requirements.
  • Pro-version detection relies on the v2pro_set constant in module/models.py and the is_v2Pro flag in module/data_utils.py for specialized 16 kHz processing.
  • Checkpoint integrity is preserved through binary ID tags (b"05", b"06") in process_ckpt.py and strict API validation in TTS_infer_pack/TTS.py.

Frequently Asked Questions

What distinguishes v2Pro and v2ProPlus from standard versions in GPT-SoVITS?

v2Pro and v2ProPlus utilize a custom SV-CN synthesis backend and require 16 kHz sampling rates, whereas standard versions (v1, v2, v3, v4) typically employ BigVGAN or Hifigan vocoders with different spectral processing parameters. The framework detects these versions through the v2pro_set identifier and applies specialized data preprocessing accordingly.

Can I interchange checkpoints between different GPT-SoVITS versions?

No. The TTS class in TTS_infer_pack/TTS.py explicitly validates version strings and prevents mixing incompatible SoVITS and GPT checkpoints. Additionally, binary ID tags in process_ckpt.py ensure that checkpoint files carry embedded version metadata, allowing the system to detect mismatches even if file paths do not indicate version information.

How does the system select the correct configuration file during training?

The open1Ba function in webui.py implements conditional logic that maps version strings to specific JSON configurations. Standard versions load s2.json, while v2Pro loads s2v2Pro.json and v2ProPlus loads s2v2ProPlus.json. This mechanism ensures that training hyperparameters align with the architectural requirements of the target model version.

Where is the version compatibility logic centralized for maintenance?

While version detection logic is distributed across multiple files for performance, the canonical version registry resides in config.py, which defines the mapping between version identifiers and resource paths. The v2pro_set in module/models.py serves as the single source of truth for Pro-version membership testing throughout the codebase.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →