GPT-SoVITS Memory Requirements by Version: Training vs Inference VRAM Guide

GPT-SoVITS V1 and V2 require 4–10 GB VRAM for training, while V3 and V4 need 14 GB unless gradient checkpointing or LoRA reduces it to 8–12 GB, with inference memory automatically capped by the UI based on model version.

Understanding GPT-SoVITS memory requirements is essential for selecting appropriate hardware and avoiding out-of-memory errors during voice cloning tasks. The RVC-Boss/GPT-SoVITS repository implements distinct memory profiles across its model versions, with automatic batch-size calculations in webui.py that adapt to your GPU's available VRAM.

Training Memory Requirements by Model Version

V1 and V2 (4–8 GB)

The original V1 architecture and the updated V2 variant represent the most memory-efficient training options. According to the repository documentation, V1 typically requires 4 GB to 6 GB VRAM during fine-tuning, while V2 uses "a bit more VRAM than V1" but remains ≤ 8 GB as noted in README.md【L336-L337】. These versions fit comfortably on most consumer GPUs like the RTX 3060 or GTX 1080 Ti.

V2Pro and V2Pro Plus (≤10 GB)

The V2Pro and V2Pro Plus variants introduce enhanced capabilities with modest memory overhead. Release notes in docs/en/README.md indicate these versions require "slightly higher VRAM usage than V2" but stay within the ≤ 10 GB range【L56-L57】. This makes them suitable for GPUs like the RTX 3070 or RTX 2080 Ti without requiring memory optimization techniques.

V3 and V4 (14 GB Baseline, Optimizable to 8 GB)

V3 and V4 represent a significant architectural leap with substantially higher base memory requirements. The changelog in docs/en/Changelog_EN.md explicitly documents that V3 requires 14 GB VRAM for ordinary fine-tuning【L411】. However, the repository provides two critical memory reduction techniques:

  • Gradient checkpointing: Reduces peak training memory to 12 GB VRAM【L427】
  • LoRA fine-tuning: Further drops requirements to approximately 8 GB VRAM【L35】

V4 maintains similar requirements to V3, with the README noting it has "slightly higher VRAM usage than V2" while surpassing V4's performance【L356】, placing it in the same 14 GB-class range as V3.

Inference Memory and Batch Size Logic

How the UI Calculates Safe Batch Sizes

The Gradio interface in webui.py implements automatic memory management through dynamic batch size calculation. The system queries available GPU memory and applies version-specific divisors:


# Logic from webui.py【L19-L22】

if version in {"v3", "v4"}:
    default_batch_size = int(minmem // 8)
else:
    default_batch_size = int(minmem // 2)

For V1 through V2Pro, the UI reserves half of the free GPU memory (minmem // 2)【L19-L21】. For V3 and V4, it becomes significantly more conservative, dividing available memory by eight (minmem // 8)【L19-L20】. This explains why an 8 GB GPU defaults to batch size 1 for V3, while a 24 GB GPU achieves batch size 3.

Half-Precision (FP16) Optimization

The configuration system in config.py automatically detects GPU capabilities to enable half-precision inference. The is_half flag determines whether the model runs in FP16 mode, which roughly halves activation memory:


# From config.py【L58-L66】

tmp = []
for i in range(torch.cuda.device_count()):
    tmp.append(torch.cuda.get_device_properties(i))
    tmp[i].total_memory / 1024**3  # Convert to GB

is_half = any(dtype == torch.float16 for _, dtype, _, _ in tmp)

Setting is_half=True (typically via environment variable or UI toggle) reduces inference VRAM from approximately 3–4 GB to 1.5–2 GB for most model versions.

Memory Optimization Techniques

Gradient Checkpointing

For V3 and V4 training, enabling gradient checkpointing trades computation time for memory savings. This technique recalculates intermediate activations during the backward pass rather than storing them, reducing the V3 training requirement from 14 GB to 12 GB as documented in docs/en/Changelog_EN.md【L427】. Activate this through the training UI checkbox labeled "Gradient checkpointing."

LoRA Fine-Tuning

Low-Rank Adaptation (LoRA) represents the most aggressive memory reduction strategy for newer models. By training only low-rank decomposition matrices instead of full model parameters, LoRA reduces V3 training requirements to approximately 8 GB VRAM【L35】. This makes V3 accessible on GPUs like the RTX 3070 or RTX 4060 Ti. Enable via the "LoRA training" mode in the UI or by setting lora=True in the JSON configuration.

Manual Batch Size Adjustment

While the UI auto-calculates safe defaults in webui.py【L34-L36】, advanced users can override these values through the batch size slider. Reducing batch size directly linearly reduces memory usage during inference, though it may increase processing time for large datasets.

Key Configuration Files

Understanding these source files helps diagnose memory issues:

  • webui.py: Contains the batch size calculation logic (default_batch_size = int(minmem // 2) or // 8)【L19-L22】 and UI sliders for manual adjustment【L1709-L1714】
  • config.py: Implements device detection and is_half flag for FP16 optimization【L58-L66】
  • docs/en/Changelog_EN.md: Documents specific VRAM requirements for V3 (14 GB), checkpointing (12 GB), and LoRA (8 GB)【L35】【L411-L428】
  • README.md: Provides comparative VRAM notes between V1, V2, and V4【L336-L337】【L356】

Summary

  • V1 and V2 train efficiently on 4–8 GB VRAM and infer on 2–4 GB, making them ideal for consumer GPUs.
  • V2Pro variants require slightly more memory (≤10 GB) but remain accessible on mid-range hardware.
  • V3 and V4 demand 14 GB for standard training, but support gradient checkpointing (12 GB) and LoRA (8 GB) to reduce requirements.
  • The UI automatically manages inference memory by calculating batch size as minmem // 2 for older models and minmem // 8 for V3/V4.
  • Enabling half-precision (is_half) roughly halves inference memory usage across all versions.

Frequently Asked Questions

Can I train GPT-SoVITS V3 on an 8 GB GPU?

Yes, but only if you use LoRA fine-tuning. According to the changelog, standard V3 training requires 14 GB VRAM, but LoRA mode reduces this to approximately 8 GB【L35】. You cannot train the full model on an 8 GB card without LoRA.

Why does V3 use more memory than V2?

V3 introduces a larger architecture and more parameters than V2, resulting in higher activation memory during both forward and backward passes. The repository documentation notes that V3 specifically requires 14 GB compared to V2's ≤8 GB【L411】【L336-L337】. This architectural expansion improves voice quality but demands more VRAM.

How do I enable half-precision to save VRAM?

Half-precision (FP16) activates automatically when config.py detects compatible hardware, setting is_half = True【L58-L66】. You can verify or force this by checking the is_half variable in your configuration. When enabled, inference memory drops by roughly 50%, allowing V3 inference on GPUs with as little as 4–5 GB VRAM.

Where is the batch size calculation logic located?

The automatic batch size calculation resides in webui.py around lines 19–22. The code checks your model version and computes default_batch_size = int(minmem // 2) for V1/V2 or int(minmem // 8) for V3/V4【L19-L22】. You can override this calculation manually using the batch size slider in the Gradio interface.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →