How UVR5 Vocal Separation Integrates with the GPT-SoVITS Preprocessing Workflow

GPT-SoVITS isolates clean vocal tracks from mixed audio by launching UVR5 as a separate subprocess, writing separated files to output/uvr5_opt before the audio slicer and ASR stages run.

The RVC-Boss/GPT-SoVITS repository embeds a dedicated UVR5 (Ultimate Vocal Remover) module directly into its data preparation pipeline. This integration ensures that raw audio mixtures are split into clean vocals and instrumental accompaniment before semantic token extraction and acoustic model training begin.

Subprocess Architecture for GPU Isolation

The main Gradio interface delegates heavy inference to an isolated Python process to prevent UI freezes. In webui.py (lines 3,001–3,025), the change_uvr5() function constructs a shell command that spawns the UVR5 web interface as a standalone subprocess:

cmd = f'"{python_exec}" -s tools/uvr5/webui.py "{infer_device}" {is_half} {webui_port_uvr5} {is_share}'
p_uvr5 = Popen(cmd, shell=True)

This approach allows the separation model to claim dedicated GPU memory without blocking the primary training UI. When the user clicks Close, change_uvr5() invokes kill_process(p_uvr5.pid, process_name_uvr5) to walk the child-process tree on Unix or issue taskkill on Windows, ensuring the heavy inference worker terminates cleanly.

Web UI Integration Points

The primary interface exposes UVR5 through a dedicated accordion “UVR5 人声分离&去混响去延迟工具” defined around lines 1,580–1,630 in webui.py. The Open and Close buttons trigger change_uvr5(), which returns status tuples via process_info() (lines 44–64) to signal whether the separation worker is opened, running, or finished.

Default configuration values reside in config.py, where webui_port_uvr5 = 9873 sets the dedicated service port. The Gradio widgets for model selection, output directories, and half-precision flags are rendered in tools/uvr5/webui.py (lines 173–192) and map directly to command-line arguments passed downstream.

The UVR5 Backend Pipeline

The subprocess entry point tools/uvr5/webui.py loads available checkpoints from tools/uvr5/uvr5_weights and exposes a single API endpoint uvr_convert (line 215). The core inference logic lives in tools/uvr5/vr.py inside the uvr() function, which orchestrates model loading via helper utilities in tools/uvr5/lib/utils.py and executes separation using architectures defined in tools/uvr5/bsroformer.py and tools/uvr5/mdxnet.py.

Data flow follows this strict pattern:

  • Input: A directory of mixed audio files (vocals + instruments) selected through the Gradio file picker.
  • Processing: The selected model runs inference on each file, optionally applying dereverberation.
  • Output: Two folders are created under output/uvr5_opt:
    • opt_vocal_root for the isolated vocal stem.
    • opt_ins_root for the accompaniment instrumental.

Placement in the Preprocessing Workflow

The “前置数据集获取工具” tab (lines 1,514–1,538) in the main UI establishes a rigid execution order:


UVR5 → Slicer → ASR → Text-audio alignment → Training

By performing UVR5 vocal separation first, the subsequent slicer (tools/slice_audio.py) processes clean vocal tracks rather than mixed signals. This reduces memory spikes during segmentation and improves the accuracy of automatic speech recognition (ASR) because the model receives speech-only audio without instrumental interference. The resulting semantic tokens and acoustic features therefore reflect pure vocal characteristics, leading to higher-quality voice cloning.

Programmatic Usage Outside the UI

You can trigger the separation stage manually without launching the full Gradio interface. The command structure mirrors the internal subprocess call:

import sys
from subprocess import Popen

cmd = (
    f'"{sys.executable}" -s tools/uvr5/webui.py '
    f'"cuda:0" False 9873 False'   # device, half-precision, port, share-mode

)
p = Popen(cmd, shell=True)
p.wait()   # Block until separation completes

After execution, retrieve the vocal files from output/uvr5_opt and pass them directly to tools/slice_audio.py or any custom preprocessing script.

Configuration and Model Management

All tunable parameters are exposed through the tools/uvr5/webui.py interface:

  • Model selection: Choose between BS-RoFormer, MDX-Net, or other architectures stored in tools/uvr5/uvr5_weights.
  • Hardware settings: Toggle is_half for FP16 inference and specify infer_device (CPU or CUDA device).
  • I/O paths: Override opt_vocal_root and opt_ins_root to redirect output away from the default output/uvr5_opt directory.

Changing any widget updates the argument vector passed to vr.py, allowing real-time reconfiguration without restarting the main training UI.

Summary

  • UVR5 runs as a subprocess launched from webui.py via change_uvr5() to isolate GPU memory from the main interface.
  • Default port 9873 is defined in config.py and used by tools/uvr5/webui.py to serve the separation API.
  • Core logic resides in tools/uvr5/vr.py (function uvr), with model definitions in bsroformer.py and mdxnet.py.
  • Output structure generates separate vocal and accompaniment folders under output/uvr5_opt.
  • Pipeline order places vocal separation before slicing and ASR to ensure clean training data.

Frequently Asked Questions

Where does UVR5 output the separated vocal files?

By default, the system writes isolated vocals to output/uvr5_opt (controlled by the opt_vocal_root parameter) and instrumental accompaniment to the same parent directory under opt_ins_root. You can override these paths in the Gradio widgets defined in tools/uvr5/webui.py (lines 173–192).

Why does GPT-SoVITS launch UVR5 as a separate process instead of a function call?

The change_uvr5() function in webui.py (lines 3,001–3,025) uses Popen to spawn tools/uvr5/webui.py in its own shell. This prevents the heavy neural network inference from freezing the main Gradio event loop and allows independent GPU allocation. The parent process can then kill the worker cleanly via kill_process() when separation completes.

Can I use UVR5 vocal separation without the full GPT-SoVITS interface?

Yes. Execute the same command the UI uses: python -s tools/uvr5/webui.py "{device}" {half_precision} {port} {share_mode}. This starts the standalone UVR5 service on the specified port (default 9873). Processed files appear in output/uvr5_opt and can be fed manually into tools/slice_audio.py or your own preprocessing scripts.

Which model architectures does the integrated UVR5 support?

The backend loads checkpoints from tools/uvr5/uvr5_weights and supports BS-RoFormer (defined in tools/uvr5/bsroformer.py) and MDX-Net (defined in tools/uvr5/mdxnet.py), among others. Model-specific loading utilities reside in tools/uvr5/lib/utils.py, and the active architecture is selected through the Gradio dropdown in tools/uvr5/webui.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →