How the GPT‑SoVITS WebUI Handles Multiple Concurrent Inference Requests

The GPT‑SoVITS WebUI serializes concurrent inference requests by maintaining a single global subprocess handle that terminates any running process before spawning a new one, ensuring only one inference task executes at any given moment.

The open‑source speech synthesis toolkit RVC‑Boss/GPT‑SoVITS provides a Gradio‑based WebUI for text‑to‑speech generation. While the interface appears asynchronous, the underlying architecture strictly manages multiple concurrent inference requests through process‑level serialization to prevent GPU contention and memory exhaustion.

Global Process Serialization in webui.py

The core mechanism lives in webui.py, where a global variable p_tts_inference tracks the currently running inference subprocess. When a user initiates synthesis by clicking the “推理” (Inference) button, the system invokes change_tts_inference() to evaluate whether an existing job is already active.


# webui.py

p_tts_inference = None  # global handle for the TTS subprocess

def change_tts_inference(...):
    global p_tts_inference
    # ... environment variable setup ...

    
    if p_tts_inference is None:
        # ✅ No inference running → start fresh

        p_tts_inference = Popen(cmd, shell=True)
    else:
        # ❌ Inference active → kill it before restart

        kill_process(p_tts_inference.pid, process_name_tts)
        p_tts_inference = None
        # UI reverts to "closed" state

This check guarantees that only one inference process (inference_webui_fast.py or inference_webui.py) exists at any time. If a new request arrives while p_tts_inference is not None, the UI immediately terminates the old process and clears the handle before launching the replacement.

The Inference Lifecycle: Starting and Replacing Processes

The change_tts_inference() function dynamically constructs the command line based on whether fast‑batch inference is enabled. It selects between the standard pipeline and the optimized fast pipeline, then applies the singleton‑process rule.

def change_tts_inference(...):
    global p_tts_inference
    
    # Build command based on fast-inference flag

    cmd = (
        f'"{python_exec}" -s GPT_SoVITS/inference_webui_fast.py "{language}"'
        if batched_infer_enabled else
        f'"{python_exec}" -s GPT_SoVITS/inference_webui.py "{language}"'
    )
    
    if p_tts_inference is None:
        p_tts_inference = Popen(cmd, shell=True)  # Launch new subprocess

    else:
        kill_process(p_tts_inference.pid, process_name_tts)  # Stop old one

        p_tts_inference = None

By killing the existing subprocess before starting a new one, the WebUI effectively queues requests sequentially rather than parallelizing them, protecting the host from uncontrolled resource consumption.

Terminating Running Subprocesses

Users can manually abort inference through the “关闭” (Close) button, which triggers close_tts(). This function mirrors the replacement logic by invoking the same kill_process() helper to ensure clean teardown.

def close_tts():
    global p_tts_inference
    if p_tts_inference is not None:
        kill_process(p_tts_inference.pid, process_name_tts)  # Terminate

        p_tts_inference = None
    return (
        process_info(process_name_tts, "closed"),
        {"__type__": "update", "visible": True},
        {"__type__": "update", "visible": False},
    )

The actual process termination is handled by kill_process() in the utility module, which adapts to the host operating system:

def kill_process(pid, process_name=""):
    if platform.system() == "Windows":
        subprocess.run(
            f"taskkill /t /f /pid {pid}", 
            shell=True,
            stdout=subprocess.DEVNULL, 
            stderr=subprocess.DEVNULL
        )
    else:
        kill_proc_tree(pid)  # Recursively kills children via psutil

    print(process_name + i18n("进程已终止"))

On Windows, the function executes taskkill /t /f to forcibly destroy the process tree. On Unix‑like systems, it uses psutil to recursively kill child processes, ensuring no orphaned GPU processes remain.

Consistent Pattern Across Heavy‑Weight Operations

The serialization strategy is not limited to TTS inference. The WebUI applies identical singleton‑process guards to other resource‑intensive tasks:

  • p_uvr5 – UVR5 audio separation
  • p_label – Dataset labeling and preparation
  • ps_slice – Audio slicing preprocessing

Each variable follows the same lifecycle: check for None, spawn subprocess on first use, kill and reset on subsequent requests. This unified approach prevents users from accidentally launching multiple model pipelines that would compete for VRAM and CPU threads.

Summary

  • Single subprocess policy: The WebUI stores the current inference process ID in p_tts_inference and rejects parallel starts by terminating existing jobs first.
  • Automatic cleanup: The kill_process() helper uses platform‑specific commands (taskkill on Windows, psutil on Linux/macOS) to ensure complete process tree destruction.
  • UI state synchronization: Functions like change_tts_inference() and close_tts() toggle button visibility and status messages to reflect whether a job is active or idle.
  • Universal pattern: Heavy‑weight operations (UVR5, slicing, labeling) reuse the same serialization logic via dedicated global handles (p_uvr5, ps_slice, etc.).

Frequently Asked Questions

Does GPT‑SoVITS support parallel inference requests in the WebUI?

No. According to the source code in webui.py, the WebUI explicitly prevents parallel inference by killing any running subprocess before starting a new one. While the underlying TTS_infer_pack/TTS.py engine exposes a parallel_infer flag, the WebUI layer enforces a single active process to avoid GPU contention.

What happens if I click the inference button while a job is running?

The WebUI immediately terminates the existing inference process via kill_process(), clears the p_tts_inference handle, and spawns a fresh subprocess for the new request. The previous job is aborted, and the UI updates to show the new task status.

How does the WebUI handle process cleanup on different operating systems?

The kill_process() function in the utility module detects the platform at runtime. On Windows, it executes taskkill /t /f /pid to forcefully kill the process tree. On Unix systems, it uses psutil to recursively terminate child processes, ensuring clean GPU memory release regardless of the host OS.

Can I modify the code to allow true concurrent inference?

Yes, but it requires significant refactoring. You would need to remove the global p_tts_inference checks in webui.py, implement request queuing or worker pools, and ensure your GPU has sufficient VRAM to load multiple model instances simultaneously. The current serialization is intentionally defensive to protect consumer‑grade hardware from out‑of‑memory crashes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →