# How the GPT‑SoVITS WebUI Handles Multiple Concurrent Inference Requests

> Discover how the GPT-SoVITS WebUI manages concurrent inference requests by serializing tasks, ensuring efficient processing and preventing conflicts.

- Repository: [RVC-Boss/GPT-SoVITS](https://github.com/RVC-Boss/GPT-SoVITS)
- Tags: internals
- Published: 2026-03-07

---

**The GPT‑SoVITS WebUI serializes concurrent inference requests by maintaining a single global subprocess handle that terminates any running process before spawning a new one, ensuring only one inference task executes at any given moment.**

The open‑source speech synthesis toolkit **RVC‑Boss/GPT‑SoVITS** provides a Gradio‑based WebUI for text‑to‑speech generation. While the interface appears asynchronous, the underlying architecture strictly manages **multiple concurrent inference requests** through process‑level serialization to prevent GPU contention and memory exhaustion.

## Global Process Serialization in webui.py

The core mechanism lives in [`webui.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/webui.py), where a global variable `p_tts_inference` tracks the currently running inference subprocess. When a user initiates synthesis by clicking the *“推理”* (Inference) button, the system invokes `change_tts_inference()` to evaluate whether an existing job is already active.

```python

# webui.py

p_tts_inference = None  # global handle for the TTS subprocess

def change_tts_inference(...):
    global p_tts_inference
    # ... environment variable setup ...

    
    if p_tts_inference is None:
        # ✅ No inference running → start fresh

        p_tts_inference = Popen(cmd, shell=True)
    else:
        # ❌ Inference active → kill it before restart

        kill_process(p_tts_inference.pid, process_name_tts)
        p_tts_inference = None
        # UI reverts to "closed" state

```

This check guarantees that **only one inference process** ([`inference_webui_fast.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/inference_webui_fast.py) or [`inference_webui.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/inference_webui.py)) exists at any time. If a new request arrives while `p_tts_inference` is not `None`, the UI immediately terminates the old process and clears the handle before launching the replacement.

## The Inference Lifecycle: Starting and Replacing Processes

The `change_tts_inference()` function dynamically constructs the command line based on whether fast‑batch inference is enabled. It selects between the standard pipeline and the optimized fast pipeline, then applies the singleton‑process rule.

```python
def change_tts_inference(...):
    global p_tts_inference
    
    # Build command based on fast-inference flag

    cmd = (
        f'"{python_exec}" -s GPT_SoVITS/inference_webui_fast.py "{language}"'
        if batched_infer_enabled else
        f'"{python_exec}" -s GPT_SoVITS/inference_webui.py "{language}"'
    )
    
    if p_tts_inference is None:
        p_tts_inference = Popen(cmd, shell=True)  # Launch new subprocess

    else:
        kill_process(p_tts_inference.pid, process_name_tts)  # Stop old one

        p_tts_inference = None

```

By killing the existing subprocess before starting a new one, the WebUI effectively **queues requests sequentially** rather than parallelizing them, protecting the host from uncontrolled resource consumption.

## Terminating Running Subprocesses

Users can manually abort inference through the *“关闭”* (Close) button, which triggers `close_tts()`. This function mirrors the replacement logic by invoking the same `kill_process()` helper to ensure clean teardown.

```python
def close_tts():
    global p_tts_inference
    if p_tts_inference is not None:
        kill_process(p_tts_inference.pid, process_name_tts)  # Terminate

        p_tts_inference = None
    return (
        process_info(process_name_tts, "closed"),
        {"__type__": "update", "visible": True},
        {"__type__": "update", "visible": False},
    )

```

The actual process termination is handled by `kill_process()` in the utility module, which adapts to the host operating system:

```python
def kill_process(pid, process_name=""):
    if platform.system() == "Windows":
        subprocess.run(
            f"taskkill /t /f /pid {pid}", 
            shell=True,
            stdout=subprocess.DEVNULL, 
            stderr=subprocess.DEVNULL
        )
    else:
        kill_proc_tree(pid)  # Recursively kills children via psutil

    print(process_name + i18n("进程已终止"))

```

On Windows, the function executes `taskkill /t /f` to forcibly destroy the process tree. On Unix‑like systems, it uses `psutil` to recursively kill child processes, ensuring no orphaned GPU processes remain.

## Consistent Pattern Across Heavy‑Weight Operations

The serialization strategy is not limited to TTS inference. The WebUI applies identical singleton‑process guards to other resource‑intensive tasks:

- **`p_uvr5`** – UVR5 audio separation
- **`p_label`** – Dataset labeling and preparation  
- **`ps_slice`** – Audio slicing preprocessing

Each variable follows the same lifecycle: check for `None`, spawn subprocess on first use, kill and reset on subsequent requests. This unified approach prevents users from accidentally launching multiple model pipelines that would compete for VRAM and CPU threads.

## Summary

- **Single subprocess policy**: The WebUI stores the current inference process ID in `p_tts_inference` and rejects parallel starts by terminating existing jobs first.
- **Automatic cleanup**: The `kill_process()` helper uses platform‑specific commands (`taskkill` on Windows, `psutil` on Linux/macOS) to ensure complete process tree destruction.
- **UI state synchronization**: Functions like `change_tts_inference()` and `close_tts()` toggle button visibility and status messages to reflect whether a job is active or idle.
- **Universal pattern**: Heavy‑weight operations (UVR5, slicing, labeling) reuse the same serialization logic via dedicated global handles (`p_uvr5`, `ps_slice`, etc.).

## Frequently Asked Questions

### Does GPT‑SoVITS support parallel inference requests in the WebUI?

No. According to the source code in [`webui.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/webui.py), the WebUI explicitly prevents parallel inference by killing any running subprocess before starting a new one. While the underlying [`TTS_infer_pack/TTS.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/TTS_infer_pack/TTS.py) engine exposes a `parallel_infer` flag, the WebUI layer enforces a single active process to avoid GPU contention.

### What happens if I click the inference button while a job is running?

The WebUI immediately terminates the existing inference process via `kill_process()`, clears the `p_tts_inference` handle, and spawns a fresh subprocess for the new request. The previous job is aborted, and the UI updates to show the new task status.

### How does the WebUI handle process cleanup on different operating systems?

The `kill_process()` function in the utility module detects the platform at runtime. On Windows, it executes `taskkill /t /f /pid` to forcefully kill the process tree. On Unix systems, it uses `psutil` to recursively terminate child processes, ensuring clean GPU memory release regardless of the host OS.

### Can I modify the code to allow true concurrent inference?

Yes, but it requires significant refactoring. You would need to remove the global `p_tts_inference` checks in [`webui.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/webui.py), implement request queuing or worker pools, and ensure your GPU has sufficient VRAM to load multiple model instances simultaneously. The current serialization is intentionally defensive to protect consumer‑grade hardware from out‑of‑memory crashes.