# How the `memory_budget_gib` Clamp Works with Per-Process Memory Fraction and Safety Reserve in YuE

> Learn how YuE's memory_budget_gib clamp utilizes per-process memory fraction and safety reserve to prevent exceeding GPU memory limits and avoid negative values.

- Repository: [multimodal-art-projection/YuE](https://github.com/multimodal-art-projection/YuE)
- Tags: internals
- Published: 2026-09-14

---

**The `memory_budget_gib` clamp in YuE sets an absolute GPU memory ceiling by subtracting a 0.5 GiB safety reserve from the requested budget and converting the remainder into a per-process memory fraction, ensuring the allocator never exceeds the limit or returns a negative value.**

YuE, the open-source multimodal music generation model by Multimodal Art Projection, implements sophisticated GPU memory management to prevent out-of-memory crashes during inference. The `memory_budget_gib` parameter acts as a hard limit that interacts with PyTorch's per-process memory fraction and a built-in safety reserve to guarantee stable allocation behavior across different hardware configurations.

## Understanding the Three Memory Management Components

YuE's memory system relies on three interconnected mechanisms to guard GPU resources. Understanding how each component functions is essential for tuning performance on constrained hardware.

### The `memory_budget_gib` Clamp

The **clamp** serves as the top-level guardrail that prevents the model from exhausting GPU memory. It represents the absolute maximum memory (in gibibytes) that the YuE process is permitted to allocate. According to the source code in [`src/yue2/storage.py`](https://github.com/multimodal-art-projection/YuE/blob/main/src/yue2/storage.py), this value is user-configurable via command-line arguments or programmatic API calls, and it acts as the input to the downstream fraction calculation.

### Per-Process Memory Fraction

YuE translates the absolute `memory_budget_gib` value into a **relative fraction** of total device memory using PyTorch's `torch.cuda.set_per_process_memory_fraction` API. This conversion happens in the `set_memory_fraction_from_budget` function, where the usable budget is divided by the device's total memory capacity. By setting this fraction, YuE instructs the CUDA allocator to trigger an out-of-memory error if internal allocations would exceed the computed threshold.

### The Safety Reserve Buffer

Before calculating the fraction, YuE deducts a **safety reserve** of 0.5 GiB (defined as `SAFETY_RESERVE_GIB = 0.5` in [`src/yue2/storage.py`](https://github.com/multimodal-art-projection/YuE/blob/main/src/yue2/storage.py)). This buffer guarantees that a small amount of GPU memory remains free for kernel launches, stream synchronization, and driver-level bookkeeping, preventing hard crashes that occur when memory is exhausted to the last byte.

## How the Clamp Interacts with Memory Allocation

The clamp mechanism ensures that the calculated per-process fraction never becomes negative, even when users specify impossibly small budgets. The logic follows this exact sequence:

1. Retrieve total GPU memory using `torch.cuda.get_device_properties(device).total_memory`
2. Compute usable budget with `max(0.0, memory_budget_gib - SAFETY_RESERVE_GIB)`
3. Derive fraction by dividing usable budget by total memory
4. Apply the fraction via `torch.cuda.set_per_process_memory_fraction`

The `max(0.0, ...)` clamp is critical because it prevents passing a negative value to PyTorch if the requested budget is smaller than the 0.5 GiB safety reserve. Without this guard, the program would attempt to set a negative memory fraction, leading to undefined behavior or opaque runtime errors.

## Implementation Details in [`src/yue2/storage.py`](https://github.com/multimodal-art-projection/YuE/blob/main/src/yue2/storage.py)

The core logic resides in [`src/yue2/storage.py`](https://github.com/multimodal-art-projection/YuE/blob/main/src/yue2/storage.py), where the `set_memory_fraction_from_budget` function orchestrates the interaction between the clamp, safety reserve, and PyTorch's allocator:

```python

# src/yue2/storage.py

import torch

SAFETY_RESERVE_GIB = 0.5

def set_memory_fraction_from_budget(memory_budget_gib: float, device):
    """
    Convert an absolute GiB budget into a per-process memory fraction.
    """
    # Convert bytes to GiB

    total_mem = torch.cuda.get_device_properties(device).total_memory / (1024 ** 3)
    
    # Apply clamp to ensure non-negative usable budget

    usable_budget = max(0.0, memory_budget_gib - SAFETY_RESERVE_GIB)
    
    # Calculate fraction of total device memory

    per_process_fraction = usable_budget / total_mem
    
    # Set PyTorch's internal limit

    torch.cuda.set_per_process_memory_fraction(per_process_fraction, device)

```

During pipeline initialization, [`src/yue2/pipeline.py`](https://github.com/multimodal-art-projection/YuE/blob/main/src/yue2/pipeline.py) invokes this function to establish memory boundaries before any model weights are loaded. This initialization order ensures that subsequent tensor allocations—whether for audio generation buffers or attention caches—respect the enforced limit immediately.

## Practical Usage Examples

You can control the `memory_budget_gib` clamp through command-line interfaces or direct Python imports.

### Command-Line Configuration

When running generation scripts, pass the budget flag to constrain GPU usage:

```bash
python examples/generate.py --memory_budget_gib 10.0

```

This command limits the process to 10 GiB of GPU memory (effectively 9.5 GiB after the safety reserve).

### Programmatic Configuration

For custom scripts, import the configuration function directly:

```python
from yue2.storage import set_memory_fraction_from_budget
import torch

device = torch.device("cuda:0")

# Set a 12 GiB budget (11.5 GiB usable)

set_memory_fraction_from_budget(memory_budget_gib=12.0, device=device)

# Subsequent allocations respect the limit

latents = torch.randn(1, 100, 1024, device=device, dtype=torch.float16)

```

### Verifying the Calculated Fraction

To inspect the effective memory fraction that YuE calculates:

```python
import torch

device = torch.device("cuda:0")
total_gib = torch.cuda.get_device_properties(device).total_memory / (1024 ** 3)
requested_budget = 12.0
reserve = 0.5

fraction = max(0.0, (requested_budget - reserve) / total_gib)
print(f"Per-process memory fraction: {fraction:.4f}")

```

## Why the Safety Reserve Matters

The **0.5 GiB safety reserve** is not merely a conservative estimate—it is a defensive programming measure against GPU driver behavior. CUDA kernels require scratch space for stream operations, and the PyTorch caching allocator maintains metadata that consumes additional memory beyond tensor payloads. By hardcoding `SAFETY_RESERVE_GIB` and deducting it before the clamp calculation, YuE guarantees that even when users maximize their `memory_budget_gib` setting, the system retains headroom for these essential operations. This design prevents the "death spiral" of out-of-memory errors that occur when allocation requests fail due to driver overhead rather than tensor size.

## Summary

- The **`memory_budget_gib` clamp** sets an absolute ceiling on GPU memory consumption by converting GiB values into per-process fractions.
- **Calculation logic** in [`src/yue2/storage.py`](https://github.com/multimodal-art-projection/YuE/blob/main/src/yue2/storage.py) subtracts the 0.5 GiB `SAFETY_RESERVE_GIB` before converting to a fraction, using `max(0.0, ...)` to prevent negative values.
- **PyTorch integration** occurs through `torch.cuda.set_per_process_memory_fraction`, which enforces the limit at the allocator level.
- **Pipeline initialization** in [`src/yue2/pipeline.py`](https://github.com/multimodal-art-projection/YuE/blob/main/src/yue2/pipeline.py) applies these settings before model loading to ensure immediate enforcement.
- The **safety reserve** guarantees kernel and driver stability by preserving a small memory buffer regardless of user configuration.

## Frequently Asked Questions

### What happens if I set `memory_budget_gib` smaller than the safety reserve?

If you request a budget below 0.5 GiB, the clamp forces `usable_budget` to 0.0 via the `max(0.0, ...)` operation. This results in a per-process memory fraction of zero, meaning **PyTorch will immediately raise an out-of-memory error on the first allocation attempt**. This "fail fast" behavior prevents the process from entering an unstable state with insufficient memory for basic operations.

### How does YuE calculate the per-process memory fraction from the budget?

YuE retrieves the device's total memory capacity, subtracts the 0.5 GiB `SAFETY_RESERVE_GIB`, and divides the remaining value by the total memory. The result is passed to `torch.cuda.set_per_process_memory_fraction` in [`src/yue2/storage.py`](https://github.com/multimodal-art-projection/YuE/blob/main/src/yue2/storage.py). For example, on a 24 GiB GPU with a 12 GiB budget, the effective fraction is approximately 0.48 (11.5 GiB usable divided by 24 GiB total).

### Can I modify or disable the safety reserve constant?

The **0.5 GiB reserve** is hardcoded as `SAFETY_RESERVE_GIB = 0.5` at the module level in [`src/yue2/storage.py`](https://github.com/multimodal-art-projection/YuE/blob/main/src/yue2/storage.py). While you can edit the source file to change this value, it is not exposed as a runtime configuration parameter. Modifying it requires changing the constant and reinstalling the package, as the design treats this buffer as a non-negotiable system requirement.

### Where is the `memory_budget_gib` parameter initialized in the pipeline?

The parameter is first ingested in [`src/yue2/pipeline.py`](https://github.com/multimodal-art-projection/YuE/blob/main/src/yue2/pipeline.py) during pipeline instantiation, where user-supplied arguments (such as `--memory_budget_gib` from the CLI) are passed to `set_memory_fraction_from_budget` in [`src/yue2/storage.py`](https://github.com/multimodal-art-projection/YuE/blob/main/src/yue2/storage.py). This initialization occurs **before** any model weights are downloaded or loaded into GPU memory, ensuring the limit is active from the first allocation.