How the `memory_budget_gib` Clamp Works with Per-Process Memory Fraction and Safety Reserve in YuE
The memory_budget_gib clamp in YuE sets an absolute GPU memory ceiling by subtracting a 0.5 GiB safety reserve from the requested budget and converting the remainder into a per-process memory fraction, ensuring the allocator never exceeds the limit or returns a negative value.
YuE, the open-source multimodal music generation model by Multimodal Art Projection, implements sophisticated GPU memory management to prevent out-of-memory crashes during inference. The memory_budget_gib parameter acts as a hard limit that interacts with PyTorch's per-process memory fraction and a built-in safety reserve to guarantee stable allocation behavior across different hardware configurations.
Understanding the Three Memory Management Components
YuE's memory system relies on three interconnected mechanisms to guard GPU resources. Understanding how each component functions is essential for tuning performance on constrained hardware.
The memory_budget_gib Clamp
The clamp serves as the top-level guardrail that prevents the model from exhausting GPU memory. It represents the absolute maximum memory (in gibibytes) that the YuE process is permitted to allocate. According to the source code in src/yue2/storage.py, this value is user-configurable via command-line arguments or programmatic API calls, and it acts as the input to the downstream fraction calculation.
Per-Process Memory Fraction
YuE translates the absolute memory_budget_gib value into a relative fraction of total device memory using PyTorch's torch.cuda.set_per_process_memory_fraction API. This conversion happens in the set_memory_fraction_from_budget function, where the usable budget is divided by the device's total memory capacity. By setting this fraction, YuE instructs the CUDA allocator to trigger an out-of-memory error if internal allocations would exceed the computed threshold.
The Safety Reserve Buffer
Before calculating the fraction, YuE deducts a safety reserve of 0.5 GiB (defined as SAFETY_RESERVE_GIB = 0.5 in src/yue2/storage.py). This buffer guarantees that a small amount of GPU memory remains free for kernel launches, stream synchronization, and driver-level bookkeeping, preventing hard crashes that occur when memory is exhausted to the last byte.
How the Clamp Interacts with Memory Allocation
The clamp mechanism ensures that the calculated per-process fraction never becomes negative, even when users specify impossibly small budgets. The logic follows this exact sequence:
- Retrieve total GPU memory using
torch.cuda.get_device_properties(device).total_memory - Compute usable budget with
max(0.0, memory_budget_gib - SAFETY_RESERVE_GIB) - Derive fraction by dividing usable budget by total memory
- Apply the fraction via
torch.cuda.set_per_process_memory_fraction
The max(0.0, ...) clamp is critical because it prevents passing a negative value to PyTorch if the requested budget is smaller than the 0.5 GiB safety reserve. Without this guard, the program would attempt to set a negative memory fraction, leading to undefined behavior or opaque runtime errors.
Implementation Details in src/yue2/storage.py
The core logic resides in src/yue2/storage.py, where the set_memory_fraction_from_budget function orchestrates the interaction between the clamp, safety reserve, and PyTorch's allocator:
# src/yue2/storage.py
import torch
SAFETY_RESERVE_GIB = 0.5
def set_memory_fraction_from_budget(memory_budget_gib: float, device):
"""
Convert an absolute GiB budget into a per-process memory fraction.
"""
# Convert bytes to GiB
total_mem = torch.cuda.get_device_properties(device).total_memory / (1024 ** 3)
# Apply clamp to ensure non-negative usable budget
usable_budget = max(0.0, memory_budget_gib - SAFETY_RESERVE_GIB)
# Calculate fraction of total device memory
per_process_fraction = usable_budget / total_mem
# Set PyTorch's internal limit
torch.cuda.set_per_process_memory_fraction(per_process_fraction, device)
During pipeline initialization, src/yue2/pipeline.py invokes this function to establish memory boundaries before any model weights are loaded. This initialization order ensures that subsequent tensor allocations—whether for audio generation buffers or attention caches—respect the enforced limit immediately.
Practical Usage Examples
You can control the memory_budget_gib clamp through command-line interfaces or direct Python imports.
Command-Line Configuration
When running generation scripts, pass the budget flag to constrain GPU usage:
python examples/generate.py --memory_budget_gib 10.0
This command limits the process to 10 GiB of GPU memory (effectively 9.5 GiB after the safety reserve).
Programmatic Configuration
For custom scripts, import the configuration function directly:
from yue2.storage import set_memory_fraction_from_budget
import torch
device = torch.device("cuda:0")
# Set a 12 GiB budget (11.5 GiB usable)
set_memory_fraction_from_budget(memory_budget_gib=12.0, device=device)
# Subsequent allocations respect the limit
latents = torch.randn(1, 100, 1024, device=device, dtype=torch.float16)
Verifying the Calculated Fraction
To inspect the effective memory fraction that YuE calculates:
import torch
device = torch.device("cuda:0")
total_gib = torch.cuda.get_device_properties(device).total_memory / (1024 ** 3)
requested_budget = 12.0
reserve = 0.5
fraction = max(0.0, (requested_budget - reserve) / total_gib)
print(f"Per-process memory fraction: {fraction:.4f}")
Why the Safety Reserve Matters
The 0.5 GiB safety reserve is not merely a conservative estimate—it is a defensive programming measure against GPU driver behavior. CUDA kernels require scratch space for stream operations, and the PyTorch caching allocator maintains metadata that consumes additional memory beyond tensor payloads. By hardcoding SAFETY_RESERVE_GIB and deducting it before the clamp calculation, YuE guarantees that even when users maximize their memory_budget_gib setting, the system retains headroom for these essential operations. This design prevents the "death spiral" of out-of-memory errors that occur when allocation requests fail due to driver overhead rather than tensor size.
Summary
- The
memory_budget_gibclamp sets an absolute ceiling on GPU memory consumption by converting GiB values into per-process fractions. - Calculation logic in
src/yue2/storage.pysubtracts the 0.5 GiBSAFETY_RESERVE_GIBbefore converting to a fraction, usingmax(0.0, ...)to prevent negative values. - PyTorch integration occurs through
torch.cuda.set_per_process_memory_fraction, which enforces the limit at the allocator level. - Pipeline initialization in
src/yue2/pipeline.pyapplies these settings before model loading to ensure immediate enforcement. - The safety reserve guarantees kernel and driver stability by preserving a small memory buffer regardless of user configuration.
Frequently Asked Questions
What happens if I set memory_budget_gib smaller than the safety reserve?
If you request a budget below 0.5 GiB, the clamp forces usable_budget to 0.0 via the max(0.0, ...) operation. This results in a per-process memory fraction of zero, meaning PyTorch will immediately raise an out-of-memory error on the first allocation attempt. This "fail fast" behavior prevents the process from entering an unstable state with insufficient memory for basic operations.
How does YuE calculate the per-process memory fraction from the budget?
YuE retrieves the device's total memory capacity, subtracts the 0.5 GiB SAFETY_RESERVE_GIB, and divides the remaining value by the total memory. The result is passed to torch.cuda.set_per_process_memory_fraction in src/yue2/storage.py. For example, on a 24 GiB GPU with a 12 GiB budget, the effective fraction is approximately 0.48 (11.5 GiB usable divided by 24 GiB total).
Can I modify or disable the safety reserve constant?
The 0.5 GiB reserve is hardcoded as SAFETY_RESERVE_GIB = 0.5 at the module level in src/yue2/storage.py. While you can edit the source file to change this value, it is not exposed as a runtime configuration parameter. Modifying it requires changing the constant and reinstalling the package, as the design treats this buffer as a non-negotiable system requirement.
Where is the memory_budget_gib parameter initialized in the pipeline?
The parameter is first ingested in src/yue2/pipeline.py during pipeline instantiation, where user-supplied arguments (such as --memory_budget_gib from the CLI) are passed to set_memory_fraction_from_budget in src/yue2/storage.py. This initialization occurs before any model weights are downloaded or loaded into GPU memory, ensuring the limit is active from the first allocation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →