The Difference Between Automatic and Explicit Expert Cache Budgets for Metal SSD Streaming

The ds4 engine determines the size of its expert cache through either an automatic mode that calculates a dynamic budget based on host RAM and GPU memory availability, or an explicit mode that uses a fixed user-defined value set via environment variables like DS4_METAL_EXPERT_CACHE_BUDGET_GIB.

The antirez/ds4 project streams Mixture-of-Experts (MoE) model weights directly from SSD storage to Apple Metal GPUs, minimizing memory overhead during inference. Central to this architecture is the expert cache, a memory buffer that holds recently accessed expert tensors. The maximum size of this cache—the expert cache budget—can be determined through two distinct mechanisms implemented in ds4_metal.m.

How the Expert Cache Budget Works

When ds4 initializes its Metal GPU streaming subsystem, it must decide how much system memory to allocate for caching expert weights. This decision prevents the cache from consuming excessive RAM while ensuring sufficient space to avoid redundant SSD reads. The budget controls the g_stream_expert_cache array, which stores the cached expert tensors. Depending on configuration, the budget calculation either adapts to current system conditions or remains fixed to a user-specified value.

Automatic Budget Mode: Dynamic System-Based Calculation

In automatic mode, the engine computes the cache budget at runtime by inspecting the host’s physical memory and the GPU’s available memory. If the user does not supply an explicit value, the function ds4_gpu_stream_expert_cache_configured_budget() in ds4_metal.m falls back to a calculation based on system memory obtained via ds4_gpu_system_memory_bytes().

This approach allows the cache to grow or shrink automatically as system memory pressure changes. The automatic calculation includes built-in safety caps to ensure the cache never exceeds a reasonable fraction of available GPU or system memory, protecting other applications and system stability.

// Simplified logic from ds4_metal.m
size_t ds4_gpu_stream_expert_cache_configured_budget(void) {
    if (g_stream_expert_cache_budget_override > 0) {
        return g_stream_expert_cache_budget_override; // Explicit mode
    }
    // Automatic mode: calculate from system memory
    size_t system_memory = ds4_gpu_system_memory_bytes();
    return calculate_default_budget(system_memory);
}

Explicit Budget Mode: User-Defined Fixed Limits

Explicit mode bypasses the automatic calculation entirely. Users can override the dynamic budgeting by setting the DS4_METAL_EXPERT_CACHE_BUDGET_GIB or DS4_METAL_EXPERT_CACHE_BUDGET_MIB environment variables before launching the application. When present, these values are parsed early in ds4_metal.m and stored in the global variable g_stream_expert_cache_budget_override.

The function ds4_gpu_stream_expert_cache_configured_budget() then returns this override value directly, fixing the cache size to the user-provided limit regardless of current system memory availability or pressure. This mode is useful when developers require deterministic memory usage or need to allocate more (or less) cache space than the automatic heuristic would select.


# Set explicit 16 GiB cache budget

export DS4_METAL_EXPERT_CACHE_BUDGET_GIB=16
./ds4_inference_app

# Alternative using MiB

export DS4_METAL_EXPERT_CACHE_BUDGET_MIB=16384
./ds4_inference_app

Key Differences Between Automatic and Explicit Budgets

Feature Automatic Budget Explicit Budget
Determination Method Runtime calculation via ds4_gpu_system_memory_bytes() User-supplied constant from environment variables
Storage Location Calculated value returned directly from ds4_gpu_stream_expert_cache_configured_budget() Stored in g_stream_expert_cache_budget_override global variable
Adaptability Adjusts to system memory pressure and GPU availability Fixed for the application lifetime
Use Case General deployment where system conditions vary Testing, deterministic behavior, or specific hardware constraints

Both modes ultimately populate and manage the same g_stream_expert_cache data structure, but they differ fundamentally in who controls the sizing: the runtime heuristic (automatic) or the user configuration (explicit).

Configuring the Expert Cache in Practice

To verify which mode is active, inspect the return value of ds4_gpu_stream_expert_cache_configured_budget() at runtime. If the global g_stream_expert_cache_budget_override is non-zero, the explicit mode is engaged; otherwise, the function computes the budget dynamically from system metrics.

// Check configuration mode in ds4_metal.m
size_t budget = ds4_gpu_stream_expert_cache_configured_budget();
if (g_stream_expert_cache_budget_override > 0) {
    printf("Explicit budget mode: %zu bytes\n", budget);
} else {
    printf("Automatic budget mode: %zu bytes (calculated from system RAM)\n", budget);
}

When deploying ds4 on systems with varying memory configurations, automatic mode provides resilience against out-of-memory errors. For performance benchmarking or constrained environments, explicit mode ensures consistent cache behavior across runs.

Summary

  • Automatic expert cache budgets are calculated dynamically in ds4_metal.m using ds4_gpu_system_memory_bytes(), adapting to available host RAM and GPU memory without user intervention.
  • Explicit expert cache budgets are set via DS4_METAL_EXPERT_CACHE_BUDGET_GIB or DS4_METAL_EXPERT_CACHE_BUDGET_MIB environment variables and stored in g_stream_expert_cache_budget_override, providing fixed, deterministic cache sizing.
  • The function ds4_gpu_stream_expert_cache_configured_budget() serves as the single point of resolution, checking for explicit overrides before falling back to automatic calculation.
  • Both methods control the g_stream_expert_cache array, but automatic mode is preferred for general use while explicit mode suits specialized deployment scenarios.

Frequently Asked Questions

How does ds4 determine the expert cache size if no environment variables are set?

When DS4_METAL_EXPERT_CACHE_BUDGET_GIB and DS4_METAL_EXPERT_CACHE_BUDGET_MIB are absent, the global g_stream_expert_cache_budget_override remains zero. The function ds4_gpu_stream_expert_cache_configured_budget() in ds4_metal.m detects this condition and computes the budget automatically by calling ds4_gpu_system_memory_bytes() to query the host's physical memory, applying internal heuristics to determine a safe cache limit.

What environment variables control the explicit expert cache budget?

According to the source code in ds4_metal.m, users can set DS4_METAL_EXPERT_CACHE_BUDGET_GIB to specify the budget in gigabytes or DS4_METAL_EXPERT_CACHE_BUDGET_MIB to specify it in mebibytes. The initialization code checks these variables using getenv() and converts the string values to numeric bytes stored in g_stream_expert_cache_budget_override.

Where is the expert cache budget logic implemented in the ds4 source code?

The budget determination logic resides in ds4_metal.m. The key function ds4_gpu_stream_expert_cache_configured_budget() handles both automatic calculation and explicit override retrieval. The global variable g_stream_expert_cache_budget_override is declared in the same file, and the environment variable parsing occurs early in the Metal GPU initialization sequence.

Can the expert cache budget change after the application starts?

No. Once ds4_gpu_stream_expert_cache_configured_budget() is called during initialization, the budget is fixed for the application lifetime. In automatic mode, the calculation uses a snapshot of system memory at startup. In explicit mode, the value is read from the environment variable once and stored in g_stream_expert_cache_budget_override. The cache size does not dynamically resize in response to changing system conditions after this initial configuration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →