How the VRAM Cache‑Pressure Penalty Affects MoE Throughput Estimation in LLMFIT
The VRAM cache-pressure penalty is a scaling factor applied in LLMFIT's throughput calculations that reduces estimated tokens-per-second as VRAM utilization approaches physical limits during MoE-Offload execution.
LLMFIT is a Rust-based framework for determining whether large language models can run efficiently on specific hardware configurations. When evaluating Mixture-of-Experts (MoE) models in MoE-Offload mode, the tool applies a VRAM cache-pressure penalty to account for memory bandwidth contention between active expert weights and the KV-cache. This penalty ensures throughput estimates reflect real-world performance degradation when VRAM is saturated.
MoE-Offload Mode and Memory Contention
MoE models activate only a subset of their expert layers during each forward pass. In MoE-Offload mode, LLMFIT keeps the currently active experts resident in VRAM while storing inactive experts in system RAM. This strategy reduces the baseline VRAM footprint but introduces a dynamic memory requirement: the KV-cache, which holds per-token activations, must also reside in VRAM during inference.
As the model weights and KV-cache compete for limited VRAM capacity, memory bus saturation occurs. The cache-pressure penalty quantifies this contention by measuring the ratio of required memory to available memory, then scaling the raw throughput estimate inversely to that pressure.
Calculating the Cache-Pressure Penalty
The penalty calculation resides in llmfit-core/src/fit.rs within the VRAM accounting logic for RunMode::MoeOffload. The implementation computes usable VRAM, adds the KV-cache size, and derives a pressure factor that grows as utilization exceeds safety margins.
// llmfit-core/src/fit.rs – lines ≈520-540
let vram_needed = model_vram_gb + kv_cache_gb;
let vram_available = system_vram_gb - vram_reserve;
let pressure = (vram_needed / vram_available).max(1.0);
let cache_pressure_penalty = 1.0 / pressure; // ≤ 1.0
let adjusted_throughput = raw_throughput * cache_pressure_penalty;
The pressure variable represents the utilization ratio of available VRAM. When vram_needed approaches vram_available, the pressure value exceeds 1.0, causing cache_pressure_penalty to drop below 1.0 and linearly reduce the adjusted_throughput. This inverse relationship reflects the throughput cost of managing cache evictions and memory stalls on a saturated GPU.
Impact on Throughput Estimates and Scoring
The magnitude of the penalty directly determines how LLMFIT categorizes hardware-model compatibility. The system uses the adjusted throughput value in the score_fit function to assign one of four fit ratings:
- Perfect: VRAM is loosely used (
pressure≈ 1.0), penalty is near 1.0, and throughput remains at the raw theoretical maximum. - Good: Moderate pressure from the KV-cache causes a small penalty factor (0.8–0.99), slightly reducing effective tokens-per-second.
- Marginal: High pressure (pressure > 1.2) triggers a significant penalty (0.5–0.8), indicating the system can run the model but with noticeable slowdowns.
- TooTight: When
vram_neededexceedsvram_availableafter reserves, the pressure factor spikes, the penalty collapses toward zero, and LLMFIT rejects the configuration as unfit.
This categorization prevents users from deploying MoE models on hardware where cache thrashing would render inference impractical.
Implementation in the Codebase
The penalty logic integrates across multiple core modules. The ModelFit::analyze_with_forced_runtime method in llmfit-core/src/fit.rs orchestrates the calculation, while llmfit-core/src/plan.rs determines whether to invoke RunMode::MoeOffload based on model metadata from llmfit-core/src/models.rs.
use llmfit_core::fit::{ModelFit, RunMode};
use llmfit_core::models::Model;
// Load a MoE model (e.g., Qwen3-MoE-80B)
let model = Model::load("Qwen3-MoE-80B")?;
// Force MoE-Offload evaluation on 24 GiB VRAM hardware
let fit = ModelFit::analyze_with_forced_runtime(&model, RunMode::MoeOffload)?;
println!("Estimated throughput: {:.2} tokens/s", fit.throughput);
println!("Cache pressure penalty: {:.2}", fit.cache_pressure_penalty);
The llmfit-tui/src/display.rs module renders these values in the terminal interface, showing users both the raw theoretical throughput and the penalty-adjusted realistic estimate. Non-MoE models bypass this logic entirely, using standard GPU-only or CPU-offload calculations instead.
Summary
- VRAM cache-pressure penalty: A scaling factor (0.0–1.0) in
llmfit-core/src/fit.rsthat reduces MoE throughput estimates based on VRAM saturation levels. - Trigger condition: Applied exclusively during
RunMode::MoeOffloadwhen active expert weights and KV-cache compete for GPU memory bandwidth. - Calculation method: Inverse of the ratio between required VRAM (
model_vram_gb + kv_cache_gb) and available VRAM (system_vram_gb - vram_reserve). - Outcome impact: Determines fit scores (Perfect, Good, Marginal, TooTight) by adjusting raw throughput to reflect real-world cache contention slowdowns.
- Code locations: Core logic in
fit.rs, run-mode selection inplan.rs, model classification inmodels.rs, and UI display indisplay.rs.
Frequently Asked Questions
What is the VRAM cache-pressure penalty in LLMFIT?
The VRAM cache-pressure penalty is a throughput scaling factor that accounts for performance degradation when the KV-cache and active expert weights compete for limited GPU memory during MoE inference. It is calculated as the inverse of the VRAM utilization ratio and applied to raw throughput estimates in llmfit-core/src/fit.rs.
How does MoE-Offload mode trigger the penalty calculation?
RunMode::MoeOffload activates the penalty branch in the fitting logic because this mode specifically splits expert storage between VRAM and system RAM, creating dynamic memory pressure from the KV-cache. Standard GPU-only modes do not apply this penalty because they assume static memory allocation without runtime cache contention.
Can the cache-pressure penalty value exceed 1.0?
No. The penalty is clamped to a maximum of 1.0 using the calculation 1.0 / pressure, where pressure is the result of (vram_needed / vram_available).max(1.0). This ensures the penalty never artificially inflates throughput; it only reduces or maintains the estimate.
Where is the penalty logic located in the source code?
The primary implementation resides in llmfit-core/src/fit.rs within the VRAM requirement calculation functions, specifically around lines 520–540. Supporting logic for determining when to apply the penalty exists in llmfit-core/src/plan.rs, and the model metadata that identifies MoE architectures is defined in llmfit-core/src/models.rs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →