# How the VRAM Cache‑Pressure Penalty Affects MoE Throughput Estimation in LLMFIT

> Discover how VRAM cache-pressure penalty impacts MoE throughput estimation in LLMFIT. Learn how utilization limits affect tokens-per-second calculations.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: deep-dive
- Published: 2026-08-21

---

**The VRAM cache-pressure penalty is a scaling factor applied in LLMFIT's throughput calculations that reduces estimated tokens-per-second as VRAM utilization approaches physical limits during MoE-Offload execution.**

LLMFIT is a Rust-based framework for determining whether large language models can run efficiently on specific hardware configurations. When evaluating **Mixture-of-Experts (MoE)** models in **MoE-Offload** mode, the tool applies a **VRAM cache-pressure penalty** to account for memory bandwidth contention between active expert weights and the KV-cache. This penalty ensures throughput estimates reflect real-world performance degradation when VRAM is saturated.

## MoE-Offload Mode and Memory Contention

MoE models activate only a subset of their expert layers during each forward pass. In **MoE-Offload** mode, LLMFIT keeps the currently active experts resident in VRAM while storing inactive experts in system RAM. This strategy reduces the baseline VRAM footprint but introduces a dynamic memory requirement: the **KV-cache**, which holds per-token activations, must also reside in VRAM during inference.

As the model weights and KV-cache compete for limited VRAM capacity, memory bus saturation occurs. The cache-pressure penalty quantifies this contention by measuring the ratio of required memory to available memory, then scaling the raw throughput estimate inversely to that pressure.

## Calculating the Cache-Pressure Penalty

The penalty calculation resides in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) within the VRAM accounting logic for `RunMode::MoeOffload`. The implementation computes usable VRAM, adds the KV-cache size, and derives a pressure factor that grows as utilization exceeds safety margins.

```rust
// llmfit-core/src/fit.rs – lines ≈520-540
let vram_needed = model_vram_gb + kv_cache_gb;
let vram_available = system_vram_gb - vram_reserve;
let pressure = (vram_needed / vram_available).max(1.0);
let cache_pressure_penalty = 1.0 / pressure;          // ≤ 1.0
let adjusted_throughput = raw_throughput * cache_pressure_penalty;

```

The `pressure` variable represents the utilization ratio of available VRAM. When `vram_needed` approaches `vram_available`, the `pressure` value exceeds 1.0, causing `cache_pressure_penalty` to drop below 1.0 and linearly reduce the `adjusted_throughput`. This inverse relationship reflects the throughput cost of managing cache evictions and memory stalls on a saturated GPU.

## Impact on Throughput Estimates and Scoring

The magnitude of the penalty directly determines how LLMFIT categorizes hardware-model compatibility. The system uses the adjusted throughput value in the `score_fit` function to assign one of four fit ratings:

- **Perfect**: VRAM is loosely used (`pressure` ≈ 1.0), penalty is near 1.0, and throughput remains at the raw theoretical maximum.
- **Good**: Moderate pressure from the KV-cache causes a small penalty factor (0.8–0.99), slightly reducing effective tokens-per-second.
- **Marginal**: High pressure (pressure > 1.2) triggers a significant penalty (0.5–0.8), indicating the system can run the model but with noticeable slowdowns.
- **TooTight**: When `vram_needed` exceeds `vram_available` after reserves, the pressure factor spikes, the penalty collapses toward zero, and LLMFIT rejects the configuration as unfit.

This categorization prevents users from deploying MoE models on hardware where cache thrashing would render inference impractical.

## Implementation in the Codebase

The penalty logic integrates across multiple core modules. The `ModelFit::analyze_with_forced_runtime` method in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) orchestrates the calculation, while [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs) determines whether to invoke `RunMode::MoeOffload` based on model metadata from [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs).

```rust
use llmfit_core::fit::{ModelFit, RunMode};
use llmfit_core::models::Model;

// Load a MoE model (e.g., Qwen3-MoE-80B)
let model = Model::load("Qwen3-MoE-80B")?;

// Force MoE-Offload evaluation on 24 GiB VRAM hardware
let fit = ModelFit::analyze_with_forced_runtime(&model, RunMode::MoeOffload)?;
println!("Estimated throughput: {:.2} tokens/s", fit.throughput);
println!("Cache pressure penalty: {:.2}", fit.cache_pressure_penalty);

```

The [`llmfit-tui/src/display.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/display.rs) module renders these values in the terminal interface, showing users both the raw theoretical throughput and the penalty-adjusted realistic estimate. Non-MoE models bypass this logic entirely, using standard GPU-only or CPU-offload calculations instead.

## Summary

- **VRAM cache-pressure penalty**: A scaling factor (0.0–1.0) in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) that reduces MoE throughput estimates based on VRAM saturation levels.
- **Trigger condition**: Applied exclusively during `RunMode::MoeOffload` when active expert weights and KV-cache compete for GPU memory bandwidth.
- **Calculation method**: Inverse of the ratio between required VRAM (`model_vram_gb + kv_cache_gb`) and available VRAM (`system_vram_gb - vram_reserve`).
- **Outcome impact**: Determines fit scores (Perfect, Good, Marginal, TooTight) by adjusting raw throughput to reflect real-world cache contention slowdowns.
- **Code locations**: Core logic in [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs), run-mode selection in [`plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/plan.rs), model classification in [`models.rs`](https://github.com/AlexsJones/llmfit/blob/main/models.rs), and UI display in [`display.rs`](https://github.com/AlexsJones/llmfit/blob/main/display.rs).

## Frequently Asked Questions

### What is the VRAM cache-pressure penalty in LLMFIT?

The VRAM cache-pressure penalty is a throughput scaling factor that accounts for performance degradation when the KV-cache and active expert weights compete for limited GPU memory during MoE inference. It is calculated as the inverse of the VRAM utilization ratio and applied to raw throughput estimates in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs).

### How does MoE-Offload mode trigger the penalty calculation?

`RunMode::MoeOffload` activates the penalty branch in the fitting logic because this mode specifically splits expert storage between VRAM and system RAM, creating dynamic memory pressure from the KV-cache. Standard GPU-only modes do not apply this penalty because they assume static memory allocation without runtime cache contention.

### Can the cache-pressure penalty value exceed 1.0?

No. The penalty is clamped to a maximum of 1.0 using the calculation `1.0 / pressure`, where `pressure` is the result of `(vram_needed / vram_available).max(1.0)`. This ensures the penalty never artificially inflates throughput; it only reduces or maintains the estimate.

### Where is the penalty logic located in the source code?

The primary implementation resides in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) within the VRAM requirement calculation functions, specifically around lines 520–540. Supporting logic for determining when to apply the penalty exists in [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs), and the model metadata that identifies MoE architectures is defined in [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs).