# How llmfit Estimates Inference Speed for Mixture-of-Experts (MoE) Models

> llmfit estimates MoE model inference speed by analyzing active expert weights and memory bandwidth, providing physics-based formulas for GPU and offloaded configurations.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: deep-dive
- Published: 2026-08-22

---

**`llmfit` calculates tokens-per-second (TPS) for MoE models by measuring active-expert weights against available memory bandwidth, applying distinct physics-based formulas for GPU-loaded versus DDR-streamed (offloaded) configurations.**

`llmfit` is a Rust-based inference profiler that delivers grounded speed predictions for large language models. For **Mixture-of-Experts (MoE) speed estimation**, the tool accounts for the fact that only a subset of experts is active per token, modeling two distinct execution paths: one where inactive experts stream from system RAM and another where all experts reside in VRAM.

## Core Speed Estimation Logic in `estimate_tps`

The heart of the calculation lives in the **`estimate_tps`** function inside [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) (lines [112‑119](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs#L112-L119)). This routine aggregates three sequential stages to produce a final throughput number that reflects real-world data-movement constraints.

## Stage 1: Active Parameter Selection for MoE Models

MoE architectures route each token through only a fraction of the total expert weights. Because of this, `llmfit` uses the **active parameter count** rather than the full model size when calculating memory traffic.

According to the source code at lines [192‑199](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs#L192-L199), the estimator first attempts to resolve the number of active experts. If this value is unknown, the code conservatively falls back to the full model size to avoid underestimating resource requirements.

## Stage 2: Bandwidth Source Selection

The estimator selects a bottleneck bandwidth based on the execution strategy decided by the path-selection logic. There are two primary sources:

### DDR Bandwidth for MoE Offloading

When running in **`RunMode::MoeOffload`**, the active experts live in VRAM while inactive experts remain in system RAM and are streamed on demand. This makes **DDR memory bandwidth** the limiting factor.

At lines [154‑161](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs#L154-L161), `ddr_bandwidth_gbps` resolves through a priority chain:
1. User-provided `CalcConfig`
2. Environment variable override
3. Measured system value
4. Conservative 50 GB/s fallback

### GPU Memory Bandwidth for Full VRAM

If the model fits entirely in GPU memory (`RunMode::Gpu`), the bottleneck shifts to the graphics card’s memory subsystem. Lines [218‑220](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs#L218-L220) pull from an internal lookup table (`gpu_memory_bandwidth_gbps`) indexed by the detected GPU architecture.

## Stage 3: Tokens-Per-Second Calculation

With the active parameter count and bandwidth identified, `estimate_tps` branches into mode-specific arithmetic.

### MoE Offload Mode Calculation

For the offload path (lines [262‑268](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs#L262-L268)), the time per token equals the sum of:
- **DDR read time** for the active-expert weights (bytes ÷ DDR bandwidth)
- **GPU compute time** for those same weights

The raw TPS is `1 / total_time`. This value is then multiplied by a tunable efficiency factor (`config.run_mode_factors`, default **0.8**) at lines [277‑279](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs#L277-L279) to account for kernel overhead and pipeline bubbles.

### GPU Mode Calculation

When all experts reside in VRAM (lines [306‑322](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs#L306-L322)), the byte-count per token splits into two components:
- **Scalable part**: Active-expert FFN weights, scaled by quantization level (e.g., Q4, Q8)
- **Fixed part**: Attention layers, router, shared experts, and LM head, captured by the constant **`MOE_FIXED_EFFECTIVE_BPP`**

The estimator divides GPU bandwidth by the total bytes and applies the same 0.8 run-mode factor.

### Cache-Pressure Penalty

GPU mode applies a **cache-pressure penalty** when VRAM utilization exceeds 60 %. Lines [336‑354](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs#L336-L354) calculate this penalty as a function of the ratio of inactive experts to total experts, reducing the TPS estimate for heavily loaded MoE models that stress the memory hierarchy.

## MoE Path Selection Logic

The speed model depends on accurate memory-path selection. Two helper functions govern this:

- **`moe_offload_path`** (lines [232‑260](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs#L232-L260)): Attempts to fit active-expert VRAM plus offloaded RAM; falls back to CPU-offload or pure GPU if capacity is exceeded.
- **`moe_memory_for_quant`** (lines [284‑293](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs#L284-L293)): Computes exact VRAM and RAM requirements for a given quantization scheme, feeding the data into the bandwidth-selection logic.

## Practical Usage Examples

You can access these estimates programmatically or via the CLI.

```rust
// Example: programmatically obtain a speed estimate for a MoE model.
use llmfit_core::{fit::ModelFit, hardware::SystemSpecs, models::LlmModel};

let system = SystemSpecs::detect();               // detects GPUs, RAM, etc.
let model  = LlmModel::load_from_name("Qwen3-Next-80B")?; // MoE model entry

// Run the analysis with defaults; the returned ModelFit contains `estimated_tps`.
let fit = ModelFit::analyze(&model, &system);

println!(
    "MoE model '{}' estimated speed: {:.1} tok/s ({:?} mode)",
    fit.model.name,
    fit.estimated_tps,
    fit.run_mode,
);

```

```bash

# CLI usage – the `fit` subcommand automatically selects MoE‑offload when needed.

$ llmfit fit "Qwen3-Next-80B" --max-context 8192
Model          Run‑Mode   Speed (tok/s)   Notes
-------------------------------------------------------------
Qwen3‑Next‑80B  MoE‑Offload   15.2          MoE: 8/64 experts active in VRAM

```

## Summary

- **`estimate_tps`** in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) drives all MoE speed calculations, beginning at line 112.
- **Active-expert counts** replace full parameter counts at lines 192‑199, reflecting sparse MoE execution.
- **DDR bandwidth** (lines 154‑161) limits speed in `MoeOffload` mode, while **GPU bandwidth** (lines 218‑220) limits full-VRAM mode.
- **Offload TPS** equals `1 / (DDR_read_time + GPU_compute_time)` multiplied by a 0.8 efficiency factor (lines 262‑279).
- **GPU-mode TPS** divides bandwidth by scalable plus fixed byte counts, applying a cache-pressure penalty when VRAM utilization exceeds 60 % (lines 306‑354).

## Frequently Asked Questions

### How does llmfit determine whether to use MoE offloading or full GPU mode?

The `moe_offload_path` function (lines [232‑260](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs#L232-L260)) evaluates whether the active experts fit in VRAM while the remaining weights stay in system RAM. If this allocation fails, the logic falls back to CPU offloading or pure GPU execution, whichever satisfies memory constraints.

### What is the default DDR bandwidth assumption if the system cannot be probed?

When `llmfit` cannot detect DDR bandwidth from the `CalcConfig`, environment variables, or hardware introspection, it defaults to **50 GB/s** as a conservative floor (lines [154‑161](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs#L154-L161)).

### Why does the GPU mode calculation use a fixed bytes-per-parameter constant?

The constant **`MOE_FIXED_EFFECTIVE_BPP`** accounts for memory traffic from attention layers, the expert router, shared experts, and the LM head—components that do not scale with the number of active experts. This ensures the bandwidth calculation captures all data movement, not just the sparse FFN weights (lines [306‑322](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs#L306-L322)).

### At what VRAM utilization does llmfit apply a cache-pressure penalty?

The penalty activates when VRAM utilization exceeds **60 %**. The magnitude grows with the ratio of inactive to total experts, reducing the TPS estimate to reflect cache thrashing and eviction overhead (lines [336‑354](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs#L336-L354)).