How llmfit Estimates Inference Speed for Mixture-of-Experts (MoE) Models
llmfit calculates tokens-per-second (TPS) for MoE models by measuring active-expert weights against available memory bandwidth, applying distinct physics-based formulas for GPU-loaded versus DDR-streamed (offloaded) configurations.
llmfit is a Rust-based inference profiler that delivers grounded speed predictions for large language models. For Mixture-of-Experts (MoE) speed estimation, the tool accounts for the fact that only a subset of experts is active per token, modeling two distinct execution paths: one where inactive experts stream from system RAM and another where all experts reside in VRAM.
Core Speed Estimation Logic in estimate_tps
The heart of the calculation lives in the estimate_tps function inside llmfit-core/src/fit.rs (lines 112‑119). This routine aggregates three sequential stages to produce a final throughput number that reflects real-world data-movement constraints.
Stage 1: Active Parameter Selection for MoE Models
MoE architectures route each token through only a fraction of the total expert weights. Because of this, llmfit uses the active parameter count rather than the full model size when calculating memory traffic.
According to the source code at lines 192‑199, the estimator first attempts to resolve the number of active experts. If this value is unknown, the code conservatively falls back to the full model size to avoid underestimating resource requirements.
Stage 2: Bandwidth Source Selection
The estimator selects a bottleneck bandwidth based on the execution strategy decided by the path-selection logic. There are two primary sources:
DDR Bandwidth for MoE Offloading
When running in RunMode::MoeOffload, the active experts live in VRAM while inactive experts remain in system RAM and are streamed on demand. This makes DDR memory bandwidth the limiting factor.
At lines 154‑161, ddr_bandwidth_gbps resolves through a priority chain:
- User-provided
CalcConfig - Environment variable override
- Measured system value
- Conservative 50 GB/s fallback
GPU Memory Bandwidth for Full VRAM
If the model fits entirely in GPU memory (RunMode::Gpu), the bottleneck shifts to the graphics card’s memory subsystem. Lines 218‑220 pull from an internal lookup table (gpu_memory_bandwidth_gbps) indexed by the detected GPU architecture.
Stage 3: Tokens-Per-Second Calculation
With the active parameter count and bandwidth identified, estimate_tps branches into mode-specific arithmetic.
MoE Offload Mode Calculation
For the offload path (lines 262‑268), the time per token equals the sum of:
- DDR read time for the active-expert weights (bytes ÷ DDR bandwidth)
- GPU compute time for those same weights
The raw TPS is 1 / total_time. This value is then multiplied by a tunable efficiency factor (config.run_mode_factors, default 0.8) at lines 277‑279 to account for kernel overhead and pipeline bubbles.
GPU Mode Calculation
When all experts reside in VRAM (lines 306‑322), the byte-count per token splits into two components:
- Scalable part: Active-expert FFN weights, scaled by quantization level (e.g., Q4, Q8)
- Fixed part: Attention layers, router, shared experts, and LM head, captured by the constant
MOE_FIXED_EFFECTIVE_BPP
The estimator divides GPU bandwidth by the total bytes and applies the same 0.8 run-mode factor.
Cache-Pressure Penalty
GPU mode applies a cache-pressure penalty when VRAM utilization exceeds 60 %. Lines 336‑354 calculate this penalty as a function of the ratio of inactive experts to total experts, reducing the TPS estimate for heavily loaded MoE models that stress the memory hierarchy.
MoE Path Selection Logic
The speed model depends on accurate memory-path selection. Two helper functions govern this:
moe_offload_path(lines 232‑260): Attempts to fit active-expert VRAM plus offloaded RAM; falls back to CPU-offload or pure GPU if capacity is exceeded.moe_memory_for_quant(lines 284‑293): Computes exact VRAM and RAM requirements for a given quantization scheme, feeding the data into the bandwidth-selection logic.
Practical Usage Examples
You can access these estimates programmatically or via the CLI.
// Example: programmatically obtain a speed estimate for a MoE model.
use llmfit_core::{fit::ModelFit, hardware::SystemSpecs, models::LlmModel};
let system = SystemSpecs::detect(); // detects GPUs, RAM, etc.
let model = LlmModel::load_from_name("Qwen3-Next-80B")?; // MoE model entry
// Run the analysis with defaults; the returned ModelFit contains `estimated_tps`.
let fit = ModelFit::analyze(&model, &system);
println!(
"MoE model '{}' estimated speed: {:.1} tok/s ({:?} mode)",
fit.model.name,
fit.estimated_tps,
fit.run_mode,
);
# CLI usage – the `fit` subcommand automatically selects MoE‑offload when needed.
$ llmfit fit "Qwen3-Next-80B" --max-context 8192
Model Run‑Mode Speed (tok/s) Notes
-------------------------------------------------------------
Qwen3‑Next‑80B MoE‑Offload 15.2 MoE: 8/64 experts active in VRAM
Summary
estimate_tpsinllmfit-core/src/fit.rsdrives all MoE speed calculations, beginning at line 112.- Active-expert counts replace full parameter counts at lines 192‑199, reflecting sparse MoE execution.
- DDR bandwidth (lines 154‑161) limits speed in
MoeOffloadmode, while GPU bandwidth (lines 218‑220) limits full-VRAM mode. - Offload TPS equals
1 / (DDR_read_time + GPU_compute_time)multiplied by a 0.8 efficiency factor (lines 262‑279). - GPU-mode TPS divides bandwidth by scalable plus fixed byte counts, applying a cache-pressure penalty when VRAM utilization exceeds 60 % (lines 306‑354).
Frequently Asked Questions
How does llmfit determine whether to use MoE offloading or full GPU mode?
The moe_offload_path function (lines 232‑260) evaluates whether the active experts fit in VRAM while the remaining weights stay in system RAM. If this allocation fails, the logic falls back to CPU offloading or pure GPU execution, whichever satisfies memory constraints.
What is the default DDR bandwidth assumption if the system cannot be probed?
When llmfit cannot detect DDR bandwidth from the CalcConfig, environment variables, or hardware introspection, it defaults to 50 GB/s as a conservative floor (lines 154‑161).
Why does the GPU mode calculation use a fixed bytes-per-parameter constant?
The constant MOE_FIXED_EFFECTIVE_BPP accounts for memory traffic from attention layers, the expert router, shared experts, and the LM head—components that do not scale with the number of active experts. This ensures the bandwidth calculation captures all data movement, not just the sparse FFN weights (lines 306‑322).
At what VRAM utilization does llmfit apply a cache-pressure penalty?
The penalty activates when VRAM utilization exceeds 60 %. The magnitude grows with the ratio of inactive to total experts, reducing the TPS estimate to reflect cache thrashing and eviction overhead (lines 336‑354).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →