How LLMFIT Estimates MoE Model Speed Using Active vs. Full Parameters
LLMFIT's throughput estimator uses active parameters for DDR-bound MoE offload calculations and full model parameters for VRAM-bound GPU execution, delivering accurate tokens-per-second predictions for sparse expert architectures.
LLMFIT (AlexsJones/llmfit) provides precise throughput predictions for large language models by adapting its estimation logic to model architecture. For Mixture-of-Experts (MoE) architectures, the tool differentiates between active parameters—the subset of weights activated per token—and full model parameters to compute realistic tokens-per-second (TPS) estimates based on hardware constraints and execution mode.
Parameter Selection Logic in estimate_tps
The core estimation function estimate_tps in llmfit-core/src/fit.rs implements conditional parameter counting for MoE models at lines 1191–1198. The estimator first attempts to retrieve the active_parameters field from the model metadata, which represents the actual number of parameters engaged during a single forward pass.
If active_parameters is present, the converter transforms this value into billions of parameters (active_gb) for subsequent bandwidth calculations. When this field is absent, the system falls back to model.params_b(), utilizing the full model size as the bandwidth denominator. This branching logic ensures that sparse MoE models receive appropriate throughput estimates that reflect their dynamic computation graphs rather than their static storage footprint.
Execution Path Divergence for MoE Architectures
LLMFIT branches into two distinct computational paths based on the RunMode configuration, each handling parameter counts differently to reflect hardware reality.
MoeOffload Mode: DDR-Bandwidth Limited
When operating in RunMode::MoeOffload, inactive expert weights reside in system RAM while only active experts stream across the memory bus. At lines 1222–1230 and 1262–1270 of llmfit-core/src/fit.rs, the estimator calculates per-token latency using the active parameter count divided by measured DDR bandwidth:
// Conceptual representation of the offload calculation
let active_gb = active_parameters as f64 / 1e9;
let tps = ddr_bandwidth / active_gb * efficiency_factor;
This approach acknowledges that offloaded MoE inference is constrained by how quickly active expert weights can transfer from DDR memory, not by the total model size residing on disk or partially in RAM.
GPU Mode: VRAM Cache Pressure
When the model fits entirely within VRAM (RunMode::Gpu), inference runtimes typically load all expert weights into GPU memory simultaneously to minimize latency variation. At lines 1281–1294, LLMFIT reverts to the dense-model calculation formula using the full parameter count:
// VRAM-bound calculation uses total model size
let full_model_gb = model.params_b();
let tps = vram_bandwidth / full_model_gb * efficiency;
This conservative estimate accounts for cache pressure and the bandwidth cost of maintaining the complete parameter set in fast memory, regardless of which experts activate per token.
Performance Impact of Parameter Accounting
Using active parameters for offload scenarios yields realistic throughput estimates that often project 4× higher TPS than full-parameter calculations for highly sparse MoE architectures. For example, Qwen-3-Next-80B achieves approximately 15 TPS in offload mode according to LLMFIT's active-parameter model, whereas a dense calculation would significantly underestimate performance by assuming all 80 billion parameters transfer per token.
Conversely, GPU mode estimations prevent over-optimistic projections by recognizing that VRAM bandwidth serves the entire model weight matrix, not just the active subset. This dual-mode approach aligns LLMFIT's predictions with observed benchmarks across different deployment configurations.
Implementation Examples
The following Rust examples demonstrate the API usage for both execution modes:
// Example: estimating TPS for a MoE model in offload mode
let tps = estimate_tps(
&moe_model, // LlmModel with active_parameters = Some(3_000_000_000)
"q4_k_m", // quantization level
&system_specs, // SystemSpecs with measured DDR bandwidth
RunMode::MoeOffload, // Offload mode
InferenceRuntime::Ollama,
&calc_config,
);
println!("Estimated TPS (offload): {:.1}", tps);
For GPU deployment of the same model:
// Example: estimating TPS for the same MoE model in GPU mode
let tps = estimate_tps(
&moe_model,
"q4_k_m",
&system_specs,
RunMode::Gpu, // GPU mode – full-model VRAM bandwidth applies
InferenceRuntime::Ollama,
&calc_config,
);
println!("Estimated TPS (GPU): {:.1}", tps);
Summary
- Active parameters represent the subset of MoE weights activated per token and are used for DDR-bandwidth calculations in offload mode.
- The
estimate_tpsfunction inllmfit-core/src/fit.rs(lines 1191–1198) prioritizesactive_parametersmetadata when available, falling back tomodel.params_b()only when necessary. - MoeOffload mode calculates throughput using
active_gb / ddr_bandwidthto reflect streaming expert loading from system RAM (lines 1222–1230). - GPU mode uses full model size for bandwidth calculations due to complete parameter residency in VRAM (lines 1281–1294).
- This distinction prevents underestimation of offload performance while avoiding overestimation of GPU throughput for sparse architectures.
Frequently Asked Questions
What are active parameters in MoE models?
Active parameters refer to the specific subset of expert weights that the gating network selects and loads for processing an individual input token. In sparse Mixture-of-Experts architectures, only a fraction of the total parameter count participates in each forward pass, unlike dense models where all parameters activate for every token.
How does LLMFIT handle missing active parameter metadata?
When the active_parameters field is absent from the model configuration, LLMFIT automatically falls back to model.params_b() in llmfit-core/src/fit.rs. This ensures the estimator remains functional for older model definitions or dense architectures while potentially producing conservative throughput estimates for MoE models lacking metadata.
Why does GPU mode use full parameters instead of active parameters?
GPU resident execution typically loads all expert weights into VRAM simultaneously to eliminate the latency of dynamic loading. Even though only active experts compute, the full parameter set occupies memory bandwidth during cache maintenance and potential prefetching operations. LLMFIT therefore uses the full model size to account for VRAM bandwidth pressure and cache coherency costs.
Where is the MoE speed estimation logic located in the codebase?
The primary implementation resides in llmfit-core/src/fit.rs within the estimate_tps function. Supporting definitions exist in llmfit-core/src/models.rs where LlmModel stores the active_parameters field, while llmfit-core/src/plan.rs references these estimates during deployment planning. The CLI display logic appears in llmfit-tui/src/display.rs for user-facing throughput reports.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →