# How LLMFIT Estimates MoE Model Speed Using Active vs. Full Parameters

> LLMFIT estimates MoE model speed using active parameters for DDR offload and full parameters for GPU execution, providing accurate tokens per second predictions for sparse architectures.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: deep-dive
- Published: 2026-08-21

---

**LLMFIT's throughput estimator uses active parameters for DDR-bound MoE offload calculations and full model parameters for VRAM-bound GPU execution, delivering accurate tokens-per-second predictions for sparse expert architectures.**

LLMFIT (AlexsJones/llmfit) provides precise throughput predictions for large language models by adapting its estimation logic to model architecture. For **Mixture-of-Experts (MoE)** architectures, the tool differentiates between **active parameters**—the subset of weights activated per token—and **full model parameters** to compute realistic tokens-per-second (TPS) estimates based on hardware constraints and execution mode.

## Parameter Selection Logic in `estimate_tps`

The core estimation function `estimate_tps` in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) implements conditional parameter counting for MoE models at lines 1191–1198. The estimator first attempts to retrieve the `active_parameters` field from the model metadata, which represents the actual number of parameters engaged during a single forward pass.

If `active_parameters` is present, the converter transforms this value into billions of parameters (`active_gb`) for subsequent bandwidth calculations. When this field is absent, the system falls back to `model.params_b()`, utilizing the full model size as the bandwidth denominator. This branching logic ensures that sparse MoE models receive appropriate throughput estimates that reflect their dynamic computation graphs rather than their static storage footprint.

## Execution Path Divergence for MoE Architectures

LLMFIT branches into two distinct computational paths based on the `RunMode` configuration, each handling parameter counts differently to reflect hardware reality.

### MoeOffload Mode: DDR-Bandwidth Limited

When operating in `RunMode::MoeOffload`, inactive expert weights reside in system RAM while only active experts stream across the memory bus. At lines 1222–1230 and 1262–1270 of [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs), the estimator calculates per-token latency using the **active parameter count** divided by measured DDR bandwidth:

```rust
// Conceptual representation of the offload calculation
let active_gb = active_parameters as f64 / 1e9;
let tps = ddr_bandwidth / active_gb * efficiency_factor;

```

This approach acknowledges that offloaded MoE inference is constrained by how quickly active expert weights can transfer from DDR memory, not by the total model size residing on disk or partially in RAM.

### GPU Mode: VRAM Cache Pressure

When the model fits entirely within VRAM (`RunMode::Gpu`), inference runtimes typically load **all** expert weights into GPU memory simultaneously to minimize latency variation. At lines 1281–1294, LLMFIT reverts to the dense-model calculation formula using the full parameter count:

```rust
// VRAM-bound calculation uses total model size
let full_model_gb = model.params_b();
let tps = vram_bandwidth / full_model_gb * efficiency;

```

This conservative estimate accounts for cache pressure and the bandwidth cost of maintaining the complete parameter set in fast memory, regardless of which experts activate per token.

## Performance Impact of Parameter Accounting

Using active parameters for offload scenarios yields realistic throughput estimates that often project **4× higher TPS** than full-parameter calculations for highly sparse MoE architectures. For example, Qwen-3-Next-80B achieves approximately 15 TPS in offload mode according to LLMFIT's active-parameter model, whereas a dense calculation would significantly underestimate performance by assuming all 80 billion parameters transfer per token.

Conversely, GPU mode estimations prevent over-optimistic projections by recognizing that VRAM bandwidth serves the entire model weight matrix, not just the active subset. This dual-mode approach aligns LLMFIT's predictions with observed benchmarks across different deployment configurations.

## Implementation Examples

The following Rust examples demonstrate the API usage for both execution modes:

```rust
// Example: estimating TPS for a MoE model in offload mode
let tps = estimate_tps(
    &moe_model,               // LlmModel with active_parameters = Some(3_000_000_000)
    "q4_k_m",                 // quantization level
    &system_specs,            // SystemSpecs with measured DDR bandwidth
    RunMode::MoeOffload,      // Offload mode
    InferenceRuntime::Ollama,
    &calc_config,
);
println!("Estimated TPS (offload): {:.1}", tps);

```

For GPU deployment of the same model:

```rust
// Example: estimating TPS for the same MoE model in GPU mode
let tps = estimate_tps(
    &moe_model,
    "q4_k_m",
    &system_specs,
    RunMode::Gpu,             // GPU mode – full-model VRAM bandwidth applies
    InferenceRuntime::Ollama,
    &calc_config,
);
println!("Estimated TPS (GPU): {:.1}", tps);

```

## Summary

- **Active parameters** represent the subset of MoE weights activated per token and are used for DDR-bandwidth calculations in offload mode.
- The `estimate_tps` function in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) (lines 1191–1198) prioritizes `active_parameters` metadata when available, falling back to `model.params_b()` only when necessary.
- **MoeOffload mode** calculates throughput using `active_gb / ddr_bandwidth` to reflect streaming expert loading from system RAM (lines 1222–1230).
- **GPU mode** uses full model size for bandwidth calculations due to complete parameter residency in VRAM (lines 1281–1294).
- This distinction prevents underestimation of offload performance while avoiding overestimation of GPU throughput for sparse architectures.

## Frequently Asked Questions

### What are active parameters in MoE models?

Active parameters refer to the specific subset of expert weights that the gating network selects and loads for processing an individual input token. In sparse Mixture-of-Experts architectures, only a fraction of the total parameter count participates in each forward pass, unlike dense models where all parameters activate for every token.

### How does LLMFIT handle missing active parameter metadata?

When the `active_parameters` field is absent from the model configuration, LLMFIT automatically falls back to `model.params_b()` in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs). This ensures the estimator remains functional for older model definitions or dense architectures while potentially producing conservative throughput estimates for MoE models lacking metadata.

### Why does GPU mode use full parameters instead of active parameters?

GPU resident execution typically loads all expert weights into VRAM simultaneously to eliminate the latency of dynamic loading. Even though only active experts compute, the full parameter set occupies memory bandwidth during cache maintenance and potential prefetching operations. LLMFIT therefore uses the full model size to account for VRAM bandwidth pressure and cache coherency costs.

### Where is the MoE speed estimation logic located in the codebase?

The primary implementation resides in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) within the `estimate_tps` function. Supporting definitions exist in [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs) where `LlmModel` stores the `active_parameters` field, while [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs) references these estimates during deployment planning. The CLI display logic appears in [`llmfit-tui/src/display.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/display.rs) for user-facing throughput reports.