# How llmfit Handles Mixture of Experts (MoE) Models: A Bandwidth-Aware Approach

> Discover how llmfit manages Mixture of Experts MoE models with its bandwidth-aware approach, keeping active experts in VRAM and streaming inactive ones from system RAM.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: deep-dive
- Published: 2026-08-23

---

**llmfit treats a Mixture of Experts (MoE) model as a special LLM class that can split its weights between GPU VRAM and system RAM, using a dedicated `RunMode::MoeOffload` execution path that keeps active experts in VRAM while streaming inactive experts from memory.**

The [AlexsJones/llmfit](https://github.com/AlexsJones/llmfit) repository is an open-source Rust application that scores and plans LLM inference across local hardware. Its MoE handling is a standout feature: instead of treating MoE models like dense models that must fully reside in VRAM, llmfit performs bandwidth-aware memory budgeting, throughput estimation, and user-visible labeling. This article examines the exact implementation details and code paths, showing you how MoE handling works inside `llmfit-core`, the TUI, and the HTTP API.

## The Three Pillars of MoE Handling in llmfit

llmfit's Mixture of Experts (MoE) support rests on three core concepts, each implemented in distinct modules:

| Concept | What it does | Where it lives |
|---|---|---|
| **RunMode::MoeOffload** — A dedicated execution path that loads active experts into VRAM while inactive experts remain in system RAM. | Defined in [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) and used throughout runtime-selection logic. | [`llmfit-core/src/fit.rs#L188-L194`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) |
| **Model metadata** — `is_moe`, `num_experts`, `active_experts`, `active_parameters`, plus bandwidth-related fields (`hidden_size`, `moe_intermediate_size`, `shared_expert_intermediate_size`). | Populated from the Hugging Face catalog and scraper; drives memory- and bandwidth-based calculations. | [`llmfit-core/src/models.rs#L100-L110`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs) |
| **MoE-aware fit and speed estimation** — The fit engine and planner compute how much RAM can be off-loaded, incorporate DDR bandwidth limits, and apply a dedicated run-mode factor (`moe_offload`). | Implemented in [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) (memory calculation, notes, speed factor) and [`plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/plan.rs) (bandwidth-aware TPS estimate). | [`llmfit-core/src/fit.rs#L527-L590`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) and [`llmfit-core/src/plan.rs#L90-L102`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs) |

## Detecting a MoE Model

When the model catalog is loaded into [`models.rs`](https://github.com/AlexsJones/llmfit/blob/main/models.rs), the boolean `is_moe` is set whenever the entry matches an MoE signature—such as "8x7B" or "30b-a3b". The struct also captures the total number of experts and how many are active per forward pass:

```rust
pub struct LlmModel {
    // …
    pub is_moe: bool,
    pub num_experts: Option<u32>,
    pub active_experts: Option<u32>,
    pub active_parameters: Option<u64>,
    // …
}

```

Source: [[`models.rs`](https://github.com/AlexsJones/llmfit/blob/main/models.rs)](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs)

This metadata is the foundation for everything that follows—run-mode selection, memory budgeting, and throughput estimation all read these fields.

## Selecting the Execution Path

During fit calculation, `fit.rs::analyze` inspects the GPU/VRAM state alongside the model metadata. If the model is MoE and the total VRAM cannot simultaneously hold all experts, the engine falls back to `RunMode::MoeOffload`.

```rust
if model.is_moe && min_vram <= system_vram {
    // All experts fit → normal GPU mode
    (RunMode::Gpu, min_vram, system_vram)
} else if model.is_moe {
    // Active experts in VRAM, inactive experts off-loaded to RAM
    moe_offload_path(...)
}

```

Source: [[`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) lines 444-56](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs)

When the `MoeOffload` path is chosen, the `moe_offload_path` function computes the RAM required for off-loaded experts via `model.moe_offloaded_ram_gb()` and annotates the reason as "MoE Offload" so users understand why the engine chose this mode.

## Memory Estimation for MoE Offload

Inside [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs), memory computations account for the off-loaded expert weights separately from active expert weights:

```rust
let moe_offloaded_gb = if run_mode == RunMode::MoeOffload {
    model.moe_offloaded_ram_gb()
} else {
    None
};

```

The resulting numbers feed the fit level calculation (`Perfect`, `Good`, etc.) and report both required VRAM and required RAM to the user. This split is exactly what makes llmfit's MoE scoring accurate—dense models use only VRAM, while MoE models can deliberately trade VRAM for RAM.

## Speed and Bandwidth Estimation

MoE offload introduces a fundamental bandwidth bottleneck: the CPU must stream inactive experts from RAM for every token. The planner in [`plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/plan.rs) recognizes this explicitly by forcing the speed run mode to `RunMode::MoeOffload` for any MoE model evaluated on the CPU-offload path:

```rust
fn speed_run_mode(path: PlanRunPath, model: &LlmModel) -> RunMode {
    match path {
        PlanRunPath::CpuOffload if model.is_moe => RunMode::MoeOffload,
        other => other.run_mode(),
    }
}

```

The actual tokens-per-second (TPS) is then computed by the shared estimator `fit::estimate_tps`, which applies a DDR-bandwidth roofline for MoE offload. Because both [`plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/plan.rs) and [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) share this estimator, the fit score and the plan report will never contradict each other.

## Configuring the MoE Speed Factor

Users can tune the MoE-offload speed factor through the `CalcConfig` struct. The default is **0.8**, exposed in the TUI's "Advanced Configuration" panel:

```rust
pub struct RunModeFactors {
    pub gpu: f64,
    pub tensor_parallel: f64,
    pub moe_offload: f64,
    pub cpu_offload: f64,
    pub cpu_only: f64,
}

```

Source: [[`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) lines 54-71](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs)

This factor lets you calibrate how heavily the bandwidth penalty of RAM streaming should be weighted when comparing MoE offload against GPU-only or CPU-only modes.

## UI and API Presentation

MoE models are visually distinguished wherever they appear—CLI, TUI, or API. In the CLI you'll see a tag like `MoE: 3/10 experts active`, and in the TUI a dedicated "MoE Offload:" line appears in the model details panel. The JSON API mirrors this via the `is_moe` boolean and the expert counts.

The TUI rendering lives in [[`tui_ui.rs`](https://github.com/AlexsJones/llmfit/blob/main/tui_ui.rs) lines 1322-1324](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_ui.rs), and API serialization in [[`serve_shared.rs`](https://github.com/AlexsJones/llmfit/blob/main/serve_shared.rs)](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/serve_shared.rs).

## Practical Examples

### Listing MoE models from the CLI

```sh
cargo run -- --cli list --moe

```

Output example:

```

Model                     Provider   Params   VRAM   RAM   MoE
----------------------------------------------------------------
qwen1.5-moe-7b            huggingface 7B     2.1GB  6.4GB  8/10 experts active
mixtral-8x7b-instruct-v0.1 huggingface 46.7B  13.5GB 25.3GB 8/8 experts active

```

### Using the Rust library directly

```rust
use llmfit_core::{
    hardware::SystemSpecs,
    models::LlmModel,
    fit::{self, RunMode},
    plan::PlanRunPath,
};

fn demo_moe_handling() {
    let model: LlmModel = /* ... */;

    let mut sys = SystemSpecs::default();
    sys.gpu_vram_gb = Some(16.0);
    sys.available_ram_gb = 32.0;
    sys.has_gpu = true;

    let (run_mode, mem_req, mem_avail) =
        fit::analyze(&model, &sys, fit::InferenceRuntime::LlamaCpp, None);

    println!(
        "Chosen run mode: {:?} (needs {:.1} GB VRAM, {:.1} GB RAM)",
        run_mode, mem_req, mem_avail
    );
}

```

When `model.is_moe` is true and the VRAM is insufficient for all experts, this code returns `RunMode::MoeOffload`.

### Querying via the HTTP API

```bash
curl http://localhost:8080/api/v1/models | jq '.[] | select(.is_moe) | {name, num_experts, active_experts}'

```

Sample response:

```json
{
  "name": "mixtral-8x7b-instruct-v0.1",
  "num_experts": 8,
  "active_experts": 8
}

```

## Summary

- **Detection** — Model metadata flags `is_moe` and records expert counts and active parameters.
- **Run-mode selection** — If a MoE model can't fully fit in VRAM, llmfit switches to `RunMode::MoeOffload`.
- **Memory budgeting** — Inactive experts are budgeted as RAM (`moe_offloaded_ram_gb`) alongside the VRAM requirement.
- **Bandwidth-aware speed estimation** — TPS calculations use a DDR-bandwidth roofline and a dedicated speed factor.
- **User control** — The MoE-offload speed factor (default 0.8) is tunable in advanced UI settings.
- **Presentation** — CLI, TUI, and API all explicitly label MoE models and show active-vs-total expert counts.

## Frequently Asked Questions

### What happens when an MoE model fits entirely in VRAM?

When all experts fit in VRAM, llmfit uses the normal `RunMode::Gpu` path. The `MoeOffload` run mode is reserved for scenarios where the total expert weights override the available VRAM.

### How does llmfit estimate throughput for an MoE model with offloaded experts? 

It applies a DDR-bandwidth roofline inside `fit::estimate_tps`, shared by both [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) and [`plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/plan.rs). This ensures that the reported TPS correctly factors in the memory bandwidth penalty from streaming your inactive experts from RAM each token.

### Can I tune how optimistic the MoE-offload speed estimate is?

Yes. The `CalcConfig` struct includes a `run_mode_factors.moe_offload` field (default 0.8) that lets you scale the computed TPS up or down. It's exposed in the TUI under "Advanced Configuration," and you can also set it programmatically via the Rust API.

### Where is MoE detection implemented in the codebase?

Detection happens in [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs), where patterns like "8x7B" set the `is_moe` flag. The subsequent decision logic lives in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) (`RunMode::MoeOffload`), with speed estimation in [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs).