# How the MoE Expert Offloading Path Functions in LLMFIT

> Learn how LLMFITs MoE expert offloading path streams inactive experts from RAM to GPU to enable inference on memory-constrained hardware. Optimize your LLM usage.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: internals
- Published: 2026-08-21

---

**The MoE expert offloading path activates when a Mixture-of-Experts model cannot fit entirely in GPU VRAM, keeping only active experts in GPU memory while streaming inactive experts from system RAM to enable inference on memory-constrained hardware.**

LLMFIT is an open-source model fitting framework that automatically determines optimal execution strategies for large language models. For Mixture-of-Experts (MoE) architectures, the framework implements a specialized **MoE expert offloading path** that analyzes hardware constraints and model characteristics to enable execution of models larger than available VRAM. This execution mode is defined and managed within the core fitting logic located in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs).

## How the MoE Expert Offloading Path Works

The offloading mechanism operates through a sophisticated memory analysis pipeline that evaluates whether partial GPU residency is feasible for MoE architectures.

### Detection in `ModelFit::analyze_inner`

The selection logic begins inside `ModelFit::analyze_inner` in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs). The code first attempts a pure-GPU execution path. When the model is identified as an MoE architecture via `model.is_moe` and the full model exceeds available VRAM (`min_vram > system_vram`), LLMFIT falls back to the `moe_offload_path` function. This fallback occurs between lines 44-62 of the fitting module, ensuring the system only considers offloading when full GPU residency is impossible.

### Memory Calculation with `moe_memory_for_quant`

The function `moe_memory_for_quant` (lines 6-12) derives precise memory requirements for a given quantisation level. It computes two critical values:

- **`moe_vram`**: The GPU memory required for active experts, calculated as `(active_parameters × bits_per_parameter / 8 / GiB) × 1.1` to account for activation overhead
- **`offloaded_gb`**: The system RAM required for inactive experts, calculated as `inactive_parameters × bits_per_parameter / 8 / GiB`

### The `moe_offload_path` Implementation

The `moe_offload_path` function (lines 23-60) iterates through the quantisation hierarchy—checking `QUANT_HIERARCHY`, `MLX_QUANT_HIERARCHY`, or `ONNX_QUANT_HIERARCHY`—to find a configuration where active experts fit in VRAM while inactive experts fit in system RAM. When both constraints satisfy `moe_vram ≤ system_vram` and `offloaded_gb ≤ system_ram`, the function returns `RunMode::MoeOffload` along with the computed memory figures.

## When LLMFIT Chooses the MoE Expert Offloading Path

LLMFIT selects this specialized path only when specific hardware and model conditions converge. The decision matrix evaluates four primary constraints:

- **MoE Architecture Detection**: The model must report `is_moe == true`, indicating multiple expert layers exist
- **VRAM Constraints**: The full model must exceed available GPU memory (`min_vram > system_vram`), forcing a non-standard execution path
- **Memory Pool Feasibility**: At least one quantisation level must yield `moe_vram` fitting in VRAM while `offloaded_gb` fits in system RAM
- **Non-Unified Memory**: The system must not use unified memory architecture (Apple Silicon), as offloading requires distinct RAM and VRAM pools for efficient streaming
- **Bandwidth Availability**: DDR bandwidth must be known (default ≈50 GB/s or user-provided via `LLMFIT_DDR_BANDWIDTH`) to estimate throughput for the expert stream

In practice, this path activates for large MoE models like Qwen3-Next-80B on machines with sufficient VRAM for active experts but insufficient space for the full parameter set.

## Implementation Details in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs)

The MoE expert offloading implementation spans multiple components within the core fitting module.

### RunMode Enumeration and `MoeOffload`

The `RunMode` enum (lines 88-94) defines `MoeOffload` as a distinct execution variant alongside `Gpu`, `CpuOffload`, and other modes. This enumeration value triggers specific handling throughout the codebase, including UI rendering in [`llmfit-tui/src/display.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/display.rs) where it displays the "MoE" icon.

### Validation and Scoring

The `score_fit` function (lines 65-73) grades the MoE offloading configuration as **Good** when `mem_available >= mem_required * 1.2`, ensuring at least 20% memory headroom. Configurations falling below this threshold receive a **Marginal** rating. The scoring logic considers total memory available across both VRAM and RAM pools, ensuring the system maintains performance margins during variable-length sequences.

### Fallback Hierarchy

If no quantisation level satisfies the MoE constraints, LLMFIT cascades to `RunMode::CpuOffload` (when the full model fits in RAM) or `RunMode::Gpu` (when VRAM unexpectedly accommodates the full model), ensuring robust execution across hardware configurations.

## Practical Implementation Example

The MoE expert offloading path activates automatically during model analysis. Below is a complete Rust example demonstrating how to trigger and inspect this execution mode:

```rust
use llmfit_core::fit::{ModelFit, RunMode};
use llmfit_core::hardware::SystemSpecs;
use llmfit_core::models::LlmModel;

// Load a MoE model (e.g., Qwen3-Next-80B)
let model = LlmModel::from_hf_repo("qwen/qwen3-next-80b")?;
let system = SystemSpecs::detect()?;

// Analyze automatically selects MoE offloading when appropriate
let fit = ModelFit::analyze(&model, &system);

println!("Run mode chosen: {}", fit.run_mode_text());
assert_eq!(fit.run_mode, RunMode::MoeOffload);

// Display detailed memory allocation notes
for note in &fit.notes {
    println!("• {}", note);
}

```

**Sample output:**

```

Run mode chosen: MoE
• MoE: 12/24 experts active in VRAM (3.5 GB) at q4_0
• Inactive experts offloaded to system RAM (5.2 GB)
• Baseline estimated speed: 15.4 tok/s

```

From the command line, invoke the fitting analysis using:

```bash
llmfit fit --model qwen/qwen3-next-80b

```

The resulting table reports "MoE" in the Mode column when the offloading path is selected.

## Summary

- The **MoE expert offloading path** activates when `model.is_moe` is true and full model VRAM residency is impossible, as determined in `ModelFit::analyze_inner`
- **Memory calculations** in `moe_memory_for_quant` distinguish between active expert VRAM (with 1.1x overhead) and inactive expert RAM requirements
- The path requires **distinct memory pools**, excluding unified-memory Apple Silicon devices from using this specific execution mode
- **Scoring logic** in `score_fit` demands 20% memory headroom (`mem_available >= mem_required * 1.2`) to achieve a "Good" fit rating
- Implementation resides primarily in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs), with UI representation in [`llmfit-tui/src/display.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/display.rs)

## Frequently Asked Questions

### What triggers the MoE expert offloading path in LLMFIT?

The path triggers when a model reports itself as an MoE architecture via `is_moe == true` and the full parameter set exceeds available GPU VRAM. Specifically, when `min_vram > system_vram` in `ModelFit::analyze_inner`, LLMFIT invokes `moe_offload_path` to evaluate whether active experts can reside in VRAM while inactive experts stream from system RAM.

### How does LLMFIT calculate memory requirements for MoE offloading?

LLMFIT uses `moe_memory_for_quant` to compute two values: `moe_vram` for active experts (including a 1.1x activation overhead multiplier) and `offloaded_gb` for inactive experts. The function iterates through available quantisation hierarchies to find a level where active parameters fit in GPU memory and inactive parameters fit in system RAM.

### Why won't LLMFIT use MoE offloading on Apple Silicon devices?

Apple Silicon uses unified memory architecture where RAM and VRAM share the same physical pool. The MoE expert offloading path requires distinct memory pools to stream inactive experts from system RAM to GPU VRAM. On unified memory systems, LLMFIT uses the standard `RunMode::Gpu` path instead, as the offloading strategy provides no benefit when memory is not physically segregated.

### What happens if no quantisation level fits the MoE constraints?

If `moe_offload_path` cannot find a quantisation level where both active experts fit in VRAM and inactive experts fit in RAM, LLMFIT falls back to `RunMode::CpuOffload` (if the full model fits in system RAM) or attempts `RunMode::Gpu` if VRAM constraints unexpectedly relax. This ensures the model executes regardless of offloading feasibility.