How the MoE Expert Offloading Path Functions in LLMFIT

The MoE expert offloading path activates when a Mixture-of-Experts model cannot fit entirely in GPU VRAM, keeping only active experts in GPU memory while streaming inactive experts from system RAM to enable inference on memory-constrained hardware.

LLMFIT is an open-source model fitting framework that automatically determines optimal execution strategies for large language models. For Mixture-of-Experts (MoE) architectures, the framework implements a specialized MoE expert offloading path that analyzes hardware constraints and model characteristics to enable execution of models larger than available VRAM. This execution mode is defined and managed within the core fitting logic located in llmfit-core/src/fit.rs.

How the MoE Expert Offloading Path Works

The offloading mechanism operates through a sophisticated memory analysis pipeline that evaluates whether partial GPU residency is feasible for MoE architectures.

Detection in ModelFit::analyze_inner

The selection logic begins inside ModelFit::analyze_inner in llmfit-core/src/fit.rs. The code first attempts a pure-GPU execution path. When the model is identified as an MoE architecture via model.is_moe and the full model exceeds available VRAM (min_vram > system_vram), LLMFIT falls back to the moe_offload_path function. This fallback occurs between lines 44-62 of the fitting module, ensuring the system only considers offloading when full GPU residency is impossible.

Memory Calculation with moe_memory_for_quant

The function moe_memory_for_quant (lines 6-12) derives precise memory requirements for a given quantisation level. It computes two critical values:

  • moe_vram: The GPU memory required for active experts, calculated as (active_parameters × bits_per_parameter / 8 / GiB) × 1.1 to account for activation overhead
  • offloaded_gb: The system RAM required for inactive experts, calculated as inactive_parameters × bits_per_parameter / 8 / GiB

The moe_offload_path Implementation

The moe_offload_path function (lines 23-60) iterates through the quantisation hierarchy—checking QUANT_HIERARCHY, MLX_QUANT_HIERARCHY, or ONNX_QUANT_HIERARCHY—to find a configuration where active experts fit in VRAM while inactive experts fit in system RAM. When both constraints satisfy moe_vram ≤ system_vram and offloaded_gb ≤ system_ram, the function returns RunMode::MoeOffload along with the computed memory figures.

When LLMFIT Chooses the MoE Expert Offloading Path

LLMFIT selects this specialized path only when specific hardware and model conditions converge. The decision matrix evaluates four primary constraints:

  • MoE Architecture Detection: The model must report is_moe == true, indicating multiple expert layers exist
  • VRAM Constraints: The full model must exceed available GPU memory (min_vram > system_vram), forcing a non-standard execution path
  • Memory Pool Feasibility: At least one quantisation level must yield moe_vram fitting in VRAM while offloaded_gb fits in system RAM
  • Non-Unified Memory: The system must not use unified memory architecture (Apple Silicon), as offloading requires distinct RAM and VRAM pools for efficient streaming
  • Bandwidth Availability: DDR bandwidth must be known (default ≈50 GB/s or user-provided via LLMFIT_DDR_BANDWIDTH) to estimate throughput for the expert stream

In practice, this path activates for large MoE models like Qwen3-Next-80B on machines with sufficient VRAM for active experts but insufficient space for the full parameter set.

Implementation Details in llmfit-core/src/fit.rs

The MoE expert offloading implementation spans multiple components within the core fitting module.

RunMode Enumeration and MoeOffload

The RunMode enum (lines 88-94) defines MoeOffload as a distinct execution variant alongside Gpu, CpuOffload, and other modes. This enumeration value triggers specific handling throughout the codebase, including UI rendering in llmfit-tui/src/display.rs where it displays the "MoE" icon.

Validation and Scoring

The score_fit function (lines 65-73) grades the MoE offloading configuration as Good when mem_available >= mem_required * 1.2, ensuring at least 20% memory headroom. Configurations falling below this threshold receive a Marginal rating. The scoring logic considers total memory available across both VRAM and RAM pools, ensuring the system maintains performance margins during variable-length sequences.

Fallback Hierarchy

If no quantisation level satisfies the MoE constraints, LLMFIT cascades to RunMode::CpuOffload (when the full model fits in RAM) or RunMode::Gpu (when VRAM unexpectedly accommodates the full model), ensuring robust execution across hardware configurations.

Practical Implementation Example

The MoE expert offloading path activates automatically during model analysis. Below is a complete Rust example demonstrating how to trigger and inspect this execution mode:

use llmfit_core::fit::{ModelFit, RunMode};
use llmfit_core::hardware::SystemSpecs;
use llmfit_core::models::LlmModel;

// Load a MoE model (e.g., Qwen3-Next-80B)
let model = LlmModel::from_hf_repo("qwen/qwen3-next-80b")?;
let system = SystemSpecs::detect()?;

// Analyze automatically selects MoE offloading when appropriate
let fit = ModelFit::analyze(&model, &system);

println!("Run mode chosen: {}", fit.run_mode_text());
assert_eq!(fit.run_mode, RunMode::MoeOffload);

// Display detailed memory allocation notes
for note in &fit.notes {
    println!("• {}", note);
}

Sample output:


Run mode chosen: MoE
• MoE: 12/24 experts active in VRAM (3.5 GB) at q4_0
• Inactive experts offloaded to system RAM (5.2 GB)
• Baseline estimated speed: 15.4 tok/s

From the command line, invoke the fitting analysis using:

llmfit fit --model qwen/qwen3-next-80b

The resulting table reports "MoE" in the Mode column when the offloading path is selected.

Summary

  • The MoE expert offloading path activates when model.is_moe is true and full model VRAM residency is impossible, as determined in ModelFit::analyze_inner
  • Memory calculations in moe_memory_for_quant distinguish between active expert VRAM (with 1.1x overhead) and inactive expert RAM requirements
  • The path requires distinct memory pools, excluding unified-memory Apple Silicon devices from using this specific execution mode
  • Scoring logic in score_fit demands 20% memory headroom (mem_available >= mem_required * 1.2) to achieve a "Good" fit rating
  • Implementation resides primarily in llmfit-core/src/fit.rs, with UI representation in llmfit-tui/src/display.rs

Frequently Asked Questions

What triggers the MoE expert offloading path in LLMFIT?

The path triggers when a model reports itself as an MoE architecture via is_moe == true and the full parameter set exceeds available GPU VRAM. Specifically, when min_vram > system_vram in ModelFit::analyze_inner, LLMFIT invokes moe_offload_path to evaluate whether active experts can reside in VRAM while inactive experts stream from system RAM.

How does LLMFIT calculate memory requirements for MoE offloading?

LLMFIT uses moe_memory_for_quant to compute two values: moe_vram for active experts (including a 1.1x activation overhead multiplier) and offloaded_gb for inactive experts. The function iterates through available quantisation hierarchies to find a level where active parameters fit in GPU memory and inactive parameters fit in system RAM.

Why won't LLMFIT use MoE offloading on Apple Silicon devices?

Apple Silicon uses unified memory architecture where RAM and VRAM share the same physical pool. The MoE expert offloading path requires distinct memory pools to stream inactive experts from system RAM to GPU VRAM. On unified memory systems, LLMFIT uses the standard RunMode::Gpu path instead, as the offloading strategy provides no benefit when memory is not physically segregated.

What happens if no quantisation level fits the MoE constraints?

If moe_offload_path cannot find a quantisation level where both active experts fit in VRAM and inactive experts fit in RAM, LLMFIT falls back to RunMode::CpuOffload (if the full model fits in system RAM) or attempts RunMode::Gpu if VRAM constraints unexpectedly relax. This ensures the model executes regardless of offloading feasibility.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →