How llmfit Handles Mixture of Experts (MoE) Models: A Bandwidth-Aware Approach

llmfit treats a Mixture of Experts (MoE) model as a special LLM class that can split its weights between GPU VRAM and system RAM, using a dedicated RunMode::MoeOffload execution path that keeps active experts in VRAM while streaming inactive experts from memory.

The AlexsJones/llmfit repository is an open-source Rust application that scores and plans LLM inference across local hardware. Its MoE handling is a standout feature: instead of treating MoE models like dense models that must fully reside in VRAM, llmfit performs bandwidth-aware memory budgeting, throughput estimation, and user-visible labeling. This article examines the exact implementation details and code paths, showing you how MoE handling works inside llmfit-core, the TUI, and the HTTP API.

The Three Pillars of MoE Handling in llmfit

llmfit's Mixture of Experts (MoE) support rests on three core concepts, each implemented in distinct modules:

Concept What it does Where it lives
RunMode::MoeOffload — A dedicated execution path that loads active experts into VRAM while inactive experts remain in system RAM. Defined in fit.rs and used throughout runtime-selection logic. llmfit-core/src/fit.rs#L188-L194
Model metadata — is_moe, num_experts, active_experts, active_parameters, plus bandwidth-related fields (hidden_size, moe_intermediate_size, shared_expert_intermediate_size). Populated from the Hugging Face catalog and scraper; drives memory- and bandwidth-based calculations. llmfit-core/src/models.rs#L100-L110
MoE-aware fit and speed estimation — The fit engine and planner compute how much RAM can be off-loaded, incorporate DDR bandwidth limits, and apply a dedicated run-mode factor (moe_offload). Implemented in fit.rs (memory calculation, notes, speed factor) and plan.rs (bandwidth-aware TPS estimate). llmfit-core/src/fit.rs#L527-L590 and llmfit-core/src/plan.rs#L90-L102

Detecting a MoE Model

When the model catalog is loaded into models.rs, the boolean is_moe is set whenever the entry matches an MoE signature—such as "8x7B" or "30b-a3b". The struct also captures the total number of experts and how many are active per forward pass:

pub struct LlmModel {
    // …
    pub is_moe: bool,
    pub num_experts: Option<u32>,
    pub active_experts: Option<u32>,
    pub active_parameters: Option<u64>,
    // …
}

Source: [models.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs)

This metadata is the foundation for everything that follows—run-mode selection, memory budgeting, and throughput estimation all read these fields.

Selecting the Execution Path

During fit calculation, fit.rs::analyze inspects the GPU/VRAM state alongside the model metadata. If the model is MoE and the total VRAM cannot simultaneously hold all experts, the engine falls back to RunMode::MoeOffload.

if model.is_moe && min_vram <= system_vram {
    // All experts fit → normal GPU mode
    (RunMode::Gpu, min_vram, system_vram)
} else if model.is_moe {
    // Active experts in VRAM, inactive experts off-loaded to RAM
    moe_offload_path(...)
}

Source: [fit.rs lines 444-56](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs)

When the MoeOffload path is chosen, the moe_offload_path function computes the RAM required for off-loaded experts via model.moe_offloaded_ram_gb() and annotates the reason as "MoE Offload" so users understand why the engine chose this mode.

Memory Estimation for MoE Offload

Inside fit.rs, memory computations account for the off-loaded expert weights separately from active expert weights:

let moe_offloaded_gb = if run_mode == RunMode::MoeOffload {
    model.moe_offloaded_ram_gb()
} else {
    None
};

The resulting numbers feed the fit level calculation (Perfect, Good, etc.) and report both required VRAM and required RAM to the user. This split is exactly what makes llmfit's MoE scoring accurate—dense models use only VRAM, while MoE models can deliberately trade VRAM for RAM.

Speed and Bandwidth Estimation

MoE offload introduces a fundamental bandwidth bottleneck: the CPU must stream inactive experts from RAM for every token. The planner in plan.rs recognizes this explicitly by forcing the speed run mode to RunMode::MoeOffload for any MoE model evaluated on the CPU-offload path:

fn speed_run_mode(path: PlanRunPath, model: &LlmModel) -> RunMode {
    match path {
        PlanRunPath::CpuOffload if model.is_moe => RunMode::MoeOffload,
        other => other.run_mode(),
    }
}

The actual tokens-per-second (TPS) is then computed by the shared estimator fit::estimate_tps, which applies a DDR-bandwidth roofline for MoE offload. Because both plan.rs and fit.rs share this estimator, the fit score and the plan report will never contradict each other.

Configuring the MoE Speed Factor

Users can tune the MoE-offload speed factor through the CalcConfig struct. The default is 0.8, exposed in the TUI's "Advanced Configuration" panel:

pub struct RunModeFactors {
    pub gpu: f64,
    pub tensor_parallel: f64,
    pub moe_offload: f64,
    pub cpu_offload: f64,
    pub cpu_only: f64,
}

Source: [fit.rs lines 54-71](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs)

This factor lets you calibrate how heavily the bandwidth penalty of RAM streaming should be weighted when comparing MoE offload against GPU-only or CPU-only modes.

UI and API Presentation

MoE models are visually distinguished wherever they appear—CLI, TUI, or API. In the CLI you'll see a tag like MoE: 3/10 experts active, and in the TUI a dedicated "MoE Offload:" line appears in the model details panel. The JSON API mirrors this via the is_moe boolean and the expert counts.

The TUI rendering lives in [tui_ui.rs lines 1322-1324](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_ui.rs), and API serialization in [serve_shared.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/serve_shared.rs).

Practical Examples

Listing MoE models from the CLI

cargo run -- --cli list --moe

Output example:


Model                     Provider   Params   VRAM   RAM   MoE
----------------------------------------------------------------
qwen1.5-moe-7b            huggingface 7B     2.1GB  6.4GB  8/10 experts active
mixtral-8x7b-instruct-v0.1 huggingface 46.7B  13.5GB 25.3GB 8/8 experts active

Using the Rust library directly

use llmfit_core::{
    hardware::SystemSpecs,
    models::LlmModel,
    fit::{self, RunMode},
    plan::PlanRunPath,
};

fn demo_moe_handling() {
    let model: LlmModel = /* ... */;

    let mut sys = SystemSpecs::default();
    sys.gpu_vram_gb = Some(16.0);
    sys.available_ram_gb = 32.0;
    sys.has_gpu = true;

    let (run_mode, mem_req, mem_avail) =
        fit::analyze(&model, &sys, fit::InferenceRuntime::LlamaCpp, None);

    println!(
        "Chosen run mode: {:?} (needs {:.1} GB VRAM, {:.1} GB RAM)",
        run_mode, mem_req, mem_avail
    );
}

When model.is_moe is true and the VRAM is insufficient for all experts, this code returns RunMode::MoeOffload.

Querying via the HTTP API

curl http://localhost:8080/api/v1/models | jq '.[] | select(.is_moe) | {name, num_experts, active_experts}'

Sample response:

{
  "name": "mixtral-8x7b-instruct-v0.1",
  "num_experts": 8,
  "active_experts": 8
}

Summary

  • Detection — Model metadata flags is_moe and records expert counts and active parameters.
  • Run-mode selection — If a MoE model can't fully fit in VRAM, llmfit switches to RunMode::MoeOffload.
  • Memory budgeting — Inactive experts are budgeted as RAM (moe_offloaded_ram_gb) alongside the VRAM requirement.
  • Bandwidth-aware speed estimation — TPS calculations use a DDR-bandwidth roofline and a dedicated speed factor.
  • User control — The MoE-offload speed factor (default 0.8) is tunable in advanced UI settings.
  • Presentation — CLI, TUI, and API all explicitly label MoE models and show active-vs-total expert counts.

Frequently Asked Questions

What happens when an MoE model fits entirely in VRAM?

When all experts fit in VRAM, llmfit uses the normal RunMode::Gpu path. The MoeOffload run mode is reserved for scenarios where the total expert weights override the available VRAM.

How does llmfit estimate throughput for an MoE model with offloaded experts?

It applies a DDR-bandwidth roofline inside fit::estimate_tps, shared by both fit.rs and plan.rs. This ensures that the reported TPS correctly factors in the memory bandwidth penalty from streaming your inactive experts from RAM each token.

Can I tune how optimistic the MoE-offload speed estimate is?

Yes. The CalcConfig struct includes a run_mode_factors.moe_offload field (default 0.8) that lets you scale the computed TPS up or down. It's exposed in the TUI under "Advanced Configuration," and you can also set it programmatically via the Rust API.

Where is MoE detection implemented in the codebase?

Detection happens in llmfit-core/src/models.rs, where patterns like "8x7B" set the is_moe flag. The subsequent decision logic lives in llmfit-core/src/fit.rs (RunMode::MoeOffload), with speed estimation in llmfit-core/src/plan.rs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →