Understanding llmfit Run Modes: 5 Execution Strategies for LLM Inference

llmfit classifies LLM execution into five distinct RunMode variants—Gpu, MoeOffload, CpuOffload, CpuOnly, and TensorParallel—that determine how model weights are distributed across GPU VRAM, system RAM, and multi-node clusters.

The open-source llmfit library analyzes your hardware capabilities against large language model requirements to recommend optimal execution strategies. At the heart of this analysis is the RunMode enum defined in llmfit-core/src/fit.rs, which categorizes every feasible inference path based on memory topology and compute availability. Understanding these run modes helps you predict performance characteristics and select the appropriate execution strategy for your specific hardware constraints.

The Five RunMode Variants in llmfit-core

The RunMode enum lives in [llmfit-core/src/fit.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs#L88-L94) and operates independently of the FitLevel classification. This separation allows llmfit to evaluate both memory fit and execution path separately. Each variant represents a distinct hardware utilization strategy:

Gpu Mode (Full VRAM Execution)

Gpu mode loads all model weights directly into GPU VRAM, providing the fastest inference path available. This mode requires sufficient VRAM capacity to hold the entire model and is the default target when hardware permits. In llmfit-core/src/fit.rs, this variant represents pure GPU acceleration without system RAM fallback.

MoeOffload Mode (Mixture-of-Experts Optimization)

MoeOffload targets Mixture-of-Experts (MoE) architectures by keeping active expert layers in VRAM while offloading inactive experts to system RAM. This specialized mode, defined alongside other variants in lines 88-94 of fit.rs, optimizes VRAM usage for sparse activation patterns typical of modern MoE models like Mixtral.

CpuOffload Mode (Hybrid Execution)

CpuOffload implements a hybrid approach where a portion of the model resides in GPU VRAM while the remainder spills into system RAM. This mode balances speed against memory constraints, allowing larger models to run on GPUs with limited VRAM by paging weights between devices as needed.

CpuOnly Mode (System RAM Inference)

CpuOnly executes the model entirely within system RAM, bypassing GPU acceleration completely. As the slowest but most compatible path, this mode serves as the fallback when no suitable GPU memory is available. The ModelFit struct tracks this via the memory_required_gb field defined in [llmfit-core/src/fit.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs#L40-L44).

TensorParallel Mode (Multi-Node Distribution)

TensorParallel distributes model weights across multiple nodes via NCCL, enabling large-scale parallel inference for models exceeding single-node memory capacity. This mode targets cluster environments and represents the only distributed execution path in the enum.

Integrating RunMode with Model Analysis

The RunMode classification integrates with the ModelFit struct to provide complete execution context. When llmfit evaluates a model, it populates the run_mode field alongside memory_required_gb and fit_level to create a comprehensive execution plan.

Below is an example of constructing a ModelFit with a specific RunMode in Rust:

use llmfit_core::{
    fit::{RunMode, ModelFit, FitLevel, ScoreComponents},
    models::LlmModel,
};

let model = LlmModel::mock(); // placeholder for a real model
let fit = ModelFit {
    model,
    fit_level: FitLevel::Perfect,
    run_mode: RunMode::Gpu,
    memory_required_gb: 12.0,
    memory_available_gb: 16.0,
    utilization_pct: 75.0,
    notes: vec![],
    moe_offloaded_gb: None,
    score: 92.3,
    score_components: ScoreComponents {
        quality: 95.0,
        speed: 90.0,
        fit: 80.0,
        context: 85.0,
    },
    estimated_tps: 1200.0,
    best_quant: "q4_0".into(),
    use_case: Default::default(),
    runtime: Default::default(),
};

Generating Runtime Commands from RunMode

The selected RunMode directly influences the command-line flags generated for underlying runtimes like llama.cpp. In [llmfit-tui/src/display.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/display.rs#L71-L99), llmfit translates the run mode into specific arguments:

let cmd = llmfit_tui::display::generate_llamacpp_command(&fit);
println!("{}", cmd.unwrap());
// -> "llama-cli -hf some/repo:q4_0 -ngl all -c 2048"

Additionally, the HTTP API serializes RunMode for remote clients. As implemented in [llmfit-tui/src/serve_shared.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/serve_shared.rs), you can query the execution mode via the REST endpoint:

curl http://localhost:3000/api/v1/models/llama-2-7b | jq '.run_mode'

# "gpu"

Summary

  • llmfit defines five execution strategies in the RunMode enum: Gpu, MoeOffload, CpuOffload, CpuOnly, and TensorParallel.
  • The enum resides in llmfit-core/src/fit.rs (lines 88-94) and operates independently of memory fit levels.
  • Gpu provides fastest single-device inference, while TensorParallel enables multi-node cluster execution.
  • MoeOffload specifically optimizes Mixture-of-Experts models by selectively paging expert layers.
  • CpuOffload and CpuOnly provide fallback paths for limited VRAM scenarios, using hybrid or pure CPU memory.
  • The ModelFit struct in llmfit-core/src/fit.rs (lines 40-44) tracks the selected mode alongside memory requirements and performance scores.
  • Command generation in llmfit-tui/src/display.rs and API serialization in serve_shared.rs consume these classifications to configure actual inference runtimes.

Frequently Asked Questions

What is the fastest run mode available in llmfit?

Gpu mode provides the fastest execution path by loading all model weights into GPU VRAM, eliminating the latency associated with CPU-GPU memory transfers. According to the source code in llmfit-core/src/fit.rs, this mode requires sufficient VRAM capacity but delivers maximum throughput for single-device inference.

How does llmfit handle Mixture-of-Experts models differently from dense models?

For MoE architectures, llmfit uses the MoeOffload run mode to keep active experts in VRAM while moving inactive experts to system RAM. This approach, defined in the RunMode enum, recognizes the sparse activation patterns of MoE models and optimizes memory utilization accordingly, unlike dense models which use standard Gpu or CpuOffload modes.

Can I inspect the selected run mode programmatically via the llmfit API?

Yes, the HTTP API exposes the run_mode field through endpoints defined in llmfit-tui/src/serve_shared.rs. You can query any analyzed model to receive a JSON response containing the selected execution strategy, as shown in the curl example targeting localhost:3000/api/v1/models/{model-id}.

Where does llmfit map run modes to user interface labels?

The desktop application maps RunMode variants to user-friendly display strings in [llmfit-desktop/src/main.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-desktop/src/main.rs). Meanwhile, the planning module in [llmfit-core/src/plan.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs) relates planning paths (PlanRunPath) to their corresponding RunMode classifications for the internal decision engine.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →