# Understanding llmfit Run Modes: 5 Execution Strategies for LLM Inference

> Explore the 5 llmfit Run Modes: Gpu, MoeOffload, CpuOffload, CpuOnly, and TensorParallel. Learn how to efficiently distribute LLM weights across GPU VRAM, system RAM, and clusters for optimal inference.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: deep-dive
- Published: 2026-08-22

---

**llmfit classifies LLM execution into five distinct RunMode variants—Gpu, MoeOffload, CpuOffload, CpuOnly, and TensorParallel—that determine how model weights are distributed across GPU VRAM, system RAM, and multi-node clusters.**

The open-source **llmfit** library analyzes your hardware capabilities against large language model requirements to recommend optimal execution strategies. At the heart of this analysis is the `RunMode` enum defined in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs), which categorizes every feasible inference path based on memory topology and compute availability. Understanding these run modes helps you predict performance characteristics and select the appropriate execution strategy for your specific hardware constraints.

## The Five RunMode Variants in llmfit-core

The `RunMode` enum lives in [[`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs)](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs#L88-L94) and operates independently of the **FitLevel** classification. This separation allows llmfit to evaluate both memory fit and execution path separately. Each variant represents a distinct hardware utilization strategy:

### Gpu Mode (Full VRAM Execution)

**Gpu** mode loads all model weights directly into GPU VRAM, providing the fastest inference path available. This mode requires sufficient VRAM capacity to hold the entire model and is the default target when hardware permits. In [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs), this variant represents pure GPU acceleration without system RAM fallback.

### MoeOffload Mode (Mixture-of-Experts Optimization)

**MoeOffload** targets Mixture-of-Experts (MoE) architectures by keeping active expert layers in VRAM while offloading inactive experts to system RAM. This specialized mode, defined alongside other variants in lines 88-94 of [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs), optimizes VRAM usage for sparse activation patterns typical of modern MoE models like Mixtral.

### CpuOffload Mode (Hybrid Execution)

**CpuOffload** implements a hybrid approach where a portion of the model resides in GPU VRAM while the remainder spills into system RAM. This mode balances speed against memory constraints, allowing larger models to run on GPUs with limited VRAM by paging weights between devices as needed.

### CpuOnly Mode (System RAM Inference)

**CpuOnly** executes the model entirely within system RAM, bypassing GPU acceleration completely. As the slowest but most compatible path, this mode serves as the fallback when no suitable GPU memory is available. The `ModelFit` struct tracks this via the `memory_required_gb` field defined in [[`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs)](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs#L40-L44).

### TensorParallel Mode (Multi-Node Distribution)

**TensorParallel** distributes model weights across multiple nodes via NCCL, enabling large-scale parallel inference for models exceeding single-node memory capacity. This mode targets cluster environments and represents the only distributed execution path in the enum.

## Integrating RunMode with Model Analysis

The `RunMode` classification integrates with the `ModelFit` struct to provide complete execution context. When llmfit evaluates a model, it populates the `run_mode` field alongside `memory_required_gb` and `fit_level` to create a comprehensive execution plan.

Below is an example of constructing a `ModelFit` with a specific `RunMode` in Rust:

```rust
use llmfit_core::{
    fit::{RunMode, ModelFit, FitLevel, ScoreComponents},
    models::LlmModel,
};

let model = LlmModel::mock(); // placeholder for a real model
let fit = ModelFit {
    model,
    fit_level: FitLevel::Perfect,
    run_mode: RunMode::Gpu,
    memory_required_gb: 12.0,
    memory_available_gb: 16.0,
    utilization_pct: 75.0,
    notes: vec![],
    moe_offloaded_gb: None,
    score: 92.3,
    score_components: ScoreComponents {
        quality: 95.0,
        speed: 90.0,
        fit: 80.0,
        context: 85.0,
    },
    estimated_tps: 1200.0,
    best_quant: "q4_0".into(),
    use_case: Default::default(),
    runtime: Default::default(),
};

```

## Generating Runtime Commands from RunMode

The selected `RunMode` directly influences the command-line flags generated for underlying runtimes like [`llama.cpp`](https://github.com/AlexsJones/llmfit/blob/main/llama.cpp). In [[`llmfit-tui/src/display.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/display.rs)](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/display.rs#L71-L99), llmfit translates the run mode into specific arguments:

```rust
let cmd = llmfit_tui::display::generate_llamacpp_command(&fit);
println!("{}", cmd.unwrap());
// -> "llama-cli -hf some/repo:q4_0 -ngl all -c 2048"

```

Additionally, the HTTP API serializes `RunMode` for remote clients. As implemented in [[`llmfit-tui/src/serve_shared.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/serve_shared.rs)](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/serve_shared.rs), you can query the execution mode via the REST endpoint:

```bash
curl http://localhost:3000/api/v1/models/llama-2-7b | jq '.run_mode'

# "gpu"

```

## Summary

- **llmfit** defines five execution strategies in the `RunMode` enum: **Gpu**, **MoeOffload**, **CpuOffload**, **CpuOnly**, and **TensorParallel**.
- The enum resides in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) (lines 88-94) and operates independently of memory fit levels.
- **Gpu** provides fastest single-device inference, while **TensorParallel** enables multi-node cluster execution.
- **MoeOffload** specifically optimizes Mixture-of-Experts models by selectively paging expert layers.
- **CpuOffload** and **CpuOnly** provide fallback paths for limited VRAM scenarios, using hybrid or pure CPU memory.
- The `ModelFit` struct in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) (lines 40-44) tracks the selected mode alongside memory requirements and performance scores.
- Command generation in [`llmfit-tui/src/display.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/display.rs) and API serialization in [`serve_shared.rs`](https://github.com/AlexsJones/llmfit/blob/main/serve_shared.rs) consume these classifications to configure actual inference runtimes.

## Frequently Asked Questions

### What is the fastest run mode available in llmfit?

**Gpu** mode provides the fastest execution path by loading all model weights into GPU VRAM, eliminating the latency associated with CPU-GPU memory transfers. According to the source code in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs), this mode requires sufficient VRAM capacity but delivers maximum throughput for single-device inference.

### How does llmfit handle Mixture-of-Experts models differently from dense models?

For MoE architectures, llmfit uses the **MoeOffload** run mode to keep active experts in VRAM while moving inactive experts to system RAM. This approach, defined in the `RunMode` enum, recognizes the sparse activation patterns of MoE models and optimizes memory utilization accordingly, unlike dense models which use standard **Gpu** or **CpuOffload** modes.

### Can I inspect the selected run mode programmatically via the llmfit API?

Yes, the HTTP API exposes the `run_mode` field through endpoints defined in [`llmfit-tui/src/serve_shared.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/serve_shared.rs). You can query any analyzed model to receive a JSON response containing the selected execution strategy, as shown in the `curl` example targeting `localhost:3000/api/v1/models/{model-id}`.

### Where does llmfit map run modes to user interface labels?

The desktop application maps `RunMode` variants to user-friendly display strings in [[`llmfit-desktop/src/main.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-desktop/src/main.rs)](https://github.com/AlexsJones/llmfit/blob/main/llmfit-desktop/src/main.rs). Meanwhile, the planning module in [[`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs)](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs) relates planning paths (`PlanRunPath`) to their corresponding `RunMode` classifications for the internal decision engine.