# RunMode Options in llmfit: 5 Execution Strategies for LLM Inference

> Explore llmfit RunMode options: Gpu, TensorParallel, MoeOffload, CpuOffload, and CpuOnly. Optimize LLM inference by distributing model weights across GPU VRAM, system RAM, and CPU.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: deep-dive
- Published: 2026-09-13

---

**The `RunMode` enum in llmfit defines five execution strategies—`Gpu`, `TensorParallel`, `MoeOffload`, `CpuOffload`, and `CpuOnly`—that determine how model weights and activations are distributed across GPU VRAM, system RAM, and CPU compute resources.**

The `RunMode` type is the central hardware abstraction in the llmfit ecosystem that dictates how large language models execute on available hardware. Defined in the core library at [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs), this enum orchestrates memory placement decisions ranging from single-GPU inference to distributed tensor parallelism, directly impacting throughput and latency characteristics according to the source code analysis.

## What Is RunMode in llmfit?

Located in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs), the `RunMode` enum encodes the execution strategy selected during the model fitting phase. Each variant represents a distinct hardware utilization pattern optimized for specific memory constraints and computational topologies.

The enum works alongside the `RunModeFactors` struct in the same file to adjust throughput estimates based on the selected strategy. During initialization, the library's hardware detection module ([`hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/hardware.rs)) scans available resources, then the planning logic in [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs) maps specific execution paths to concrete `RunMode` values via the `PlanRunPath` enum.

## The Five RunMode Variants Explained

### Gpu

The `Gpu` variant keeps all model weights and activations resident in GPU VRAM. This mode delivers maximum inference speed when GPU memory capacity exceeds model requirements. According to the source code in [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs), this is the preferred default when `min_vram_gb` constraints are satisfied by available hardware.

### TensorParallel

`TensorParallel` distributes tensor operations across multiple GPUs or nodes, partitioning the model horizontally. This strategy activates when models exceed single-GPU memory limits but fit within aggregate cluster VRAM. The planning logic in [`plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/plan.rs) explicitly maps `PlanRunPath::TensorParallel` to this `RunMode` variant.

### MoeOffload

Specifically designed for Mixture-of-Experts (MoE) architectures, `MoeOffload` maintains active expert layers in VRAM while spilling inactive experts to system RAM. This selective offloading reduces VRAM pressure for MoE models where only a fraction of experts activate per token, as detected by the model metadata analysis in the fitting pipeline.

### CpuOffload

The `CpuOffload` strategy executes core computations on GPU but spills large activation buffers to system RAM. This mode suits GPUs with limited VRAM that can accept moderate performance penalties for memory-intensive operations. The implementation tracks these spillover costs in the `RunModeFactors` calculations.

### CpuOnly

`CpuOnly` runs the entire inference workload in system RAM without GPU acceleration. This fallback mode activates automatically on machines lacking compatible GPUs or when models exceed all available GPU memory. While sacrificing speed, it ensures model accessibility across heterogeneous hardware configurations.

## How RunMode Selection Works

The selection pipeline follows a deterministic hierarchy defined in [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs):

1. **Hardware Detection**: The system scans available GPUs and RAM via [`hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/hardware.rs).
2. **Constraint Analysis**: `fit::ModelFit::analyze_*` methods in [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) evaluate model metadata including `min_vram_gb` and `is_moe` flags.
3. **Path Mapping**: `PlanRunPath` variants map to specific `RunMode` values (e.g., `PlanRunPath::Gpu.run_mode()` returns `RunMode::Gpu`).
4. **Factor Application**: The system applies `RunModeFactors` to adjust throughput predictions based on memory bandwidth limitations inherent to each strategy.

## Using RunMode in Your Code

When integrating llmfit into applications, you interact with `RunMode` through the `ModelFit` structure:

```rust
use llmfit_core::fit::{RunMode, ModelFit};

// After hardware detection and analysis
let chosen_mode = RunMode::Gpu;

let fit = ModelFit {
    run_mode: chosen_mode,
    // Additional fields for scoring and memory tracking
    ..Default::default()
};

println!("Selected run mode: {:?}", fit.run_mode);

```

This pattern appears throughout the codebase, from the terminal interface in [`llmfit-tui/src/tui_ui.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_ui.rs) to the Tauri desktop application in [`llmfit-desktop/src/main.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-desktop/src/main.rs).

## Where RunMode Appears in the Codebase

The enum propagates through multiple system layers:

- **Core Analysis**: [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) defines the enum and `RunModeFactors`.
- **Planning Engine**: [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs) maps execution paths to modes.
- **Terminal UI**: [`llmfit-tui/src/tui_ui.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_ui.rs) renders human-readable labels with color coding.
- **Desktop Interface**: [`llmfit-desktop/src/main.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-desktop/src/main.rs) displays mode selection to end users.
- **API Serialization**: [`llmfit-tui/src/serve_shared.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/serve_shared.rs) converts variants to short strings via `run_mode_code` for JSON API responses.

## Summary

- **`RunMode`** controls hardware utilization strategy in llmfit, defined in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs).
- **Five variants** cover the spectrum from pure GPU execution (`Gpu`) to CPU-only fallback (`CpuOnly`), including distributed (`TensorParallel`) and MoE-specific (`MoeOffload`) strategies.
- **Selection logic** in [`plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/plan.rs) automatically chooses modes based on hardware detection and model metadata constraints.
- **Integration points** span the TUI, desktop app, and HTTP API via serialization in [`serve_shared.rs`](https://github.com/AlexsJones/llmfit/blob/main/serve_shared.rs).

## Frequently Asked Questions

### Where is the RunMode enum defined in llmfit?

The `RunMode` enum is defined in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs). This core module also contains the `RunModeFactors` struct used to calculate performance characteristics for each execution strategy. The planning module at [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs) contains the mapping logic that selects appropriate modes based on hardware constraints.

### How does llmfit choose between Gpu and CpuOffload modes?

The library evaluates available VRAM against model requirements during the `fit::ModelFit::analyze_*` phase. If the model's `min_vram_gb` exceeds available GPU memory but the system has sufficient RAM, `CpuOffload` activates to spill activation buffers while keeping computation on GPU. If no GPU is available or the model exceeds all GPU memory, it falls back to `CpuOnly`.

### What is the difference between TensorParallel and MoeOffload?

`TensorParallel` horizontally partitions model layers across multiple GPUs for general large models, while `MoeOffload` specifically optimizes Mixture-of-Experts architectures by keeping only active expert layers in VRAM and offloading inactive ones to RAM. The latter requires MoE-specific metadata detection during model loading, as implemented in the core fitting logic.

### Can I manually specify a RunMode instead of using automatic selection?

Yes. While the default `PlanRunPath` logic in [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs) provides automatic selection, you can construct a `ModelFit` directly with a specific `RunMode` variant as shown in the implementation examples. This bypasses the hardware detection heuristics when you need explicit control over execution strategy.