# What Are the Five Execution Paths for LLMs in llmfit? A Deep Dive into RunMode

> Explore the five LLM execution paths in llmfit: Gpu, TensorParallel, MoeOffload, CpuOffload, and CpuOnly. Understand the RunMode enum and optimize your LLM performance with this deep dive.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: deep-dive
- Published: 2026-08-20

---

**The five execution paths for LLMs in llmfit are Gpu, TensorParallel, MoeOffload, CpuOffload, and CpuOnly, defined as variants of the `RunMode` enum in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs).**

The `llmfit` crate determines how large language models execute on your hardware through a centralized `RunMode` enumeration. These five execution paths control where model weights reside—GPU VRAM, system RAM, or distributed across multiple devices—and directly impact inference speed and memory efficiency.

## Understanding the RunMode Enum in llmfit

At the heart of `llmfit`'s execution logic sits the `RunMode` enum. The source code defines this enum at **line 188 in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs)**, with each variant representing a distinct hardware placement strategy.

The five `RunMode` variants are:

- **Gpu** — Full GPU acceleration with all weights in VRAM
- **TensorParallel** — Model parallelism across multiple GPUs
- **MoeOffload** — Selective GPU residency for mixture-of-experts models
- **CpuOffload** — Hybrid GPU/CPU execution with partial VRAM usage
- **CpuOnly** — CPU-only execution in system RAM

These variants appear in the string mapping at **lines 731–735 of [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs)**, where `llmfit` converts runtime selections into display-friendly names.

## Execution Path 1: Gpu Mode

**Gpu** represents the fastest execution path. The entire model loads into GPU VRAM, eliminating data transfer bottlenecks between host and device memory.

`llmfit` selects this mode when hardware detection reports sufficient VRAM to hold the complete model weights, activations, and KV cache. The selection logic in [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) prioritizes this path whenever feasible.

```rust
use llmfit_core::fit::RunMode;

// Example: forcing GPU execution for a model that fits in VRAM
let selected_mode = RunMode::Gpu;
println!("Chosen execution path: {:?}", selected_mode);

```

## Execution Path 2: TensorParallel Mode

**TensorParallel** distributes model layers across multiple GPUs or nodes. Instead of replicating weights on each device, `llmfit` shards tensors horizontally or vertically to pool available VRAM.

This path activates when a single GPU cannot accommodate the model, but multiple GPUs present sufficient aggregate memory. The [`hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/hardware.rs) module detects multi-GPU configurations and feeds this data to the fit analyzer.

## Execution Path 3: MoeOffload Mode

**MoeOffload** optimizes mixture-of-experts (MoE) architectures. MoE models contain numerous "expert" subnetworks, but only a subset activates per token.

`llmfit` keeps active experts resident in GPU memory while offloading inactive experts to system RAM. This selective residency dramatically reduces VRAM requirements for models like Mixtral without sacrificing the GPU acceleration of the active computation path.

## Execution Path 4: CpuOffload Mode

**CpuOffload** creates a tiered memory hierarchy. Hot weights and active compute graphs remain in GPU VRAM; cold weights spill to CPU-addressable system RAM.

This hybrid approach benefits systems with mid-range GPUs that lack sufficient VRAM for full model residency. The [`plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/plan.rs) module calculates precise memory partitions and estimates throughput degradation from cross-bus transfers.

## Execution Path 5: CpuOnly Mode

**CpuOnly** serves as the universal fallback. The entire model executes in system RAM using CPU compute, requiring no GPU presence.

`llmfit` selects this path when no compatible GPU exists, when drivers are unavailable, or when explicit CPU-only constraints are applied. While slower, this mode guarantees model execution on any hardware configuration.

## How llmfit Selects Execution Paths

The execution path selection follows a structured pipeline across core modules:

1. **Hardware detection** — [`hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/hardware.rs) probes GPUs, VRAM capacity, and unified memory support
2. **Fit analysis** — [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) compares model requirements against detected capabilities
3. **Path ranking** — The analyzer prioritizes `Gpu` > `TensorParallel` > `MoeOffload` > `CpuOffload` > `CpuOnly` based on estimated throughput
4. **Output propagation** — Selected `RunMode` flows to [`tui_app.rs`](https://github.com/AlexsJones/llmfit/blob/main/tui_app.rs), [`display.rs`](https://github.com/AlexsJones/llmfit/blob/main/display.rs), and CLI formatting

The [`AGENTS.md`](https://github.com/AlexsJones/llmfit/blob/main/AGENTS.md) documentation confirms this five-variant design, listing `Gpu`, `MoeOffload`, `TensorParallel`, `CpuOffload`, and `CpuOnly` as the complete execution strategy set.

## Practical Usage Examples

**CLI output showing assigned execution paths:**

```bash
$ cargo run -- fit --list
MODEL                RUN MODE
llama-7b            Gpu
phi-2               CpuOffload
mixtral-8x7b        MoeOffload
opt-125m            CpuOnly

```

**Filtering the TUI by execution mode:**

```rust
app.filter.run_mode = Some(RunMode::CpuOnly);
app.apply_filters();   // Shows only models that will run on CPU only

```

## Key Source Files for Execution Path Logic

| File | Responsibility |
|------|--------------|
| [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) | `RunMode` enum definition and selection algorithm |
| [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs) | GPU/VRAM detection for feasibility checks |
| [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs) | Memory and throughput estimation per mode |
| [`llmfit-tui/src/display.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/display.rs) | User-facing `RunMode` formatting |
| [`AGENTS.md`](https://github.com/AlexsJones/llmfit/blob/main/AGENTS.md) | Architectural documentation of the five paths |

## Summary

- **Five execution paths** — Gpu, TensorParallel, MoeOffload, CpuOffload, and CpuOnly — govern all LLM execution in `llmfit`
- **RunMode enum** centralizes these variants at line 188 of [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs)
- **Hardware-aware selection** automatically picks the fastest feasible path based on VRAM, GPU count, and model architecture
- **MoE specialization** via MoeOffload provides optimized support for mixture-of-experts models
- **Graceful degradation** from Gpu through CpuOnly ensures models run on any hardware

## Frequently Asked Questions

### How does llmfit decide which execution path to use?

`llmfit` queries system hardware through [`hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/hardware.rs), then compares available resources against model memory requirements in [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs). The algorithm prioritizes GPU-resident modes and falls back through TensorParallel, MoeOffload, CpuOffload, and finally CpuOnly until finding a feasible configuration.

### Can I force a specific execution path instead of auto-selection?

Yes. The core library exposes `RunMode` as a public enum, and the CLI/TUI layers accept runtime filters. Set `app.filter.run_mode` to your preferred variant before calling the fit analysis to override automatic selection.

### What makes MoeOffload different from standard CpuOffload?

**MoeOffload** specializes in mixture-of-experts architectures by dynamically swapping experts based on activation patterns, keeping only active experts in GPU memory. **CpuOffload** uses a static partitioning scheme without expert-aware logic, making MoeOffload substantially more efficient for MoE models.