What Are the Five Execution Paths for LLMs in llmfit? A Deep Dive into RunMode

The five execution paths for LLMs in llmfit are Gpu, TensorParallel, MoeOffload, CpuOffload, and CpuOnly, defined as variants of the RunMode enum in llmfit-core/src/fit.rs.

The llmfit crate determines how large language models execute on your hardware through a centralized RunMode enumeration. These five execution paths control where model weights reside—GPU VRAM, system RAM, or distributed across multiple devices—and directly impact inference speed and memory efficiency.

Understanding the RunMode Enum in llmfit

At the heart of llmfit's execution logic sits the RunMode enum. The source code defines this enum at line 188 in llmfit-core/src/fit.rs, with each variant representing a distinct hardware placement strategy.

The five RunMode variants are:

  • Gpu — Full GPU acceleration with all weights in VRAM
  • TensorParallel — Model parallelism across multiple GPUs
  • MoeOffload — Selective GPU residency for mixture-of-experts models
  • CpuOffload — Hybrid GPU/CPU execution with partial VRAM usage
  • CpuOnly — CPU-only execution in system RAM

These variants appear in the string mapping at lines 731–735 of fit.rs, where llmfit converts runtime selections into display-friendly names.

Execution Path 1: Gpu Mode

Gpu represents the fastest execution path. The entire model loads into GPU VRAM, eliminating data transfer bottlenecks between host and device memory.

llmfit selects this mode when hardware detection reports sufficient VRAM to hold the complete model weights, activations, and KV cache. The selection logic in fit.rs prioritizes this path whenever feasible.

use llmfit_core::fit::RunMode;

// Example: forcing GPU execution for a model that fits in VRAM
let selected_mode = RunMode::Gpu;
println!("Chosen execution path: {:?}", selected_mode);

Execution Path 2: TensorParallel Mode

TensorParallel distributes model layers across multiple GPUs or nodes. Instead of replicating weights on each device, llmfit shards tensors horizontally or vertically to pool available VRAM.

This path activates when a single GPU cannot accommodate the model, but multiple GPUs present sufficient aggregate memory. The hardware.rs module detects multi-GPU configurations and feeds this data to the fit analyzer.

Execution Path 3: MoeOffload Mode

MoeOffload optimizes mixture-of-experts (MoE) architectures. MoE models contain numerous "expert" subnetworks, but only a subset activates per token.

llmfit keeps active experts resident in GPU memory while offloading inactive experts to system RAM. This selective residency dramatically reduces VRAM requirements for models like Mixtral without sacrificing the GPU acceleration of the active computation path.

Execution Path 4: CpuOffload Mode

CpuOffload creates a tiered memory hierarchy. Hot weights and active compute graphs remain in GPU VRAM; cold weights spill to CPU-addressable system RAM.

This hybrid approach benefits systems with mid-range GPUs that lack sufficient VRAM for full model residency. The plan.rs module calculates precise memory partitions and estimates throughput degradation from cross-bus transfers.

Execution Path 5: CpuOnly Mode

CpuOnly serves as the universal fallback. The entire model executes in system RAM using CPU compute, requiring no GPU presence.

llmfit selects this path when no compatible GPU exists, when drivers are unavailable, or when explicit CPU-only constraints are applied. While slower, this mode guarantees model execution on any hardware configuration.

How llmfit Selects Execution Paths

The execution path selection follows a structured pipeline across core modules:

  1. Hardware detection — hardware.rs probes GPUs, VRAM capacity, and unified memory support
  2. Fit analysis — fit.rs compares model requirements against detected capabilities
  3. Path ranking — The analyzer prioritizes Gpu > TensorParallel > MoeOffload > CpuOffload > CpuOnly based on estimated throughput
  4. Output propagation — Selected RunMode flows to tui_app.rs, display.rs, and CLI formatting

The AGENTS.md documentation confirms this five-variant design, listing Gpu, MoeOffload, TensorParallel, CpuOffload, and CpuOnly as the complete execution strategy set.

Practical Usage Examples

CLI output showing assigned execution paths:

$ cargo run -- fit --list
MODEL                RUN MODE
llama-7b            Gpu
phi-2               CpuOffload
mixtral-8x7b        MoeOffload
opt-125m            CpuOnly

Filtering the TUI by execution mode:

app.filter.run_mode = Some(RunMode::CpuOnly);
app.apply_filters();   // Shows only models that will run on CPU only

Key Source Files for Execution Path Logic

File Responsibility
llmfit-core/src/fit.rs RunMode enum definition and selection algorithm
llmfit-core/src/hardware.rs GPU/VRAM detection for feasibility checks
llmfit-core/src/plan.rs Memory and throughput estimation per mode
llmfit-tui/src/display.rs User-facing RunMode formatting
AGENTS.md Architectural documentation of the five paths

Summary

  • Five execution paths — Gpu, TensorParallel, MoeOffload, CpuOffload, and CpuOnly — govern all LLM execution in llmfit
  • RunMode enum centralizes these variants at line 188 of llmfit-core/src/fit.rs
  • Hardware-aware selection automatically picks the fastest feasible path based on VRAM, GPU count, and model architecture
  • MoE specialization via MoeOffload provides optimized support for mixture-of-experts models
  • Graceful degradation from Gpu through CpuOnly ensures models run on any hardware

Frequently Asked Questions

How does llmfit decide which execution path to use?

llmfit queries system hardware through hardware.rs, then compares available resources against model memory requirements in fit.rs. The algorithm prioritizes GPU-resident modes and falls back through TensorParallel, MoeOffload, CpuOffload, and finally CpuOnly until finding a feasible configuration.

Can I force a specific execution path instead of auto-selection?

Yes. The core library exposes RunMode as a public enum, and the CLI/TUI layers accept runtime filters. Set app.filter.run_mode to your preferred variant before calling the fit analysis to override automatic selection.

What makes MoeOffload different from standard CpuOffload?

MoeOffload specializes in mixture-of-experts architectures by dynamically swapping experts based on activation patterns, keeping only active experts in GPU memory. CpuOffload uses a static partitioning scheme without expert-aware logic, making MoeOffload substantially more efficient for MoE models.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →