RunMode Options in llmfit: 5 Execution Strategies for LLM Inference
The RunMode enum in llmfit defines five execution strategies—Gpu, TensorParallel, MoeOffload, CpuOffload, and CpuOnly—that determine how model weights and activations are distributed across GPU VRAM, system RAM, and CPU compute resources.
The RunMode type is the central hardware abstraction in the llmfit ecosystem that dictates how large language models execute on available hardware. Defined in the core library at llmfit-core/src/fit.rs, this enum orchestrates memory placement decisions ranging from single-GPU inference to distributed tensor parallelism, directly impacting throughput and latency characteristics according to the source code analysis.
What Is RunMode in llmfit?
Located in llmfit-core/src/fit.rs, the RunMode enum encodes the execution strategy selected during the model fitting phase. Each variant represents a distinct hardware utilization pattern optimized for specific memory constraints and computational topologies.
The enum works alongside the RunModeFactors struct in the same file to adjust throughput estimates based on the selected strategy. During initialization, the library's hardware detection module (hardware.rs) scans available resources, then the planning logic in llmfit-core/src/plan.rs maps specific execution paths to concrete RunMode values via the PlanRunPath enum.
The Five RunMode Variants Explained
Gpu
The Gpu variant keeps all model weights and activations resident in GPU VRAM. This mode delivers maximum inference speed when GPU memory capacity exceeds model requirements. According to the source code in fit.rs, this is the preferred default when min_vram_gb constraints are satisfied by available hardware.
TensorParallel
TensorParallel distributes tensor operations across multiple GPUs or nodes, partitioning the model horizontally. This strategy activates when models exceed single-GPU memory limits but fit within aggregate cluster VRAM. The planning logic in plan.rs explicitly maps PlanRunPath::TensorParallel to this RunMode variant.
MoeOffload
Specifically designed for Mixture-of-Experts (MoE) architectures, MoeOffload maintains active expert layers in VRAM while spilling inactive experts to system RAM. This selective offloading reduces VRAM pressure for MoE models where only a fraction of experts activate per token, as detected by the model metadata analysis in the fitting pipeline.
CpuOffload
The CpuOffload strategy executes core computations on GPU but spills large activation buffers to system RAM. This mode suits GPUs with limited VRAM that can accept moderate performance penalties for memory-intensive operations. The implementation tracks these spillover costs in the RunModeFactors calculations.
CpuOnly
CpuOnly runs the entire inference workload in system RAM without GPU acceleration. This fallback mode activates automatically on machines lacking compatible GPUs or when models exceed all available GPU memory. While sacrificing speed, it ensures model accessibility across heterogeneous hardware configurations.
How RunMode Selection Works
The selection pipeline follows a deterministic hierarchy defined in llmfit-core/src/plan.rs:
- Hardware Detection: The system scans available GPUs and RAM via
hardware.rs. - Constraint Analysis:
fit::ModelFit::analyze_*methods infit.rsevaluate model metadata includingmin_vram_gbandis_moeflags. - Path Mapping:
PlanRunPathvariants map to specificRunModevalues (e.g.,PlanRunPath::Gpu.run_mode()returnsRunMode::Gpu). - Factor Application: The system applies
RunModeFactorsto adjust throughput predictions based on memory bandwidth limitations inherent to each strategy.
Using RunMode in Your Code
When integrating llmfit into applications, you interact with RunMode through the ModelFit structure:
use llmfit_core::fit::{RunMode, ModelFit};
// After hardware detection and analysis
let chosen_mode = RunMode::Gpu;
let fit = ModelFit {
run_mode: chosen_mode,
// Additional fields for scoring and memory tracking
..Default::default()
};
println!("Selected run mode: {:?}", fit.run_mode);
This pattern appears throughout the codebase, from the terminal interface in llmfit-tui/src/tui_ui.rs to the Tauri desktop application in llmfit-desktop/src/main.rs.
Where RunMode Appears in the Codebase
The enum propagates through multiple system layers:
- Core Analysis:
llmfit-core/src/fit.rsdefines the enum andRunModeFactors. - Planning Engine:
llmfit-core/src/plan.rsmaps execution paths to modes. - Terminal UI:
llmfit-tui/src/tui_ui.rsrenders human-readable labels with color coding. - Desktop Interface:
llmfit-desktop/src/main.rsdisplays mode selection to end users. - API Serialization:
llmfit-tui/src/serve_shared.rsconverts variants to short strings viarun_mode_codefor JSON API responses.
Summary
RunModecontrols hardware utilization strategy in llmfit, defined inllmfit-core/src/fit.rs.- Five variants cover the spectrum from pure GPU execution (
Gpu) to CPU-only fallback (CpuOnly), including distributed (TensorParallel) and MoE-specific (MoeOffload) strategies. - Selection logic in
plan.rsautomatically chooses modes based on hardware detection and model metadata constraints. - Integration points span the TUI, desktop app, and HTTP API via serialization in
serve_shared.rs.
Frequently Asked Questions
Where is the RunMode enum defined in llmfit?
The RunMode enum is defined in llmfit-core/src/fit.rs. This core module also contains the RunModeFactors struct used to calculate performance characteristics for each execution strategy. The planning module at llmfit-core/src/plan.rs contains the mapping logic that selects appropriate modes based on hardware constraints.
How does llmfit choose between Gpu and CpuOffload modes?
The library evaluates available VRAM against model requirements during the fit::ModelFit::analyze_* phase. If the model's min_vram_gb exceeds available GPU memory but the system has sufficient RAM, CpuOffload activates to spill activation buffers while keeping computation on GPU. If no GPU is available or the model exceeds all GPU memory, it falls back to CpuOnly.
What is the difference between TensorParallel and MoeOffload?
TensorParallel horizontally partitions model layers across multiple GPUs for general large models, while MoeOffload specifically optimizes Mixture-of-Experts architectures by keeping only active expert layers in VRAM and offloading inactive ones to RAM. The latter requires MoE-specific metadata detection during model loading, as implemented in the core fitting logic.
Can I manually specify a RunMode instead of using automatic selection?
Yes. While the default PlanRunPath logic in llmfit-core/src/plan.rs provides automatic selection, you can construct a ModelFit directly with a specific RunMode variant as shown in the implementation examples. This bypasses the hardware detection heuristics when you need explicit control over execution strategy.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →