How llmfit Chooses Between GPU, MoE-Offload, CPU-Offload, CPU-Only, and Tensor-Parallel Run Paths
llmfit determines the optimal execution path through a deterministic three-stage pipeline that evaluates hardware capabilities, memory fit levels, and speed heuristics to select between GPU, CPU-offload, CPU-only, MoE-offload, and tensor-parallel modes.
The run path selection logic in the AlexsJones/llmfit repository analyzes system specifications against model requirements to automatically pick the fastest viable execution strategy. This process ensures that Mixture-of-Experts (MoE) models receive special bandwidth-aware handling while respecting cluster configurations for distributed inference. The decision engine operates entirely within the llmfit-core crate, utilizing hardware detection, memory estimation, and throughput prediction to assign the appropriate RunMode.
Hardware Detection and System Specifications
The selection process begins with comprehensive hardware introspection. The SystemSpecs::detect() function in llmfit-core/src/hardware.rs probes the host system to gather critical metrics including GPU presence (has_gpu), total VRAM (total_gpu_vram_gb), available system RAM, CPU core count, GPU backend type, and whether the architecture uses unified memory.
These specifications form the foundation for all subsequent decisions. On unified memory systems (like Apple Silicon), certain offload paths become invalid since CPU and GPU share the same physical memory pool, forcing the optimizer toward CpuOnly execution when GPU capacity is exceeded.
Memory-Fit Evaluation and FitLevel Calculation
Once hardware specs are established, llmfit calculates memory requirements using model.estimate_memory_gb_with_kv(...) and compares them against available resources. The helper function fit_level_for() in llmfit-core/src/plan.rs quantifies the compatibility between a model and a candidate path, returning a FitLevel enum variant: Perfect, Good, Marginal, or TooTight.
Paths receiving a TooTight classification are immediately discarded from consideration. The evaluation differs by execution target:
- GPU Path → Compares model memory against
total_gpu_vram_gb - CPU-Offload Path → Validates against system RAM (disabled on unified memory systems)
- CPU-Only Path → Compares against total system RAM (always available as fallback)
GPU Path Validation
For the PlanRunPath::Gpu variant, the engine requires that VRAM capacity exceed the model's memory footprint including key-value cache projections. When FitLevel::Perfect or FitLevel::Good is achieved, this path receives the highest priority in subsequent ranking stages due to superior memory bandwidth.
CPU Offload vs CPU Only
The PlanRunPath::CpuOffload option activates when the GPU lacks sufficient VRAM but system RAM can accommodate the weights. However, on unified-memory architectures, this path is disallowed because offloading provides no benefit when CPU and GPU share physical memory. In such cases, the optimizer defaults to PlanRunPath::CpuOnly, which loads all parameters into system RAM regardless of GPU availability.
Speed-Heuristic Ranking and MoE Detection
After filtering by memory constraints, llmfit estimates throughput for each viable path. The estimate_tps_with_gpu() function in plan.rs calculates tokens-per-second estimates using measured GPU memory bandwidth from resolve_gpu_bandwidth(), falling back to fixed "K-values" constants when hardware profiling is unavailable.
Each estimate is adjusted by configurable penalty factors accessed via config.run_mode_factors.for_run_mode(run_mode), allowing slower paths like CpuOffload or CpuOnly to be weighted appropriately against raw GPU performance.
Special Handling for Mixture-of-Experts
MoE models trigger a specialized execution branch within the speed_run_mode() function:
// Located in llmfit-core/src/plan.rs
fn speed_run_mode(path: PlanRunPath, model: &LlmModel) -> RunMode {
match path {
PlanRunPath::CpuOffload if model.is_moe => RunMode::MoeOffload,
other => other.run_mode(),
}
}
When model.is_moe returns true and the system selects CPU offloading, llmfit assigns RunMode::MoeOffload rather than the standard CpuOffload. This distinction is critical because inactive experts are streamed from system RAM during inference, requiring the throughput estimator to use DDR memory bandwidth factors instead of device bandwidth in its calculations.
Candidate Selection and RunMode Assignment
The final selection occurs in evaluate_current() within llmfit-core/src/fit.rs. This function collects (FitLevel, PlanRunPath, tps) triples for all feasible paths and sorts them using a hierarchical comparator:
candidates.sort_by(|a, b| {
// Higher fit level wins, then higher path priority (GPU > CPU-offload > CPU-only)
rank(b.0).cmp(&rank(a.0))
.then_with(|| p(b.1).cmp(&p(a.1)))
});
The winning candidate's PlanRunPath maps to a concrete RunMode (Gpu, CpuOffload, CpuOnly, or MoeOffload) which becomes the PlanCurrentStatus.run_mode value used for inference scheduling.
Tensor-Parallel Cluster Mode
Tensor-parallel execution follows a separate decision branch activated exclusively by cluster configuration. When system.cluster_mode == true (detected in llmfit-core/src/hardware.rs), the standard PlanRunPath enumeration is bypassed entirely.
In cluster mode, llmfit evaluates whether the model's memory requirements can be distributed across multiple nodes using NCCL. If sufficient aggregate memory exists across the cluster, RunMode::TensorParallel is assigned directly in the core analysis logic within llmfit-core/src/fit.rs, enabling distributed inference without single-device memory constraints.
Code Examples
CLI Inspection of Selected Run Path
Query llmfit to display which execution mode it selects for a specific model:
# Analyze Qwen-7B and display the selected execution strategy
llmfit fit --perfect -n 1 "Qwen-7B"
Example output for a model fitting entirely in VRAM:
Model: Qwen-7B
Fit level: Perfect
Run mode: Gpu
Estimated TPS: 38.2
When VRAM is insufficient but system RAM can accommodate the weights:
Run mode: CpuOffload
For MoE architectures with insufficient GPU memory:
Run mode: MoeOffload
Programmatic API Usage
Access the run path selection logic directly in Rust:
use llmfit_core::{models::LlmModel, hardware::SystemSpecs, plan::estimate_model_plan};
let specs = SystemSpecs::detect()?; // Detects VRAM, RAM, CPU cores
let model = llmfit_core::models::resolve_model_selector(
&llmfit_core::models::load_embedded()?,
"Qwen-7B"
)?;
let request = llmfit_core::plan::PlanRequest {
context: 8192,
quant: None,
target_tps: None,
kv_quant: None,
};
let plan = estimate_model_plan(&model, &request, &specs)?;
println!("Selected run mode: {}", plan.current.run_mode_text());
// Returns: "GPU", "CPU offload", "CPU-only", or "Tensor-parallel"
Key Source Files and Architecture
| File | Responsibility |
|---|---|
llmfit-core/src/hardware.rs |
Implements SystemSpecs::detect() for hardware introspection including VRAM, RAM, and cluster mode detection. |
llmfit-core/src/plan.rs |
Contains PlanRunPath definitions, fit_level_for() memory evaluation, speed_run_mode() mapping, and estimate_tps_with_gpu() throughput calculation. |
llmfit-core/src/fit.rs |
Defines RunMode and FitLevel enums, implements evaluate_current() for candidate ranking and final selection. |
llmfit-core/src/models.rs |
Provides model metadata including is_moe boolean for MoE detection and memory estimation methods. |
llmfit-core/src/hwprofile.rs |
Stores CalcConfig including user-tunable run_mode_factors for adjusting speed heuristics. |
Summary
- Hardware detection via
SystemSpecs::detect()establishes VRAM, RAM, and architectural constraints before evaluation begins. - Memory-fit evaluation assigns
FitLevelrankings to each path, immediately disqualifyingTooTightcandidates that exceed available resources. - MoE specialization transforms
CpuOffloadpaths intoMoeOffloadmodes whenmodel.is_moeis true, ensuring bandwidth calculations account for expert streaming overhead. - Throughput estimation uses
estimate_tps_with_gpu()with configurable penalty factors to compare execution speeds across viable paths. - Tensor-parallel execution bypasses standard path selection when
cluster_modeis enabled, enabling multi-node inference for models exceeding single-device memory. - All decisions are deterministic and reproducible given identical
SystemSpecs, model metadata, andCalcConfigparameters.
Frequently Asked Questions
How does llmfit decide between CPU offloading and CPU-only execution?
CPU offloading (CpuOffload) is selected when a discrete GPU exists with insufficient VRAM for the full model, but sufficient system RAM is available to host the weights. CPU-only (CpuOnly) becomes the fallback when no GPU is present, when the system uses unified memory (making offloading redundant), or when the GPU lacks sufficient VRAM and the CPU-offload path is disabled. The decision hinges on the unified_memory flag detected in hardware.rs and the FitLevel returned by fit_level_for().
What triggers MoE-offload mode instead of standard CPU offloading?
The speed_run_mode() function in plan.rs checks model.is_moe when evaluating a CpuOffload path. If the model architecture is a Mixture-of-Experts, the function maps the path to RunMode::MoeOffload rather than RunMode::CpuOffload. This distinction ensures that throughput estimates use DDR memory bandwidth calculations instead of GPU bandwidth, accurately reflecting the cost of streaming inactive experts from system RAM during inference.
When does llmfit select tensor-parallel execution?
Tensor-parallel mode is activated exclusively when system.cluster_mode == true, indicating a multi-node environment with NCCL support. In this configuration, the standard single-device path selection is bypassed. The optimizer evaluates whether the model's memory requirements can be distributed across the cluster's aggregate memory; if feasible, RunMode::TensorParallel is assigned directly in llmfit-core/src/fit.rs, enabling sharded inference across multiple GPUs and nodes.
Can users override the automatic run path selection?
Yes, users can influence selection through the CalcConfig structure defined in llmfit-core/src/hwprofile.rs. The run_mode_factors field allows configuring penalty multipliers for each RunMode, effectively changing the relative ranking of paths during the speed-heuristic phase. While the engine prevents selecting paths with FitLevel::TooTight for safety, adjusting these factors can prioritize, for example, CpuOffload over marginal GPU fits when latency requirements permit slower execution.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →