# How llmfit Chooses Between GPU, MoE-Offload, CPU-Offload, CPU-Only, and Tensor-Parallel Run Paths

> Discover how llmfit selects the best run path GPU, MoE-offload, CPU-offload, CPU-only, or tensor-parallel. Understand its hardware, memory, and speed evaluation.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: internals
- Published: 2026-09-11

---

**llmfit determines the optimal execution path through a deterministic three-stage pipeline that evaluates hardware capabilities, memory fit levels, and speed heuristics to select between GPU, CPU-offload, CPU-only, MoE-offload, and tensor-parallel modes.**

The **run path selection** logic in the `AlexsJones/llmfit` repository analyzes system specifications against model requirements to automatically pick the fastest viable execution strategy. This process ensures that Mixture-of-Experts (MoE) models receive special bandwidth-aware handling while respecting cluster configurations for distributed inference. The decision engine operates entirely within the `llmfit-core` crate, utilizing hardware detection, memory estimation, and throughput prediction to assign the appropriate `RunMode`.

## Hardware Detection and System Specifications

The selection process begins with comprehensive hardware introspection. The `SystemSpecs::detect()` function in [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs) probes the host system to gather critical metrics including GPU presence (`has_gpu`), total VRAM (`total_gpu_vram_gb`), available system RAM, CPU core count, GPU backend type, and whether the architecture uses unified memory.

These specifications form the foundation for all subsequent decisions. On unified memory systems (like Apple Silicon), certain offload paths become invalid since CPU and GPU share the same physical memory pool, forcing the optimizer toward `CpuOnly` execution when GPU capacity is exceeded.

## Memory-Fit Evaluation and FitLevel Calculation

Once hardware specs are established, `llmfit` calculates memory requirements using `model.estimate_memory_gb_with_kv(...)` and compares them against available resources. The helper function `fit_level_for()` in [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs) quantifies the compatibility between a model and a candidate path, returning a `FitLevel` enum variant: `Perfect`, `Good`, `Marginal`, or `TooTight`.

Paths receiving a `TooTight` classification are immediately discarded from consideration. The evaluation differs by execution target:

- **GPU Path** → Compares model memory against `total_gpu_vram_gb`
- **CPU-Offload Path** → Validates against system RAM (disabled on unified memory systems)
- **CPU-Only Path** → Compares against total system RAM (always available as fallback)

### GPU Path Validation

For the `PlanRunPath::Gpu` variant, the engine requires that VRAM capacity exceed the model's memory footprint including key-value cache projections. When `FitLevel::Perfect` or `FitLevel::Good` is achieved, this path receives the highest priority in subsequent ranking stages due to superior memory bandwidth.

### CPU Offload vs CPU Only

The `PlanRunPath::CpuOffload` option activates when the GPU lacks sufficient VRAM but system RAM can accommodate the weights. However, on unified-memory architectures, this path is disallowed because offloading provides no benefit when CPU and GPU share physical memory. In such cases, the optimizer defaults to `PlanRunPath::CpuOnly`, which loads all parameters into system RAM regardless of GPU availability.

## Speed-Heuristic Ranking and MoE Detection

After filtering by memory constraints, `llmfit` estimates throughput for each viable path. The `estimate_tps_with_gpu()` function in [`plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/plan.rs) calculates tokens-per-second estimates using measured GPU memory bandwidth from `resolve_gpu_bandwidth()`, falling back to fixed "K-values" constants when hardware profiling is unavailable.

Each estimate is adjusted by configurable penalty factors accessed via `config.run_mode_factors.for_run_mode(run_mode)`, allowing slower paths like `CpuOffload` or `CpuOnly` to be weighted appropriately against raw GPU performance.

### Special Handling for Mixture-of-Experts

MoE models trigger a specialized execution branch within the `speed_run_mode()` function:

```rust
// Located in llmfit-core/src/plan.rs
fn speed_run_mode(path: PlanRunPath, model: &LlmModel) -> RunMode {
    match path {
        PlanRunPath::CpuOffload if model.is_moe => RunMode::MoeOffload,
        other => other.run_mode(),
    }
}

```

When `model.is_moe` returns true and the system selects CPU offloading, `llmfit` assigns `RunMode::MoeOffload` rather than the standard `CpuOffload`. This distinction is critical because inactive experts are streamed from system RAM during inference, requiring the throughput estimator to use DDR memory bandwidth factors instead of device bandwidth in its calculations.

## Candidate Selection and RunMode Assignment

The final selection occurs in `evaluate_current()` within [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs). This function collects `(FitLevel, PlanRunPath, tps)` triples for all feasible paths and sorts them using a hierarchical comparator:

```rust
candidates.sort_by(|a, b| {
    // Higher fit level wins, then higher path priority (GPU > CPU-offload > CPU-only)
    rank(b.0).cmp(&rank(a.0))
        .then_with(|| p(b.1).cmp(&p(a.1)))
});

```

The winning candidate's `PlanRunPath` maps to a concrete `RunMode` (Gpu, CpuOffload, CpuOnly, or MoeOffload) which becomes the `PlanCurrentStatus.run_mode` value used for inference scheduling.

## Tensor-Parallel Cluster Mode

Tensor-parallel execution follows a separate decision branch activated exclusively by cluster configuration. When `system.cluster_mode == true` (detected in [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs)), the standard `PlanRunPath` enumeration is bypassed entirely.

In cluster mode, `llmfit` evaluates whether the model's memory requirements can be distributed across multiple nodes using NCCL. If sufficient aggregate memory exists across the cluster, `RunMode::TensorParallel` is assigned directly in the core analysis logic within [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs), enabling distributed inference without single-device memory constraints.

## Code Examples

### CLI Inspection of Selected Run Path

Query `llmfit` to display which execution mode it selects for a specific model:

```bash

# Analyze Qwen-7B and display the selected execution strategy

llmfit fit --perfect -n 1 "Qwen-7B"

```

Example output for a model fitting entirely in VRAM:

```

Model: Qwen-7B
Fit level: Perfect
Run mode: Gpu
Estimated TPS: 38.2

```

When VRAM is insufficient but system RAM can accommodate the weights:

```

Run mode: CpuOffload

```

For MoE architectures with insufficient GPU memory:

```

Run mode: MoeOffload

```

### Programmatic API Usage

Access the run path selection logic directly in Rust:

```rust
use llmfit_core::{models::LlmModel, hardware::SystemSpecs, plan::estimate_model_plan};

let specs = SystemSpecs::detect()?;  // Detects VRAM, RAM, CPU cores
let model = llmfit_core::models::resolve_model_selector(
    &llmfit_core::models::load_embedded()?, 
    "Qwen-7B"
)?;

let request = llmfit_core::plan::PlanRequest {
    context: 8192,
    quant: None,
    target_tps: None,
    kv_quant: None,
};

let plan = estimate_model_plan(&model, &request, &specs)?;
println!("Selected run mode: {}", plan.current.run_mode_text());
// Returns: "GPU", "CPU offload", "CPU-only", or "Tensor-parallel"

```

## Key Source Files and Architecture

| File | Responsibility |
|------|----------------|
| [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs) | Implements `SystemSpecs::detect()` for hardware introspection including VRAM, RAM, and cluster mode detection. |
| [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs) | Contains `PlanRunPath` definitions, `fit_level_for()` memory evaluation, `speed_run_mode()` mapping, and `estimate_tps_with_gpu()` throughput calculation. |
| [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) | Defines `RunMode` and `FitLevel` enums, implements `evaluate_current()` for candidate ranking and final selection. |
| [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs) | Provides model metadata including `is_moe` boolean for MoE detection and memory estimation methods. |
| [`llmfit-core/src/hwprofile.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hwprofile.rs) | Stores `CalcConfig` including user-tunable `run_mode_factors` for adjusting speed heuristics. |

## Summary

- **Hardware detection** via `SystemSpecs::detect()` establishes VRAM, RAM, and architectural constraints before evaluation begins.
- **Memory-fit evaluation** assigns `FitLevel` rankings to each path, immediately disqualifying `TooTight` candidates that exceed available resources.
- **MoE specialization** transforms `CpuOffload` paths into `MoeOffload` modes when `model.is_moe` is true, ensuring bandwidth calculations account for expert streaming overhead.
- **Throughput estimation** uses `estimate_tps_with_gpu()` with configurable penalty factors to compare execution speeds across viable paths.
- **Tensor-parallel** execution bypasses standard path selection when `cluster_mode` is enabled, enabling multi-node inference for models exceeding single-device memory.
- All decisions are deterministic and reproducible given identical `SystemSpecs`, model metadata, and `CalcConfig` parameters.

## Frequently Asked Questions

### How does llmfit decide between CPU offloading and CPU-only execution?

**CPU offloading** (`CpuOffload`) is selected when a discrete GPU exists with insufficient VRAM for the full model, but sufficient system RAM is available to host the weights. **CPU-only** (`CpuOnly`) becomes the fallback when no GPU is present, when the system uses unified memory (making offloading redundant), or when the GPU lacks sufficient VRAM and the CPU-offload path is disabled. The decision hinges on the `unified_memory` flag detected in [`hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/hardware.rs) and the `FitLevel` returned by `fit_level_for()`.

### What triggers MoE-offload mode instead of standard CPU offloading?

The `speed_run_mode()` function in [`plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/plan.rs) checks `model.is_moe` when evaluating a `CpuOffload` path. If the model architecture is a Mixture-of-Experts, the function maps the path to `RunMode::MoeOffload` rather than `RunMode::CpuOffload`. This distinction ensures that throughput estimates use DDR memory bandwidth calculations instead of GPU bandwidth, accurately reflecting the cost of streaming inactive experts from system RAM during inference.

### When does llmfit select tensor-parallel execution?

Tensor-parallel mode is activated exclusively when `system.cluster_mode == true`, indicating a multi-node environment with NCCL support. In this configuration, the standard single-device path selection is bypassed. The optimizer evaluates whether the model's memory requirements can be distributed across the cluster's aggregate memory; if feasible, `RunMode::TensorParallel` is assigned directly in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs), enabling sharded inference across multiple GPUs and nodes.

### Can users override the automatic run path selection?

Yes, users can influence selection through the `CalcConfig` structure defined in [`llmfit-core/src/hwprofile.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hwprofile.rs). The `run_mode_factors` field allows configuring penalty multipliers for each `RunMode`, effectively changing the relative ranking of paths during the speed-heuristic phase. While the engine prevents selecting paths with `FitLevel::TooTight` for safety, adjusting these factors can prioritize, for example, `CpuOffload` over marginal GPU fits when latency requirements permit slower execution.