# How the llmfit Fit Algorithm Chooses a Runtime: Hardware-Aware Execution in Rust

> Discover how the llmfit fit algorithm chooses a runtime. It ranks hardware against model needs, filters by backend, and picks the first GPU-first option for efficient execution.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: internals
- Published: 2026-09-13

---

**The llmfit fit algorithm selects a runtime by ranking hardware capabilities against model requirements, filtering by backend support, and choosing the first viable option from a GPU-first hierarchy.**

The **llmfit** project by AlexsJones implements an intelligent model-placement engine that automatically determines whether to run large language models on GPU, CPU, or distributed configurations. Understanding how the **llmfit fit algorithm** evaluates your system ensures you can predict which execution path the tool will select for any given model.

## Three Factors That Drive Runtime Selection

The algorithm weighs three primary inputs before committing to an execution strategy.

### Hardware Capability Detection

The process begins in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) with `SystemSpecs::detect()`, which probes the host machine for available RAM, VRAM (or unified memory on Apple Silicon), GPU type, and CPU core count. This structure populates fields like `ram_gb`, `vram_gb`, and a `unified_memory` boolean flag that downstream logic uses to calculate available headroom.

### Model Resource Requirements

Each model entry in the embedded catalog ([`hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/hf_models.json)) declares `min_ram_gb` and `min_vram_gb` thresholds derived from parameter counts. When `ModelDatabase::new()` loads this data in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs), it exposes these requirements to the comparison engine. The algorithm treats these values as hard constraints; any runtime that cannot satisfy both limits is immediately discarded.

### Backend Provider Constraints

Not all execution modes are available for every provider. The `providers::Provider::supported_runtimes()` method in [`llmfit-core/src/providers.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/providers.rs) returns a bitmask of feasible runtimes for backends like Ollama, llama.cpp, MLX, or vLLM. The core filtering logic intersects the hardware capabilities with this provider-specific capability set to eliminate unsupported combinations.

## The Runtime Selection Logic in fit.rs

The decision flow follows a deterministic priority queue defined by the `RunMode` enum in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs). The hierarchy prefers GPU acceleration and progressively falls back to memory-conservative alternatives:

1. **Gpu** – Full GPU execution when VRAM satisfies `min_vram_gb`
2. **MoeOffload** – Mixture-of-Experts layer offloading for partial GPU utilization
3. **CpuOffload** – CPU RAM spillover when GPU memory is insufficient (skipped on Apple Silicon unified memory)
4. **CpuOnly** – Pure CPU execution when `min_ram_gb` is met but no GPU is viable
5. **TensorParallel** – Distributed execution across multiple nodes when `SystemSpecs::cluster` is detected

The selection loop iterates over this order in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) (lines 260–300). For each `RunMode`, the code verifies that the current `SystemSpecs` meets the model’s `min_vram_gb` (for GPU paths) or `min_ram_gb` (for CPU paths) and that the chosen provider reports support via `supported_runtimes()`. The first variant satisfying both conditions is selected. If none match, the model receives a `TooTight` classification and is excluded from results.

## Special Cases: Apple Silicon and Distributed Execution

Apple Silicon systems trigger specialized logic because macOS employs unified memory where RAM and VRAM share the same physical pool. When `SystemSpecs::detect()` sets `unified_memory = true`, the algorithm skips the `CpuOffload` path—there is no distinct RAM pool to spill into—and treats the unified pool as VRAM for GPU-bound calculations.

For clustered environments, the engine considers `TensorParallel` only after exhausting single-node runtimes. This mode requires the provider to expose distributed execution capabilities and multiple nodes to be registered in the current `SystemSpecs` cluster configuration.

## Computing the Final Fit Score

Once a runtime is locked, `ModelFit::analyze_with_forced_runtime()` in [`llmfit-core/src/analysis.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/analysis.rs) (lines 78–140) calculates memory utilization percentages, estimated throughput, and a composite fit score. The algorithm assigns categorical fit levels—**Perfect**, **Good**, **Marginal**, or **TooTight**—based on how comfortably the remaining headroom exceeds the model’s declared minimums.

## Example: Manually Invoking the Runtime Selector

You can programmatically replicate the CLI’s decision logic using the core library directly:

```rust
use llmfit_core::{SystemSpecs, ModelDatabase, ModelFit};

fn main() -> Result<(), Box<dyn std::error::Error>> {
    // Detect host hardware capabilities
    let specs = SystemSpecs::detect()?;

    // Load the embedded HuggingFace model catalog
    let db = ModelDatabase::new()?;

    // Analyze all models and capture viable fits
    let fits = db.models.iter()
        .filter_map(|model| ModelFit::analyze(model, &specs).ok())
        .collect::<Vec<_>>();

    // Display the selected runtime for each compatible model
    for fit in fits {
        println!("Model: {:<30} → Runtime: {}", fit.model.name, fit.run_mode);
    }
    Ok(())
}

```

This pattern mirrors the internal behavior of [`llmfit-tui/src/main.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/main.rs), which orchestrates detection, analysis, and tabular output when running `cargo run -- --cli`.

## Summary

- **Hardware detection** via `SystemSpecs::detect()` establishes the baseline RAM, VRAM, and GPU topology.
- **Model requirements** from [`hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/hf_models.json) provide immutable `min_ram_gb` and `min_vram_gb` constraints.
- **Provider filtering** through `supported_runtimes()` eliminates execution modes the backend cannot handle.
- The **RunMode hierarchy** (Gpu → MoeOffload → CpuOffload → CpuOnly → TensorParallel) defines the selection priority.
- **Apple Silicon systems** skip `CpuOffload` due to unified memory architecture.
- The **fit score** in [`analysis.rs`](https://github.com/AlexsJones/llmfit/blob/main/analysis.rs) quantifies how well the chosen runtime matches the hardware after selection.

## Frequently Asked Questions

### How does llmfit handle machines without a GPU?

When `SystemSpecs::detect()` reports zero VRAM, the algorithm immediately skips the `Gpu` and `MoeOffload` variants. It evaluates `CpuOnly` against the available system RAM using `min_ram_gb` thresholds. If the model fits within physical memory limits, `CpuOnly` is selected; otherwise, the model is marked `TooTight`.

### What happens if my hardware barely meets the requirements?

The `ModelFit::analyze_with_forced_runtime()` function computes a headroom ratio. If utilization exceeds safe margins (typically >90% of available resources), the fit level downgrades from **Good** to **Marginal**. This signals that the runtime will function but may suffer from swapping or thermal throttling.

### Can I force a specific runtime instead of auto-selection?

While the core algorithm defaults to auto-selection via `ModelFit::analyze()`, the library exposes `analyze_with_forced_runtime()` which accepts an explicit `RunMode` parameter. This bypasses the hierarchy logic and attempts to fit the model into the specified runtime regardless of the default priority order, provided the provider supports it.

### Where does llmfit get the model requirements data?

The requirements are baked into the binary at compile time via [`hf_models.json`](https://github.com/AlexsJones/llmfit/blob/main/hf_models.json). The `ModelDatabase::new()` constructor in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) (lines 120–150) deserializes this embedded JSON file, which contains pre-calculated `min_ram_gb` and `min_vram_gb` values derived from model parameter counts and quantization profiles.