How the llmfit Fit Algorithm Chooses a Runtime: Hardware-Aware Execution in Rust
The llmfit fit algorithm selects a runtime by ranking hardware capabilities against model requirements, filtering by backend support, and choosing the first viable option from a GPU-first hierarchy.
The llmfit project by AlexsJones implements an intelligent model-placement engine that automatically determines whether to run large language models on GPU, CPU, or distributed configurations. Understanding how the llmfit fit algorithm evaluates your system ensures you can predict which execution path the tool will select for any given model.
Three Factors That Drive Runtime Selection
The algorithm weighs three primary inputs before committing to an execution strategy.
Hardware Capability Detection
The process begins in llmfit-core/src/fit.rs with SystemSpecs::detect(), which probes the host machine for available RAM, VRAM (or unified memory on Apple Silicon), GPU type, and CPU core count. This structure populates fields like ram_gb, vram_gb, and a unified_memory boolean flag that downstream logic uses to calculate available headroom.
Model Resource Requirements
Each model entry in the embedded catalog (hf_models.json) declares min_ram_gb and min_vram_gb thresholds derived from parameter counts. When ModelDatabase::new() loads this data in llmfit-core/src/fit.rs, it exposes these requirements to the comparison engine. The algorithm treats these values as hard constraints; any runtime that cannot satisfy both limits is immediately discarded.
Backend Provider Constraints
Not all execution modes are available for every provider. The providers::Provider::supported_runtimes() method in llmfit-core/src/providers.rs returns a bitmask of feasible runtimes for backends like Ollama, llama.cpp, MLX, or vLLM. The core filtering logic intersects the hardware capabilities with this provider-specific capability set to eliminate unsupported combinations.
The Runtime Selection Logic in fit.rs
The decision flow follows a deterministic priority queue defined by the RunMode enum in llmfit-core/src/fit.rs. The hierarchy prefers GPU acceleration and progressively falls back to memory-conservative alternatives:
- Gpu – Full GPU execution when VRAM satisfies
min_vram_gb - MoeOffload – Mixture-of-Experts layer offloading for partial GPU utilization
- CpuOffload – CPU RAM spillover when GPU memory is insufficient (skipped on Apple Silicon unified memory)
- CpuOnly – Pure CPU execution when
min_ram_gbis met but no GPU is viable - TensorParallel – Distributed execution across multiple nodes when
SystemSpecs::clusteris detected
The selection loop iterates over this order in llmfit-core/src/fit.rs (lines 260–300). For each RunMode, the code verifies that the current SystemSpecs meets the model’s min_vram_gb (for GPU paths) or min_ram_gb (for CPU paths) and that the chosen provider reports support via supported_runtimes(). The first variant satisfying both conditions is selected. If none match, the model receives a TooTight classification and is excluded from results.
Special Cases: Apple Silicon and Distributed Execution
Apple Silicon systems trigger specialized logic because macOS employs unified memory where RAM and VRAM share the same physical pool. When SystemSpecs::detect() sets unified_memory = true, the algorithm skips the CpuOffload path—there is no distinct RAM pool to spill into—and treats the unified pool as VRAM for GPU-bound calculations.
For clustered environments, the engine considers TensorParallel only after exhausting single-node runtimes. This mode requires the provider to expose distributed execution capabilities and multiple nodes to be registered in the current SystemSpecs cluster configuration.
Computing the Final Fit Score
Once a runtime is locked, ModelFit::analyze_with_forced_runtime() in llmfit-core/src/analysis.rs (lines 78–140) calculates memory utilization percentages, estimated throughput, and a composite fit score. The algorithm assigns categorical fit levels—Perfect, Good, Marginal, or TooTight—based on how comfortably the remaining headroom exceeds the model’s declared minimums.
Example: Manually Invoking the Runtime Selector
You can programmatically replicate the CLI’s decision logic using the core library directly:
use llmfit_core::{SystemSpecs, ModelDatabase, ModelFit};
fn main() -> Result<(), Box<dyn std::error::Error>> {
// Detect host hardware capabilities
let specs = SystemSpecs::detect()?;
// Load the embedded HuggingFace model catalog
let db = ModelDatabase::new()?;
// Analyze all models and capture viable fits
let fits = db.models.iter()
.filter_map(|model| ModelFit::analyze(model, &specs).ok())
.collect::<Vec<_>>();
// Display the selected runtime for each compatible model
for fit in fits {
println!("Model: {:<30} → Runtime: {}", fit.model.name, fit.run_mode);
}
Ok(())
}
This pattern mirrors the internal behavior of llmfit-tui/src/main.rs, which orchestrates detection, analysis, and tabular output when running cargo run -- --cli.
Summary
- Hardware detection via
SystemSpecs::detect()establishes the baseline RAM, VRAM, and GPU topology. - Model requirements from
hf_models.jsonprovide immutablemin_ram_gbandmin_vram_gbconstraints. - Provider filtering through
supported_runtimes()eliminates execution modes the backend cannot handle. - The RunMode hierarchy (Gpu → MoeOffload → CpuOffload → CpuOnly → TensorParallel) defines the selection priority.
- Apple Silicon systems skip
CpuOffloaddue to unified memory architecture. - The fit score in
analysis.rsquantifies how well the chosen runtime matches the hardware after selection.
Frequently Asked Questions
How does llmfit handle machines without a GPU?
When SystemSpecs::detect() reports zero VRAM, the algorithm immediately skips the Gpu and MoeOffload variants. It evaluates CpuOnly against the available system RAM using min_ram_gb thresholds. If the model fits within physical memory limits, CpuOnly is selected; otherwise, the model is marked TooTight.
What happens if my hardware barely meets the requirements?
The ModelFit::analyze_with_forced_runtime() function computes a headroom ratio. If utilization exceeds safe margins (typically >90% of available resources), the fit level downgrades from Good to Marginal. This signals that the runtime will function but may suffer from swapping or thermal throttling.
Can I force a specific runtime instead of auto-selection?
While the core algorithm defaults to auto-selection via ModelFit::analyze(), the library exposes analyze_with_forced_runtime() which accepts an explicit RunMode parameter. This bypasses the hierarchy logic and attempts to fit the model into the specified runtime regardless of the default priority order, provided the provider supports it.
Where does llmfit get the model requirements data?
The requirements are baked into the binary at compile time via hf_models.json. The ModelDatabase::new() constructor in llmfit-core/src/fit.rs (lines 120–150) deserializes this embedded JSON file, which contains pre-calculated min_ram_gb and min_vram_gb values derived from model parameter counts and quantization profiles.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →