How llmfit Determines the FitLevel of a Model: Memory-Aware Analysis in Rust
llmfit calculates a model's FitLevel by comparing required memory against available hardware resources, applying ratio-based thresholds to categorize fit quality, and adjusting the final verdict based on execution runtime constraints.
The open-source AlexsJones/llmfit repository implements a sophisticated hardware-aware analysis engine that determines how well a large language model fits on your local machine. Understanding how llmfit determines the FitLevel of a model reveals the deterministic logic behind its memory-ratio calculations and runtime-based capping system.
The Three-Phase FitLevel Determination Process
The determination logic resides in llmfit-core/src/fit.rs, specifically within the ModelFit::analyze implementation (lines 936-942). The process unfolds across three distinct phases:
Phase 1: Runtime and Execution Path Selection
First, llmfit selects an InferenceRuntime based on host hardware configuration. According to the source code (lines 500-507), the system prioritizes Metal for MLX, falls back to llama-cpp when a GPU is present, and selects vLLM for pre-quantized models.
Subsequently, the code determines the RunMode (lines 511-560). The available modes include:
Gpufor direct GPU executionTensorParallelfor multi-GPU setupsMoeOffloadfor Mixture-of-Experts offloadingCpuOffloadfor partial CPU fallbackCpuOnlyfor CPU-exclusive execution
Phase 2: Memory Requirement vs. Availability Calculation
The system computes mem_required based on the selected quantization parameters and mem_available from the detected hardware pool. For unified-memory architectures (such as Apple Silicon), GPU and CPU share a single memory pool. On distinct-memory systems, the algorithm prioritizes GPU VRAM first, spilling to system RAM only when necessary.
Phase 3: FitLevel Scoring and Capping
The score_fit(mem_required, mem_available, run_mode) function calculates the critical memory-ratio by dividing required memory by available memory. This raw ratio feeds into pure_ratio_verdict, which maps values to provisional FitLevel categories before cap_for_run_mode applies runtime-based adjustments.
The FitLevel Threshold System
The pure_ratio_verdict function in llmfit-core/src/fit.rs applies the following deterministic thresholds:
- ≤ 0.7: Perfect (optimal headroom)
- ≤ 1.0: Good (acceptable utilization)
- ≤ 1.5: Marginal (risk of swapping)
- > 1.5: TooTight (guaranteed failure)
However, the provisional verdict undergoes modification through cap_for_run_mode (lines 925-933). This function enforces hardware reality: offload and CPU-only execution paths cannot achieve a Perfect rating. When running in CpuOffload or CpuOnly mode, the system automatically caps any Perfect verdict to Good, reflecting the performance penalty of non-GPU execution.
The final fit_level field in the ModelFit struct (lines 329-334) stores this capped result for downstream consumption.
Runtime-Based FitLevel Capping Explained
The capping mechanism ensures honest reporting of user experience. As implemented in cap_for_run_mode (lines 925-933):
- Gpu and TensorParallel modes preserve Perfect ratings
- MoeOffload, CpuOffload, and CpuOnly modes downgrade Perfect to Good
This distinction matters because the FitLevel directly influences filtering and sorting behaviors throughout the CLI and TUI interfaces. Only GPU-resident execution paths qualify for Perfect ratings, as they alone provide the latency characteristics users expect from local LLM inference.
Practical Code Examples
Accessing the FitLevel after analysis follows this pattern:
// Detect hardware and analyze model fit
let system = SystemSpecs::detect();
let model = llmfit_core::models::LlmModel::load("gemma-2b");
let analysis = ModelFit::analyze(&model, &system);
println!("Fit level: {:?}", analysis.fit_level);
Filtering models by runnability:
// Exclude models that won't fit in memory
let runnable: Vec<_> = all_fits
.into_iter()
.filter(|f| f.fit_level != FitLevel::TooTight)
.collect();
CLI argument handling for fit filtering:
// Handle --fit perfect flag
if args.fit == FitArg::Perfect {
fits.retain(|f| f.fit_level == FitLevel::Perfect);
}
Integration Across the Codebase
The FitLevel enum propagates through multiple layers of the application:
- Core: Defined in
llmfit-core/src/fit.rswithin theModelFitstruct (lines 329-334) - TUI: Rendered with color indicators in
llmfit-tui/src/display.rs(line 256) - API: Exposed via HTTP endpoints in
llmfit-tui/src/serve_api.rs(line 150)
This consistent representation ensures that whether users interact via command line, terminal UI, or programmatic API, the FitLevel determination remains identical and hardware-accurate.
Summary
- llmfit calculates FitLevel through a memory-ratio comparing required versus available memory
- Raw ratios map to categorical thresholds: Perfect (≤0.7), Good (≤1.0), Marginal (≤1.5), and TooTight (>1.5)
- The
cap_for_run_modefunction downgrades Perfect fits to Good when using CPU offload or CPU-only execution paths - Final FitLevel values store in the
ModelFitstruct'sfit_levelfield (lines 329-334 inllmfit-core/src/fit.rs) - The system supports filtering and sorting across CLI, TUI, and API layers based on these determinations
Frequently Asked Questions
What are the exact memory ratio thresholds for each FitLevel?
The thresholds are hardcoded in the pure_ratio_verdict function within llmfit-core/src/fit.rs. A ratio of 0.7 or below yields Perfect, 1.0 or below yields Good, 1.5 or below yields Marginal, and anything exceeding 1.5 returns TooTight. These ratios represent the proportion of memory required relative to memory available on the target system.
Why can't CPU-only models achieve a Perfect FitLevel?
The cap_for_run_mode function (lines 925-933) explicitly caps Perfect ratings to Good for any execution path that cannot run entirely on GPU hardware. This reflects the performance reality that CPU inference, while functional, cannot match the latency and throughput characteristics of GPU-resident execution, regardless of adequate memory availability.
How does llmfit handle unified memory systems like Apple Silicon?
On unified-memory architectures, the system treats GPU and CPU as sharing a single memory pool. The mem_available calculation accounts for this shared resource, while the runtime selection logic (lines 500-507) prioritizes Metal/MLX for Apple Silicon hosts. The FitLevel determination then proceeds normally, though the unified pool typically provides more flexible headroom than discrete VRAM constraints.
Can users filter models by FitLevel in the CLI?
Yes. The CLI accepts fit-level arguments (such as --fit perfect) and filters the Vec<ModelFit> results accordingly. The implementation uses fits.retain() to keep only models matching the requested FitLevel variant, ensuring users see only models that meet their specific performance requirements.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →