How llmfit Determines the Best Fit Level: Perfect, Good, Marginal, or Too Tight

llmfit determines the best fit level by comparing memory requirements against available resources through a two-stage process that selects a hardware execution path and scores memory headroom using the score_fit function in llmfit-core/src/fit.rs.

The AlexsJones/llmfit repository determines the best fit level for large language models through a deterministic compatibility analysis system implemented in its core Rust library. By evaluating hardware specifications against model metadata, the framework assigns one of four FitLevel classifications that indicate whether a model will run efficiently on the target system.

The Two-Stage Fit Determination Process

Stage 1: Hardware Detection and RunMode Selection

The analysis begins in ModelFit::analyze_inner (lines 697-770 of llmfit-core/src/fit.rs), where the system interrogates the hardware configuration through SystemSpecs. Based on GPU availability, unified memory status, and MoE-specific requirements, the code selects a RunMode from the following options:

  • Gpu or TensorParallel for GPU-accelerated execution
  • MoeOffload, CpuOffload, or CpuOnly for CPU-bound or hybrid execution

For each mode, the system calculates mem_required (the memory needed for the model considering quantization and context) and identifies mem_available from the appropriate memory pool (VRAM, system RAM, or a combination).

Stage 2: Scoring Memory Headroom with score_fit

Once the execution path is established, the score_fit function (lines 744-795) performs the actual fit evaluation. This function receives mem_required, mem_available, the model's recommended_ram_gb (typically roughly twice the model size), and the selected run_mode to calculate the final classification.

Fit Level Classification Rules and Thresholds

The classification logic implemented in the match block (lines 749-794) applies the following deterministic rules:

Too Tight: Assigned immediately when mem_required > mem_available, indicating the model cannot fit in the selected memory pool regardless of execution mode.

Perfect: Exclusive to Gpu and TensorParallel modes. Requires that recommended_ram_gb <= mem_available, ensuring the system's recommended memory comfortably fits with ample headroom. This rating is unattainable on CPU-only paths.

Good: Requires at least a 20% safety margin, calculated as mem_available >= mem_required * 1.2. Available for GPU modes when Perfect criteria are not met, and serving as the maximum achievable rating for CPU modes (MoeOffload, CpuOffload, CpuOnly).

Marginal: Assigned when the model fits (mem_required <= mem_available) but lacks the 20% buffer required for Good, indicating functional but constrained execution with limited headroom.

Source Code Implementation Details

The fit level determination is implemented across several key locations in the llmfit-core crate:

  • llmfit-core/src/fit.rs (lines 749-794): Contains the score_fit match block that implements the classification logic based on RunMode and memory ratios.
  • llmfit-core/src/fit.rs (lines 176-222): Defines fit_text() and fit_emoji() helper methods that expose the FitLevel enum values as human-readable strings and colored UI indicators.
  • llmfit-core/src/hardware.rs: Supplies SystemSpecs and GPU detection utilities used during the initial analysis phase.
  • llmfit-core/src/models.rs: Provides model metadata including min_vram_gb, min_ram_gb, and quantization parameters necessary for memory estimation.

Programmatically Checking Model Compatibility

You can programmatically determine fit levels using the following Rust pattern:

use llmfit_core::fit::ModelFit;
use llmfit_core::hardware::SystemSpecs;
use llmfit_core::models::LlmModel;

// Detect system specs from the local machine
let system = SystemSpecs::detect()?;

// Load a model from the embedded catalog
let model = LlmModel::from_name("mistralai/Mistral-7B-Instruct-v0.1")?;

// Execute the full compatibility analysis
let analysis = ModelFit::analyze(&model, &system);

// Inspect the determined fit level
println!("Fit level: {}", analysis.fit_text());     // → "Perfect", "Good", "Marginal", or "Too Tight"
println!("Emoji: {}", analysis.fit_emoji());        // → "🟢", "🟡", "🟠", or "🔴"
println!("Run mode: {}", analysis.run_mode_text()); // → "GPU", "CPU Offload", etc.

On systems without GPU acceleration, the analysis will never return Perfect because the score_fit logic restricts this rating to GPU-centric RunMode variants according to the source code rules.

Summary

  • llmfit uses a two-stage process in ModelFit::analyze_inner to first select a RunMode and then score memory headroom via score_fit in llmfit-core/src/fit.rs.
  • Perfect ratings require GPU acceleration (Gpu or TensorParallel mode) and sufficient space for the model's recommended_ram_gb.
  • Good ratings require a minimum 20% memory buffer (mem_available >= 1.2 * mem_required).
  • Too Tight indicates the model cannot fit in available memory, while Marginal indicates a fit without adequate safety margins.
  • Helper methods fit_text() and fit_emoji() provide user-facing representations of the fit level at lines 176-222.

Frequently Asked Questions

What causes llmfit to return "Too Tight" for a model?

The "Too Tight" classification occurs in score_fit when mem_required exceeds mem_available for the selected memory pool. This determination happens immediately in the scoring logic at llmfit-core/src/fit.rs regardless of the selected RunMode, and indicates that the model cannot be loaded into the available VRAM or system RAM.

Why can't CPU-only execution achieve a "Perfect" fit rating?

According to the source code logic in lines 749-794, the Perfect variant is exclusive to the Gpu and TensorParallel arms of the score_fit match statement. The implementation explicitly reserves this rating for GPU-accelerated paths, meaning CpuOnly, CpuOffload, and MoeOffload modes can only achieve "Good" as their maximum rating even with abundant system memory.

How much memory headroom does llmfit require for a "Good" rating?

The score_fit function requires a 20% safety margin above the required memory, implementing the check as mem_available >= mem_required * 1.2. This threshold applies uniformly across all RunMode variants when determining whether to upgrade a fit classification from "Marginal" to "Good".

Where is the fit level logic implemented in the llmfit codebase?

The core classification algorithm resides in the score_fit function within llmfit-core/src/fit.rs (specifically lines 744-795). The entry point for the full analysis is ModelFit::analyze_inner (lines 697-770), while helper methods for text and emoji representation appear at lines 176-222 in the same file.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →