How the LLMFIT Run Mode Selection Algorithm Ranks Fit Levels (Perfect, Good, Marginal, Too Tight)
The run mode selection algorithm in LLMFIT ranks fit levels by first rejecting models that exceed available memory as Too Tight, then applying run-mode-specific logic where only GPU and TensorParallel executions can achieve Perfect fit when recommended memory requirements are met, while all other modes are capped at Good or Marginal based on a mandatory 20 percent head-room safety factor.
The LLMFIT library (available at AlexsJones/llmfit) analyzes large language model compatibility with local hardware by assigning discrete fit ratings during the run mode selection process. These ratings determine whether a model will run efficiently, run with constraints, or fail to load entirely based on memory availability and execution path constraints.
The Four Fit Levels in LLMFIT
LLMFIT categorizes hardware-model compatibility into four distinct states:
- Perfect: The model fits comfortably within GPU VRAM and meets the manufacturer's recommended memory budget.
- Good: The model fits with at least 20 percent additional head-room above minimum requirements, but does not meet the recommended threshold.
- Marginal: The model barely fits into the allocated memory pool with less than 20 percent head-room, indicating potential performance degradation.
- Too Tight: The model's minimum memory requirements exceed available resources, making execution impossible.
How the score_fit Algorithm Works
The ranking logic resides in the private function score_fit within [llmfit-core/src/fit.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs#L75). This function evaluates memory constraints after a run mode has been selected, applying a four-step hierarchy to determine the final fit level.
Step 1: Memory Exhaustion Check (Too Tight)
The algorithm first performs an absolute capacity check. If mem_required exceeds mem_available, the function returns immediately with FitLevel::TooTight and terminates further evaluation. This early exit prevents impossible configurations from consuming unnecessary computation cycles.
Step 2: Run Mode Capability Gates
The selected run_mode defines the ceiling for achievable fit levels. The algorithm discriminates between GPU-centric and CPU-dependent execution paths:
- GPU and TensorParallel modes can potentially achieve Perfect fit because they keep the model entirely in high-bandwidth VRAM.
- MoE-offload, CPU-offload, and CPU-only modes are structurally limited to Good or Marginal because they rely on lower-bandwidth CPU memory or partial offloading strategies.
Step 3: Head-Room Threshold Evaluation (20% Rule)
For modes that pass the initial capacity check, the algorithm applies a safety multiplier of 1.2 (20 percent) to the required memory. The logic follows:
- If
mem_availableis greater than or equal tomem_required * 1.2, the model receives a Good rating. - If available memory falls below this threshold, the model is marked Marginal.
This safety factor ensures that operating system overhead and activation spikes do not cause out-of-memory errors during inference.
Step 4: Recommended Memory Upgrade to Perfect
For GPU and TensorParallel modes only, the algorithm performs a final check against the model's recommended memory specification. If the system provides enough memory to satisfy the recommended budget (recommended <= mem_available), the fit level upgrades from Good to Perfect, regardless of the 20 percent head-room calculation. This reflects the assumption that GPU-first systems should accommodate the model's optimal memory profile for peak performance.
Run Mode Selection Determines Maximum Achievable Fit
The architecture of the run mode selection algorithm inherently limits fit levels based on execution strategy. Because Perfect fit requires satisfying both minimum and recommended memory constraints within GPU VRAM, only models executing via RunMode::Gpu or RunMode::TensorParallel can achieve this status. CPU-dependent paths—while potentially viable—are capped at Good even when ample system RAM exists, acknowledging the latency penalties associated with PCIe transfers and CPU inference.
The concrete implementation in score_fit uses pattern matching to enforce these constraints:
fn score_fit(
mem_required: f64,
mem_available: f64,
recommended: f64,
run_mode: RunMode,
) -> FitLevel {
if mem_required > mem_available {
return FitLevel::TooTight;
}
match run_mode {
RunMode::Gpu | RunMode::TensorParallel => {
if recommended <= mem_available {
FitLevel::Perfect
} else if mem_available >= mem_required * 1.2 {
FitLevel::Good
} else {
FitLevel::Marginal
}
}
RunMode::MoeOffload | RunMode::CpuOffload | RunMode::CpuOnly => {
if mem_available >= mem_required * 1.2 {
FitLevel::Good
} else {
FitLevel::Marginal
}
}
}
}
Code Example: Detecting Fit Levels in Practice
Applications can invoke the analysis pipeline to retrieve fit rankings programmatically:
use llmfit_core::fit::{ModelFit, FitLevel};
use llmfit_core::hardware::SystemSpecs;
use llmfit_core::models::LlmModel;
// Load model metadata and detect hardware capabilities
let model: LlmModel = /* ... */;
let system: SystemSpecs = SystemSpecs::detect().unwrap();
// Execute the run mode selection and fit ranking algorithm
let fit: ModelFit = ModelFit::analyze(&model, &system);
// Access the resulting fit level and execution path
println!("Fit level: {}", fit.fit_text()); // "Perfect", "Good", "Marginal", or "TooTight"
println!("Run mode: {}", fit.run_mode_text()); // "GPU", "TensorParallel", "CPU", etc.
// Render UI-friendly indicators
println!("{} {}", fit.fit_emoji(), fit.fit_text());
// Output: 🟢 Perfect
Implementation Files and Architecture
The fit ranking system spans multiple modules across the LLMFIT codebase:
llmfit-core/src/fit.rs– Contains thescore_fitfunction (line 75),ModelFitstruct, andrank_models_by_fit_opts_colsorting logic (line 95).llmfit-core/src/hardware.rs– DefinesSystemSpecswith VRAM detection, unified memory queries, and GPU bandwidth metrics used for run mode selection.llmfit-core/src/models.rs– Stores model metadata includingmin_vram_gb,recommended_ram_gb, and quantization hierarchies fed into the scoring algorithm.llmfit-tui/src/display.rs– Implementsfit_emoji()andfit_text()helpers for terminal rendering.llmfit-tui/src/tui_ui.rs– Presents fit ratings in the interactive table view.
How Fit Rankings Affect Model Sorting
After individual fit levels are calculated, the function rank_models_by_fit_opts_col at [line 95 of fit.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs#L95) reorganizes the results. This sorting logic pushes all FitLevel::TooTight entries to the end of the result list, ensuring that viable configurations (Perfect, Good, or Marginal) appear first in user interfaces and automated selection pipelines.
Summary
- Too Tight assignments occur immediately when
mem_requiredexceedsmem_available, blocking execution. - The 20 percent safety factor (1.2x multiplier) separates Good from Marginal fits across all run modes.
- Perfect ratings require GPU or TensorParallel execution and satisfaction of the model's recommended memory budget.
- CPU-dependent modes (MoE-offload, CPU-offload, CPU-only) are algorithmically restricted from achieving Perfect fit regardless of available system RAM.
- The
score_fitfunction inllmfit-core/src/fit.rsimplements this logic at line 75, whilerank_models_by_fit_opts_colat line 95 prioritizes display order.
Frequently Asked Questions
What determines if a model gets a Perfect fit rating in LLMFIT?
A model achieves Perfect fit only when executing under RunMode::Gpu or RunMode::TensorParallel and the system provides enough memory to satisfy the model's recommended RAM requirement. Merely exceeding the minimum memory by 20 percent results in a Good rating unless the recommended threshold is also met.
Why can't CPU-only or offloaded modes achieve Perfect fit?
The run mode selection algorithm caps CPU-dependent paths (MoE-offload, CPU-offload, and CPU-only) at Good or Marginal because these strategies involve transferring data across the PCIe bus or relying on slower DDR memory. The Perfect designation is reserved for configurations that keep the entire model in high-bandwidth GPU VRAM with optimal head-room.
What is the 20% head-room rule in the fit scoring algorithm?
The 20 percent head-room rule requires that mem_available must be at least 1.2 times (120% of) mem_required to qualify for a Good rating. If available memory falls between 100% and 120% of the requirement, the model receives a Marginal rating. This safety margin accommodates runtime memory fluctuations during inference.
How does LLMFIT handle models that don't fit in available memory?
When mem_required exceeds mem_available, the score_fit function returns FitLevel::TooTight immediately. The sorting function rank_models_by_fit_opts_col subsequently pushes these entries to the end of result lists, ensuring that user interfaces display only viable configurations prominently while still indicating which models are incompatible with current hardware.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →