When Does llmfit Report a "Good" Fit Level?
A "Good" fit level in llmfit requires at least 20% memory headroom—meaning available memory must be 1.2× or greater than the model's required memory—while falling short of the "Perfect" tier.
The llmfit crate analyzes whether large language models can run on your hardware by assigning one of four FitLevel values: Perfect, Good, Marginal, or Too Tight. This article explains exactly when the tool reports Good and how the underlying logic works according to the AlexsJones/llmfit source code.
How llmfit Determines Fit Level
The core decision happens in llmfit-core/src/fit.rs inside the score_fit function. This function compares three memory figures:
mem_required: Calculated from model size, quantization, and estimated context lengthmem_available: VRAM for GPU runs, system RAM for CPU-only pathsrecommended: The vendor's recommended RAM for optimal performance
The function first rejects models that cannot fit at all. For viable models, it branches by RunMode to assign the appropriate tier.
When "Good" Is Reported for GPU and TensorParallel Modes
For GPU and TensorParallel runs, score_fit checks three conditions in order:
- Perfect:
recommended <= mem_available— the model meets vendor recommendations - Good:
mem_available >= mem_required * 1.2— required memory plus 20% slack - Marginal: Anything tighter than 20% headroom but still fitting
// llmfit-core/src/fit.rs line 44-78
fn score_fit(mem_required: f64, mem_available: f64,
recommended: f64, run_mode: RunMode) -> FitLevel {
if mem_required > mem_available {
return FitLevel::TooTight;
}
match run_mode {
RunMode::Gpu | RunMode::TensorParallel => {
if recommended <= mem_available {
FitLevel::Perfect
} else if mem_available >= mem_required * 1.2 {
FitLevel::Good
} else {
FitLevel::Marginal
}
}
// ... MoE/CPU branches
}
}
Key insight: GPU runs only reach Good when they have sufficient headroom but miss the recommended memory threshold. If your GPU exceeds the vendor recommendation, you get Perfect instead.
When "Good" Is Reported for CPU and Offload Modes
MoE Offload, CPU Offload, and CPU Only modes follow a simplified path:
- Good:
mem_available >= mem_required * 1.2(same 20% rule) - Marginal: Anything tighter that still fits
These modes cannot achieve Perfect regardless of available RAM. As the source comment explains: "Perfect requires GPU acceleration. CPU paths cap at Good."
// llmfit-core/src/fit.rs
RunMode::MoeOffload |
RunMode::CpuOffload |
RunMode::CpuOnly => {
if mem_available >= mem_required * 1.2 {
FitLevel::Good
} else {
FitLevel::Marginal
}
}
This design reflects practical reality: even with abundant system RAM, CPU inference lacks the performance characteristics that vendors target with their "recommended" specifications.
Practical Examples: Checking for "Good" Fit
Query a Model's Fit Level Programmatically
// Detect hardware and analyze a 7B model
let system = SystemSpecs::detect(); // queries RAM, GPU, VRAM
let model = LlmModel::from_name("llama-2-7b");
let fit = ModelFit::analyze(&model, &system);
println!("Fit level: {}", fit.fit_text());
// → "Good" when VRAM is 1.2× required but below recommended
Inspect Why a Model Received "Good"
// Review detailed notes explaining the fit calculation
for note in &fit.notes {
println!("· {}", note);
}
// Example output: "GPU: model loaded into VRAM"
// "VRAM headroom: 23% (Good tier)"
Filter Models by "Good" Fit via CLI
# Terminal: show only models with Good or better fit
$ llmfit fit --good
The FitLevel Enum Definition
The four-tier system is defined in llmfit-core/src/fit.rs at lines 78-83:
pub enum FitLevel {
Perfect, // Recommended memory met on GPU
Good, // Fits with headroom (GPU tight, or CPU comfortable)
Marginal, // Minimum memory met but tight
TooTight, // Does not fit in available memory
}
The Good variant specifically covers two scenarios: GPU runs that work but lack recommended memory, and CPU/offload runs with comfortable headroom.
Summary
- "Good" requires 20% memory headroom:
mem_available >= mem_required × 1.2 - GPU limitation: Good only appears when missing the Perfect threshold (recommended memory)
- CPU ceiling: Offload and CPU-only modes max out at Good—no Perfect tier exists
- Below Good: Less than 20% headroom drops to Marginal; insufficient memory yields Too Tight
Frequently Asked Questions
What is the exact memory multiplier for "Good" fit in llmfit?
"Good" requires available memory to be at least 1.2 times the required memory—a 20% buffer. This multiplier is hardcoded in the score_fit function across all run modes in llmfit-core/src/fit.rs.
Can a CPU-only model ever achieve "Perfect" fit?
No. According to the source code comments in the fit.rs implementation, "Perfect requires GPU acceleration. CPU paths cap at Good." Even with abundant system RAM, CPU and offload modes cannot exceed the Good tier.
Why did my model show "Marginal" instead of "Good" when it fits in VRAM?
"Marginal" indicates insufficient headroom. If your available VRAM is between 100% and 120% of the model's required memory, llmfit reports Marginal rather than Good. The 20% threshold (1.2× multiplier) must be satisfied to advance to the Good tier.
How does quantization affect the "Good" fit calculation?
Quantization reduces mem_required, making "Good" easier to achieve. The mem_required value passed to score_fit already accounts for your selected quantization level. Lower-bit quantization decreases memory needs, which can push a Marginal fit to Good or enable a previously Too Tight model to fit.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →