How Does llmfit Analyze and Score LLM Models? A Deep Dive into the Core Pipeline
llmfit evaluates large language models against detected hardware specifications to generate a ModelFit containing a memory-fit verdict, throughput estimate, and weighted composite score, orchestrated through a 13-step pipeline in the llmfit-core crate.
According to the AlexsJones/llmfit source code, the analysis engine transforms static model metadata and dynamic system capabilities into ranked recommendations using Rust-based hardware detection, runtime selection, and benchmark-calibrated scoring algorithms.
The Three-Phase Analysis Architecture
The ranking pipeline lives entirely within the llmfit-core crate and follows a structured progression from hardware introspection to final score computation. The system first detects system capabilities, filters the embedded model catalog against runtime constraints, then constructs detailed fit profiles through quantization analysis and memory budgeting.
Phase 1: Hardware Detection and Model Catalog Loading
The pipeline initiates with SystemSpecs::detect() in llmfit-core/src/hardware.rs, which gathers RAM, CPU, GPU type, VRAM, and unified-memory information. Simultaneously, ModelDatabase::new() in llmfit-core/src/models.rs reads the embedded hf_models.json file (along with any custom or ONNX entries) to produce a list of LlmModel structures containing metadata, quantization hierarchies, and memory estimates.
Phase 2: Filtering and Fit Construction
Before detailed analysis, the system invokes rankable_models() in llmfit-core/src/analysis.rs (lines 174-190) to discard entries incompatible with the detected backend or those flagged through sanitization checks for draft heads and size mismatches. The build_model_fits() function (lines 241-260) then iterates over these candidates, calling ModelFit::analyze_inner() for each to construct a complete fit list.
Phase 3: Runtime Resolution and Memory Analysis
Inside analyze_inner() in llmfit-core/src/fit.rs (lines 94-108), the system resolves the inference runtime based on platform and model format.
Quantization Selection
The best_quant_for_runtime_budget() function (lines 71-88) walks the appropriate quantization hierarchy (GGUF, ONNX, or MLX) to identify the smallest quantization level that satisfies the memory budget. This selection feeds into execution path determination.
Execution Path and Memory Verdict
The system determines the RunMode—GPU, MoE-offload, CPU-only, or Tensor-Parallel—through cpu_path(), moe_offload_path(), or GPU branch logic in llmfit-core/src/fit.rs (lines 445-514). This calculates mem_required and mem_available, which score_fit() (lines 335-343) evaluates against static thresholds like FIT_PERFECT_MAX_RATIO to produce a fit_level of Perfect, Good, Marginal, or TooTight.
Throughput Estimation and Performance Scoring
Roof-Line Based Speed Prediction
The estimate_tps() function in llmfit-core/src/fit.rs (lines 778-796) calculates throughput using a roof-line model based on GPU memory bandwidth or backend-specific constants. It applies a user-configurable efficiency factor and multiplies by run-mode coefficients (GPU = 1.0, Tensor-Parallel ≈ 0.9) to generate estimated_tps, prefill_tps, and ttft_ms values.
Multi-Dimensional Component Scoring
compute_scores() (lines 828-846) derives four 0-100 component scores: quality, speed, fit, and context, based on model metadata, selected quantization, use-case inference, and throughput estimates. The weighted_score() function (lines 36-47) then multiplies these components by per-use-case weights defined in ScoringWeights to produce the final composite score.
Benchmark Calibration and Confidence Scoring
After initial formula-based estimation, apply_local_calibration() in llmfit-core/src/analysis.rs (lines 262-280) incorporates measurements from LocalBenchIndex and CommunityBenchIndex to adjust the predicted throughput against real-world data. The derive_estimate_confidence() function in llmfit-core/src/fit.rs (lines 302-322) tags each fit as MeasuredLocal, MeasuredCommunity, Calibrated, or Estimated, indicating the reliability of the speed prediction.
The ModelFit Output Structure
The resulting ModelFit structure (defined in llmfit-core/src/fit.rs) contains:
- fit_level: Memory compatibility verdict (Perfect, Good, Marginal, TooTight)
- run_mode: Execution strategy (GPU, MoE-offload, CPU-offload, CPU-only, Tensor-Parallel)
- estimated_tps and prefill_tps/ttft_ms: Throughput figures for inference and prefill stages
- score: A 0-100 composite reflecting quality, speed, memory fit, and context suitability
- notes: Human-readable diagnostics such as "Unified memory" or "MLX runtime ≈ 30% faster"
All callers (CLI, TUI, HTTP API, or MCP) receive an unsorted Vec<ModelFit> that they can organize using rank_models_by_fit_opts_* functions sorted by score, tokens-per-second, or memory utilization.
Implementation Examples
Rust API Usage
use llmfit_core::{
hardware::SystemSpecs,
models::ModelDatabase,
analysis::{InstalledIndex, build_model_fits},
fit::InferenceRuntime,
};
fn main() {
// 1️⃣ Detect the current hardware.
let specs = SystemSpecs::detect().expect("could not detect specs");
// 2️⃣ Load the model catalog.
let db = ModelDatabase::new();
// 3️⃣ Detect which models are already installed in the various providers.
let installed = InstalledIndex::detect_all();
// 4️⃣ Build the fit list (no forced runtime, default context cap).
let fits = build_model_fits(&db, &specs, &installed, None, None);
// 5️⃣ Sort by the composite score (highest first) and print the top‑5.
let mut sorted = llmfit_core::fit::rank_models_by_fit_opts(fits);
sorted.truncate(5);
for f in sorted {
println!(
"{} – {} ({:.1}% RAM) – {:.1} tok/s – score {:.1}",
f.model.name,
f.run_mode_text(),
f.utilization_pct,
f.estimated_tps,
f.score,
);
}
}
Command Line Interface
# Show a table of the best‑scoring models for the current machine
cargo run -- fit --perfect -n 10
Python Integration
import llmfit
# Returns a list of dicts with the same fields as ModelFit
results = llmfit.fit()
for m in results[:5]:
print(f"{m['model']['name']}: {m['score']:.1f} – {m['estimated_tps']:.1f} tok/s")
Summary
- llmfit analyzes LLMs through a 13-step pipeline in
llmfit-corethat compares model requirements against detected hardware specs. - The
SystemSpecs::detect()function inhardware.rscaptures RAM, CPU, GPU, and VRAM data to establish runtime constraints. rankable_models()andbuild_model_fits()inanalysis.rsfilter and construct candidate fits using quantization hierarchies frommodels.rs.- Runtime selection (MLX, llama.cpp, vLLM) and quantization optimization occur in
analyze_inner()withinfit.rs. - Memory-fit verdicts rely on
score_fit()comparing required versus available memory against static thresholds. - Throughput estimation uses
estimate_tps()with roof-line calculations and efficiency factors, calibrated byapply_local_calibration()against local and community benchmarks. - Final ModelFit structures contain composite scores (0-100), run modes, and confidence tags derived from weighted component scoring.
Frequently Asked Questions
How does llmfit determine if a specific model will fit my GPU memory?
The system calculates mem_required and mem_available in analyze_inner() through execution path logic (lines 445-514 in fit.rs), then score_fit() (lines 335-343) evaluates the ratio against static thresholds like FIT_PERFECT_MAX_RATIO to assign a fit_level of Perfect, Good, Marginal, or TooTight.
Which inference runtimes does llmfit support for analysis?
According to the source code in fit.rs (lines 94-108), the system automatically selects MLX for Apple Silicon, llama.cpp for general quantized models, or vLLM for pre-quantized deployments, with the ability to force specific backends through user configuration.
How accurate are llmfit's throughput estimates?
Accuracy depends on the confidence tier assigned by derive_estimate_confidence() (lines 302-322 in fit.rs): MeasuredLocal indicates actual benchmark data from your machine, MeasuredCommunity uses crowd-sourced benchmarks, Calibrated adjusts formula predictions based on available data, and Estimated relies solely on roof-line modeling with efficiency factors.
Can I customize the scoring weights for different use cases?
Yes. The weighted_score() function in fit.rs (lines 36-47) applies per-use-case weights defined in ScoringWeights to the four components (quality, speed, fit, context). Users can configure these weights to prioritize speed over quality or context length over memory efficiency when generating the final 0-100 composite score.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →