How llmfit Calculates Scoring Components: Quality, Speed, Fit, and Context Explained
The llmfit library evaluates every model-hardware pairing using four independent metrics—quality, speed, fit, and context—each normalized to a 0–100 scale, then combines them via use-case-specific weights into a final composite score.
The llmfit project provides intelligent model fitting for Large Language Models by analyzing hardware constraints against model requirements. Its scoring system breaks down compatibility into four distinct components that quantify intrinsic model merit, inference throughput, memory efficiency, and context window utilization. Understanding how these values are computed allows developers to interpret rankings and customize evaluations for specific deployment scenarios.
The Four Scoring Components in llmfit
The scoring architecture in llmfit-core/src/fit.rs isolates four independent factors. Each component function returns a normalized value clamped between 0 and 100.
Quality Score (Intrinsic Model Merit)
The quality score measures intrinsic model merit independent of hardware. The quality_score() function (lines 15–34 of llmfit-core/src/fit.rs) aggregates multiple metadata signals:
- Base tier: Derived from the active parameter count.
- Family bump: Reputation adjustment based on model family.
- Generation-age bonus: Newer architectures receive higher weight.
- Recency bonus: Rewards recently released checkpoints.
- Quantization penalty: Applied via
models::quant_quality_penaltyfor quantized weights. - Task alignment: Added from
task_bench::scorebased on benchmark performance for the target task.
The final value is clamped to the 0–100 range before returning.
Speed Score (Throughput Normalization)
The speed score estimates inference throughput relative to use-case demands. The speed_score(tps, use_case) function (lines 84–90) divides the estimated tokens-per-second (TPS) by a target specific to the workload:
- 40 TPS: Default target for general chat and completion tasks.
- 25 TPS: Target for reasoning-intensive workloads.
- 200 TPS: Target for embedding and retrieval tasks.
The result is scaled to a 0–100 range and clamped, meaning a model generating 80 TPS against a 40 TPS target achieves a speed score of 100.
Fit Score (Memory Utilization)
The fit score quantifies memory efficiency by comparing model requirements against available resources. The fit_score(mem_required, mem_available) function (lines 108–112) calculates utilization as:
util = (mem_required / mem_available) * 100.0
Values exceeding 100 are clamped, ensuring models that exceed the memory budget receive a low fit score. A model using exactly half of available VRAM or RAM scores 50.
Context Score (Context Window Utilization)
The context score evaluates whether the hardware can support the model’s advertised context window. The context_score(model, use_case) function (lines 114–118) compares the effective context length (used for memory estimation) against the model’s native maximum context size. The ratio is scaled to 0–100 and clamped, indicating how much of the full context window is usable given current constraints.
Computing the Composite Score
The four components are assembled into a final ranking through a weighted combination. The compute_scores() method (called at line 620 of fit.rs) invokes all four scoring functions, then passes the results to weighted_score() (lines 172–176).
Per-use-case weights are stored in CalcConfig::scoring_weights (default definitions at lines 86–92). For example, a "coding" use-case might prioritize speed over quality, while a "reasoning" use-case weights quality and context higher. The weighted_score() function multiplies each component by its corresponding weight and sums the products to produce the final 0–100 composite score.
Inspecting Scores Programmatically
You can access these calculations directly through the ModelFit API. The following example detects system specifications, analyzes a model, and prints the component breakdown:
use llmfit_core::fit::{ModelFit, CalcConfig};
use llmfit_core::hardware::SystemSpecs;
use llmfit_core::models::LlmModel;
// 1. Detect system specs (RAM, GPU, etc.)
let system = SystemSpecs::detect()?;
// 2. Load a model definition from the embedded catalog
let model: LlmModel = llmfit_core::models::load_model("meta-llama/Meta-Llama-3-8B-Instruct")?;
// 3. Run the analysis with default configuration
let fit = ModelFit::analyze(&model, &system);
// 4. Inspect the composite score and the four components
println!("Composite score: {:.1}", fit.score);
println!(" Quality: {:.1}", fit.score_components.quality);
println!(" Speed: {:.1}", fit.score_components.speed);
println!(" Fit: {:.1}", fit.score_components.fit);
println!(" Context: {:.1}", fit.score_components.context);
The score_components struct provides direct access to the individual values calculated by quality_score(), speed_score(), fit_score(), and context_score().
Customizing Component Weights
To override the default importance of each component for a specific workload, modify the CalcConfig before analysis:
let mut cfg = CalcConfig::default();
// Index 1 corresponds to the "Coding" use-case
// Weights are [quality, speed, fit, context]
cfg.scoring_weights.weights[1] = [0.30, 0.40, 0.15, 0.15];
let fit = ModelFit::analyze_with_config(&model, &system, cfg);
This configuration increases the speed weight to 40% and reduces quality to 30% for coding tasks, favoring faster inference over raw model capability.
Summary
- Quality measures intrinsic model value using parameters, family reputation, quantization penalties, and benchmark scores.
- Speed normalizes estimated throughput against use-case targets (40, 25, or 200 TPS).
- Fit calculates memory utilization percentage, clamped when requirements exceed available hardware.
- Context scores usable context length relative to the model’s native maximum.
- All components are computed in
llmfit-core/src/fit.rsand combined viaweighted_score()using per-use-case weights fromCalcConfig.
Frequently Asked Questions
How does llmfit handle models that exceed available memory?
The fit_score() function in llmfit-core/src/fit.rs calculates memory utilization as a percentage of available resources. If mem_required exceeds mem_available, the resulting value is clamped above 100, which typically drives the composite score down significantly when weights are applied. This ensures memory-hungry models rank lower unless other components are heavily prioritized.
Can I adjust the importance of speed versus quality for specific use cases?
Yes. The CalcConfig::scoring_weights array allows per-use-case customization of the four component weights. Modify the weight arrays before calling ModelFit::analyze_with_config() to increase or decrease the relative importance of speed, quality, fit, or context for any supported use-case enum value.
What quantization penalties affect the quality score?
The quality_score() function applies adjustments via models::quant_quality_penalty(), which reduces the score based on the quantization level (e.g., INT4, INT8) relative to full-precision weights. This penalty accounts for the intrinsic accuracy loss associated with compressed model formats.
Where does llmfit get the task-specific benchmark data for quality scoring?
Task alignment bonuses are retrieved from task_bench::score() in llmfit-core/src/task_bench.rs. This module provides benchmark scores that quantify model performance on specific tasks (coding, reasoning, etc.), which quality_score() incorporates as a final adjustment to the intrinsic quality metric.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →