Understanding llmfit Fit Analysis Dimensions: Quality, Speed, Fit, and Context
llmfit evaluates every model using a four-dimensional scoring system—Quality, Speed, Fit, and Context—that combines into a composite 0-100 score to rank hardware-model compatibility.
The llmfit library (AlexsJones/llmfit) provides hardware-aware model selection by analyzing how well Large Language Models match your system's specifications. At the core of this analysis are four distinct fit analysis dimensions defined in the ScoreComponents struct within llmfit-core/src/fit.rs (lines 199-206). These dimensions quantify everything from raw inference speed to memory utilization efficiency, enabling precise ranking of models for specific hardware configurations.
The Four Dimensions of llmfit Fit Analysis
The compute_scores function (lines 1499-1512 in llmfit-core/src/fit.rs) populates the ScoreComponents struct by calling four specialized helper functions. Each dimension targets a specific aspect of model-hardware alignment.
Quality (Intrinsic Model Capability)
The Quality dimension measures intrinsic model capability through parameter count, model family reputation, quantization penalties, and task-specific alignment. According to the source code, the quality_score() helper calculates this by combining a base quality tier derived from active parameters, applying family-specific bumps, generation/recency bonuses, and quantization penalties. Optional benchmark-derived task bumps further refine this score based on specific use cases.
Speed (Token Throughput)
The Speed dimension estimates token-per-second (TPS) throughput for the selected hardware and run mode. The speed_score() helper scales the raw TPS estimate (estimated_tps) to a 0-100 range, weighting the result by the chosen use-case parameters. This ensures that latency-critical applications prioritize differently than batch processing workloads when evaluating the same hardware configuration.
Fit (Memory Utilization Efficiency)
The Fit dimension evaluates memory-utilization efficiency by measuring how tightly a model’s memory requirements fill the available VRAM or RAM pool. The fit_score() helper computes the ratio of required versus available memory, rewarding models that utilize most of the memory pool without exceeding it. This prevents both under-utilization (wasted resources) and overallocation (out-of-memory errors).
Context (Usable Context Length)
The Context dimension determines usable context length after accounting for weight and KV-cache overhead. The context_score() helper examines the model’s native context window alongside the hardware’s memory headroom, capping the usable tokens accordingly. This ensures the reported context length reflects actual runtime capability rather than theoretical maximums.
Combining Dimensions into a Composite Score
While individual dimensions provide granular insight, llmfit combines them via the weighted_score function into an overall composite Score (0-100) that the UI sorts by. This weighting allows the system to prioritize different aspects based on user preferences—favoring raw speed for real-time applications or maximizing context length for document analysis tasks.
The compute_scores function orchestrates this process by sequentially invoking quality_score(), speed_score(), fit_score(), and context_score(), aggregating their results into the final ScoreComponents struct that powers the ranking algorithm.
Working with Fit Analysis in Rust
You can access these dimensions programmatically after performing a fit analysis:
// Perform a fit analysis on a model for the current system.
let model_fit = ModelFit::analyze(&model, &system_specs);
// Access the four dimensional scores:
let dims = &model_fit.score_components;
println!("Quality: {:.1}", dims.quality);
println!("Speed: {:.1}", dims.speed);
println!("Fit: {:.1}", dims.fit);
println!("Context: {:.1}", dims.context);
To inspect the composite score alongside its constituent dimensions:
// Example: printing the composite score and its dimensions.
println!("Overall score: {:.1}", model_fit.score);
println!(" – Quality: {:.1}", model_fit.score_components.quality);
println!(" – Speed: {:.1}", model_fit.score_components.speed);
println!(" – Fit: {:.1}", model_fit.score_components.fit);
println!(" – Context: {:.1}", model_fit.score_components.context);
Summary
- Quality evaluates intrinsic model capability through parameters, family reputation, and quantization impact using
quality_score(). - Speed estimates TPS throughput via
speed_score(), scaling raw performance estimates to hardware-specific expectations. - Fit measures memory utilization efficiency through
fit_score(), optimizing for full but safe memory usage. - Context calculates usable context length after overhead via
context_score(), ensuring realistic token limits. - All four dimensions are defined in
ScoreComponentsat lines 199-206 ofllmfit-core/src/fit.rsand computed bycompute_scoresat lines 1499-1512.
Frequently Asked Questions
What are the four fit analysis dimensions in llmfit?
The four dimensions are Quality (intrinsic model capability), Speed (token throughput), Fit (memory utilization efficiency), and Context (usable context length after overhead). These are defined as fields in the ScoreComponents struct within llmfit-core/src/fit.rs.
How does llmfit calculate the composite score from individual dimensions?
The compute_scores function calls four helper functions—quality_score(), speed_score(), fit_score(), and context_score()—to populate the ScoreComponents struct. These values are then fed into weighted_score to produce a composite 0-100 score that ranks models by overall hardware compatibility.
Where are the fit analysis dimensions implemented in the llmfit source code?
The dimensions are defined in the ScoreComponents struct at lines 199-206 of llmfit-core/src/fit.rs. The calculation logic resides in individual helper functions within the same file, while compute_scores (lines 1499-1512) orchestrates the scoring process by invoking these helpers.
How does the Fit dimension handle memory constraints?
The Fit dimension uses fit_score() to evaluate the ratio of required memory versus available memory, rewarding models that utilize most of the available VRAM or RAM without exceeding the pool. This prevents both resource under-utilization and out-of-memory errors during inference.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →