LLMFIT Scoring Weights Per Use Case: How Quality, Speed, Fit, and Context Are Weighted
LLMFIT calculates composite scores using a configurable 6×4 weight matrix defined in ScoringWeights that assigns different importance to quality, speed, fit, and context depending on whether the use case is General, Coding, Chat, Embedding, Rerank, or Vision.
The AlexsJones/llmfit repository implements a sophisticated scoring system that evaluates large language models across four key dimensions. Understanding these scoring weights per use case is essential for selecting the right model for your specific deployment scenario, as the framework automatically adjusts its priorities based on whether you are running inference for coding assistants, chatbots, or embedding pipelines.
The Four Scoring Dimensions
Before examining the weights, it is important to understand what each metric measures according to the source code implementation:
- Quality: How well the model's responses match a reference standard, measured by the quality-test suite implemented in
llmfit-core/src/quality.rs. - Speed: Normalized throughput (tokens per second) after accounting for the model's runtime overhead and hardware constraints.
- Fit: How comfortably the model fits within the detected hardware resources, specifically RAM and VRAM utilization limits.
- Context: The model's native context window, capped by any runtime-imposed limits and normalized against the maximum supported size.
These four metrics feed into the composite score calculation defined in llmfit-core/src/fit.rs.
The ScoringWeights Matrix
The per-use-case weighting is encoded in the ScoringWeights struct, which contains a static 6×4 matrix - one row for each UseCase enum variant and one column for each metric. According to the source code in llmfit-core/src/fit.rs:86-94, the default implementation populates the matrix with the following values:
| Use Case | Quality Weight | Speed Weight | Fit Weight | Context Weight |
|---|---|---|---|---|
| General | 1.0 | 1.0 | 1.0 | 1.0 |
| Coding | 1.0 | 0.8 | 1.0 | 0.9 |
| Chat | 1.0 | 1.2 | 1.0 | 1.0 |
| Embedding | 0.8 | 1.0 | 0.7 | 1.0 |
| Rerank | 0.9 | 1.0 | 1.0 | 1.0 |
| Vision | 1.0 | 0.9 | 1.0 | 1.0 |
The struct definition in fit.rs declares this as a fixed-size array:
pub struct ScoringWeights {
/// (quality, speed, fit, context) per use case
pub weights: [[f64; 4]; 6],
}
The order of use cases in the matrix corresponds to the UseCase enum definition, allowing the get(use_case) method to return the appropriate four-tuple of weights for any specific deployment scenario.
How Weights Are Applied
When LLMFIT computes a composite score for a particular model and use case, it follows a specific calculation pipeline implemented in the weighted_score function:
-
Collect raw metrics into a
ScoreComponentsstruct containingquality,speed_norm,fit_norm, andcontext_normvalues. -
Retrieve weights for the requested use case via
config.scoring_weights.get(use_case), which returns a tuple(wq, ws, wf, wc). -
Calculate the weighted sum using the formula found in
llmfit-core/src/fit.rs:
let (wq, ws, wf, wc) = config.scoring_weights.get(use_case);
let score = (quality * wq
+ speed_norm * ws
+ fit_norm * wf
+ context_norm * wc)
/ (wq + ws + wf + wc);
-
Normalize the result by dividing by the sum of weights, ensuring the final score always falls within the 0-100 range regardless of the weight configuration.
-
Clamp to context limits so a model cannot receive credit for context lengths larger than its native support.
This weighting system explains why Chat models receive a 1.2× multiplier on speed (prioritizing throughput for conversational interfaces) while Embedding models reduce both quality and fit weights to 0.8 and 0.7 respectively (favoring batch throughput over perfect accuracy for vector generation).
Customizing Weights for Your Deployment
The weighting matrix is deliberately configurable. Projects embedding LLMFIT as a library can replace ScoringWeights::default() with custom values to prioritize specific metrics for their infrastructure requirements.
To override the default scoring weights per use case, construct a custom CalcConfig:
use llmfit_core::fit::{CalcConfig, ScoringWeights, UseCase};
fn configure_for_latency() {
let mut cfg = CalcConfig::default();
cfg.scoring_weights = ScoringWeights {
weights: [
// General
[1.0, 1.0, 1.0, 1.0],
// Coding - boost speed for IDE autocomplete
[1.0, 1.5, 1.0, 0.9],
// Chat
[1.0, 1.2, 1.0, 1.0],
// Embedding
[0.8, 1.0, 0.7, 1.0],
// Rerank
[0.9, 1.0, 1.0, 1.0],
// Vision
[1.0, 0.9, 1.0, 1.0],
],
};
// Use cfg when calculating scores
let components = llmfit_core::fit::ScoreComponents {
quality: 85.0,
speed_norm: 0.78,
fit_norm: 0.92,
context_norm: 0.95,
};
let score = llmfit_core::fit::weighted_score(
components,
UseCase::Coding,
&cfg,
);
}
You can also access specific use case weights programmatically to display them in your own UI or logging system:
let weights = ScoringWeights::default();
let (wq, ws, wf, wc) = weights.get(UseCase::Embedding);
println!("Embedding weights: Q={:.1}, S={:.1}, F={:.1}, C={:.1}", wq, ws, wf, wc);
Summary
- LLMFIT evaluates models using four metrics: Quality, Speed, Fit, and Context, each normalized before weighting.
- The
ScoringWeightsstruct inllmfit-core/src/fit.rsstores a 6×4 matrix mapping eachUseCasevariant to its specific weight tuple. - Default weights prioritize speed for Chat (1.2×) and reduce quality emphasis for Embedding (0.8×) to match typical deployment constraints.
- The final composite score is calculated as a normalized weighted sum, ensuring results remain comparable across different weight configurations.
- Developers can override default weights by providing a custom
ScoringWeightsinstance toCalcConfigfor specialized hardware or latency requirements.
Frequently Asked Questions
How do I override the default scoring weights in LLMFIT?
Create a custom ScoringWeights struct with your desired 6×4 matrix and assign it to CalcConfig.scoring_weights before calling weighted_score(). The matrix must maintain the same use case ordering as the UseCase enum (General, Coding, Chat, Embedding, Rerank, Vision).
What is the difference between the Fit and Context metrics?
Fit measures how well the model's memory footprint matches your available hardware (RAM/VRAM), while Context measures the model's native context window length normalized against maximum supported values. A model might have excellent context support (high Context score) but poor Fit if it barely fits in your GPU memory.
Which use case prioritizes speed over quality?
The Chat use case applies a 1.2× multiplier to speed while keeping quality at 1.0×, reflecting the interactive nature of conversational AI where latency directly impacts user experience. Conversely, Coding maintains full quality weight (1.0×) but reduces speed importance to 0.8× to prioritize correctness over raw throughput.
Are the scoring weights normalized in the final calculation?
Yes. The weighted_score function divides the weighted sum by the sum of the four weights (wq + ws + wf + wc), ensuring the final score remains within a consistent 0-100 range regardless of whether you use the default weights or custom values that sum to a different total.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →