How Default Scoring Weights Vary Across Use Cases in llmfit
llmfit evaluates models using four weighted components—quality, speed, fit, and context—with specific default scoring weights that shift based on the use case, ranging from quality-focused Reasoning (0.55) to speed-optimized Embedding (0.40).
The llmfit framework calculates composite model scores through a weighted combination of four distinct metrics. Each default scoring weight is tuned to reflect the priorities of specific workloads, ensuring that a model's suitability is measured against the demands of its intended application rather than generic benchmarks.
The Four Score Components
Every evaluation in llmfit measures quality, speed, fit, and context as the foundational axes of performance. These components are combined using a weighting matrix defined in llmfit-core/src/fit.rs within the ScoringWeights struct's default implementation.
The final composite score is calculated using the formula:
raw = quality * wq + speed * ws + fit * wf + context * wc
final = round(raw * 10) / 10
Where wq, ws, wf, and wc represent the respective default scoring weights retrieved for the active use case.
Use-Case Specific Weight Matrices
The UseCase enum in llmfit-core/src/models.rs determines which weight tuple is retrieved when scoring a model. The enumeration order—General, Coding, Reasoning, Chat, Multimodal, Embedding—maps directly to rows in the weight matrix hardcoded in the ScoringWeights defaults.
| Use Case | Quality Weight | Speed Weight | Fit Weight | Context Weight |
|---|---|---|---|---|
| General | 0.45 | 0.30 | 0.15 | 0.10 |
| Coding | 0.50 | 0.20 | 0.15 | 0.15 |
| Reasoning | 0.55 | 0.15 | 0.15 | 0.15 |
| Chat | 0.40 | 0.35 | 0.15 | 0.10 |
| Multimodal | 0.50 | 0.20 | 0.15 | 0.15 |
| Embedding | 0.30 | 0.40 | 0.20 | 0.10 |
Quality-Heavy Workloads
Reasoning tasks assign the highest quality weight (0.55) while minimizing speed priority (0.15), reflecting the computational depth required for logical inference and complex problem-solving where accuracy outweighs latency.
Speed-Critical Applications
Embedding workloads invert this priority, allocating 0.40 to speed and only 0.30 to quality. This configuration acknowledges that vector generation pipelines prioritize throughput and batch processing speed over generative excellence.
Balanced Scenarios
Chat applications distribute weights more evenly between quality (0.40) and speed (0.35), optimizing for responsive conversational interfaces without sacrificing coherence. Meanwhile, Coding and Multimodal tasks share identical weight profiles (0.50 quality, 0.20 speed), emphasizing output precision while maintaining moderate latency tolerance.
How Scoring Works in Practice
When weighted_score() processes a model evaluation, it retrieves the appropriate weight tuple via ScoringWeights::get() based on the UseCase variant passed to it. This lookup occurs in llmfit-core/src/fit.rs, where the function matches the use case index to the corresponding row in the default matrix.
The weighted calculation produces a raw score that is then rounded to one decimal place to generate the final composite metric.
Implementing Custom Evaluations
The following Rust example demonstrates how to access these default scoring weights and compute a composite score for the Coding use case according to the llmfit source code:
// Retrieve the default scoring weights
let cfg = llmfit_core::fit::CalcConfig::default();
let use_case = llmfit_core::models::UseCase::Coding;
// Get the four weights for the Coding use-case
let (quality_w, speed_w, fit_w, context_w) = cfg.scoring_weights.get(use_case);
println!("Coding weights: q={:.2}, s={:.2}, f={:.2}, c={:.2}",
quality_w, speed_w, fit_w, context_w);
// Example: compute a composite score for a model
let comps = llmfit_core::fit::ScoreComponents {
quality: 80.0,
speed: 70.0,
fit: 90.0,
context: 60.0,
};
let composite = llmfit_core::fit::weighted_score(comps, use_case, &cfg);
println!("Composite score for Coding: {}", composite);
This implementation references the specific file paths llmfit-core/src/fit.rs for the scoring logic and llmfit-core/src/models.rs for the use case enumeration that drives the weight selection.
Summary
- llmfit employs four core metrics—quality, speed, fit, and context—with weights varying by workload type.
- The ScoringWeights struct in
llmfit-core/src/fit.rsdefines the default matrix indexed by the UseCase enum fromllmfit-core/src/models.rs. - Reasoning tasks prioritize quality (0.55) over speed (0.15), while Embedding tasks prioritize speed (0.40) over quality (0.30).
- The
weighted_score()function applies these weights using the formularound((quality * wq + speed * ws + fit * wf + context * wc) * 10) / 10.
Frequently Asked Questions
How does llmfit determine which scoring weights to use?
The framework matches the UseCase enum variant passed to weighted_score() against the ordered rows in the ScoringWeights default implementation. Each variant—General, Coding, Reasoning, Chat, Multimodal, or Embedding—retrieves its specific four-tuple of weights from the matrix defined in llmfit-core/src/fit.rs.
Can the default scoring weights be overridden in llmfit?
The analysis focuses on the default implementations within the ScoringWeights struct. While the CalcConfig struct holds these weights and is passed to scoring functions, customizing them would require modifying the configuration before evaluation or extending the core implementation found in llmfit-core/src/fit.rs.
Why does the Embedding use case prioritize speed over quality?
Embedding models typically process large volumes of text into vector representations where throughput and latency significantly impact pipeline performance. The weight distribution (0.40 speed, 0.30 quality, 0.20 fit, 0.10 context) reflects production requirements for rapid vector generation rather than high-fidelity generative output.
What is the significance of the rounding operation in the scoring formula?
The final calculation round(raw * 10) / 10 standardizes all composite scores to one decimal place, ensuring consistent precision across different use cases and preventing floating-point artifacts from affecting model rankings or threshold comparisons.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →