Where Local, Community, and Measured Presets Take Precedence in llmfit Benchmarks
The fitting routine prioritizes local measured benchmarks over community presets, falling back to cached community data only when local measurements are unavailable, as implemented in llmfit-core/src/fit.rs.
Understanding local, community, and measured presets precedence is essential when estimating model performance in the AlexsJones/llmfit project. The fitting logic residing in llmfit-core/src/fit.rs enforces a strict hierarchy that prefers user-generated measurements from llmfit bench commands over pre-computed community leaderboard data. This article examines exactly how the precedence logic determines which preset data drives the estimation process.
Understanding the Confidence Hierarchy
The precedence system relies on the EstimateConfidence enum defined in fit.rs, which establishes a clear priority order for data sources during model fitting.
The EstimateConfidence Enum Definition
Located at lines 276 and 287 of llmfit-core/src/fit.rs, this enum codifies the hierarchy:
// Simplified representation of the enum hierarchy (fit.rs#L276-L287)
pub enum EstimateConfidence {
Measured, // Local benchmark data (highest priority)
MeasuredCommunity, // Cached community leaderboard
Formula, // Pure theoretical calculation
}
EstimateConfidence::Measured represents data from a user-run llmfit bench command, while MeasuredCommunity indicates fallback to cached leaderboard entries. Formula serves as the lowest-confidence fallback when no empirical data exists.
How Local Benchmarks Take Precedence
The ModelFit::analyze method and its helper estimate_basis implement the primary check for local measurements. When local throughput data exists, the system bypasses community calibration entirely.
The Local Measurement Check (Lines 3049-3068)
After constructing a fit, the code explicitly checks for local benchmark availability before considering any community data:
// Inside ModelFit::analyze – precedence decision logic (fit.rs#L3049-L3068)
if let Some(local_meas) = self.measured_tps {
// Local benchmark exists → use directly with highest confidence
self.estimate_confidence = EstimateConfidence::Measured;
} else {
// No local measurement → proceed to community fallback
self.estimate_confidence = EstimateConfidence::MeasuredCommunity;
}
This block ensures that local measured benchmarks dominate the fitting process whenever present. If self.measured_tps contains valid throughput data, the routine immediately assigns EstimateConfidence::Measured and skips the community calibration loop.
Community Calibration as Fallback Mechanism
Only when measured_tps remains unset does the estimator invoke the community calibration loop. This secondary process iterates through cached preset labels to load pre-computed leaderboard data from disk.
The Community Calibration Loop (Lines 3145-3149)
The fallback logic appears in the latter sections of fit.rs, executing exclusively when local measurements are absent:
// Community calibration executes only when local measurement is absent (fit.rs#L3145-L3149)
for label in crate::benchmarks::cached_preset_labels() {
let Some(specs) = specs_for_preset_label(label) else { continue };
let Some(resp) = crate::benchmarks::cached_leaderboard_for_preset(label) else { continue };
// Compute est/measured ratio using community data...
}
This loop calls cached_preset_labels() and cached_leaderboard_for_preset() from the benchmarks module, applying MeasuredCommunity confidence only after confirming no local data exists. The iteration over cached_preset_labels at line 3145 and loading via cached_leaderboard_for_preset at line 3149 occurs solely as a secondary estimation path.
Key Source Files and Functions
The precedence logic spans multiple files in the llmfit-core crate:
llmfit-core/src/fit.rs: ContainsModelFit::analyzeandestimate_basisimplementing the priority logic (lines 3049-3068 and 3145-3149).llmfit-core/src/benchmarks.rs: Providescached_preset_labels()andcached_leaderboard_for_preset()for community data retrieval.llmfit-core/src/models.rs: Handles canonical slug lookup to match benchmark rows with model entries.
Summary
- Local benchmarks take absolute precedence: When
measured_tpscontains data from a user-runllmfit benchcommand, the system assignsEstimateConfidence::Measuredand ignores community data. - Community presets serve as fallback: The calibration loop activating at
fit.rs#L3145only executes when local measurements are absent, assigningEstimateConfidence::MeasuredCommunity. - Formula estimates rank lowest: Pure theoretical calculations via
EstimateConfidence::Formulaapply only when neither local nor community data exists. - Hierarchy is enforced in
ModelFit::analyze: The core logic resides inllmfit-core/src/fit.rswith checks at lines 3049-3068 determining which path executes.
Frequently Asked Questions
What happens when both local and community benchmarks exist for the same model?
The local measured benchmarks always win. As implemented in ModelFit::analyze at lines 3049-3068 of fit.rs, the presence of measured_tps immediately sets EstimateConfidence::Measured, causing the fitting routine to skip the community calibration loop entirely regardless of what community data exists on disk.
How does the system determine which confidence level to assign during estimation?
The assignment follows a strict cascade defined in the EstimateConfidence enum. First, the code checks for local measurements at fit.rs#L3049. If absent, it falls back to community calibration at line 3145, assigning MeasuredCommunity. Only if both sources fail does it default to Formula confidence.
Where does llmfit store the community benchmark data used for fallback?
Community benchmarks reside in the cache accessed through llmfit-core/src/benchmarks.rs. The cached_preset_labels() function retrieves available presets, while cached_leaderboard_for_preset() loads the specific leaderboard data for each label, as referenced in the calibration loop at fit.rs#L3149.
Can users force the system to use community presets instead of local benchmarks?
According to the source code in fit.rs, there is no override mechanism exposed in ModelFit::analyze. The precedence logic hardcodes the check for self.measured_tps before considering community data. To use community presets, users must ensure no local benchmark file exists for the target model.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →