Where Local, Community, and Measured Presets Take Precedence in llmfit Benchmarks

The fitting routine prioritizes local measured benchmarks over community presets, falling back to cached community data only when local measurements are unavailable, as implemented in llmfit-core/src/fit.rs.

Understanding local, community, and measured presets precedence is essential when estimating model performance in the AlexsJones/llmfit project. The fitting logic residing in llmfit-core/src/fit.rs enforces a strict hierarchy that prefers user-generated measurements from llmfit bench commands over pre-computed community leaderboard data. This article examines exactly how the precedence logic determines which preset data drives the estimation process.

Understanding the Confidence Hierarchy

The precedence system relies on the EstimateConfidence enum defined in fit.rs, which establishes a clear priority order for data sources during model fitting.

The EstimateConfidence Enum Definition

Located at lines 276 and 287 of llmfit-core/src/fit.rs, this enum codifies the hierarchy:

// Simplified representation of the enum hierarchy (fit.rs#L276-L287)
pub enum EstimateConfidence {
    Measured,         // Local benchmark data (highest priority)
    MeasuredCommunity, // Cached community leaderboard
    Formula,          // Pure theoretical calculation
}

EstimateConfidence::Measured represents data from a user-run llmfit bench command, while MeasuredCommunity indicates fallback to cached leaderboard entries. Formula serves as the lowest-confidence fallback when no empirical data exists.

How Local Benchmarks Take Precedence

The ModelFit::analyze method and its helper estimate_basis implement the primary check for local measurements. When local throughput data exists, the system bypasses community calibration entirely.

The Local Measurement Check (Lines 3049-3068)

After constructing a fit, the code explicitly checks for local benchmark availability before considering any community data:

// Inside ModelFit::analyze – precedence decision logic (fit.rs#L3049-L3068)
if let Some(local_meas) = self.measured_tps {
    // Local benchmark exists → use directly with highest confidence
    self.estimate_confidence = EstimateConfidence::Measured;
} else {
    // No local measurement → proceed to community fallback
    self.estimate_confidence = EstimateConfidence::MeasuredCommunity;
}

This block ensures that local measured benchmarks dominate the fitting process whenever present. If self.measured_tps contains valid throughput data, the routine immediately assigns EstimateConfidence::Measured and skips the community calibration loop.

Community Calibration as Fallback Mechanism

Only when measured_tps remains unset does the estimator invoke the community calibration loop. This secondary process iterates through cached preset labels to load pre-computed leaderboard data from disk.

The Community Calibration Loop (Lines 3145-3149)

The fallback logic appears in the latter sections of fit.rs, executing exclusively when local measurements are absent:

// Community calibration executes only when local measurement is absent (fit.rs#L3145-L3149)
for label in crate::benchmarks::cached_preset_labels() {
    let Some(specs) = specs_for_preset_label(label) else { continue };
    let Some(resp) = crate::benchmarks::cached_leaderboard_for_preset(label) else { continue };
    // Compute est/measured ratio using community data...
}

This loop calls cached_preset_labels() and cached_leaderboard_for_preset() from the benchmarks module, applying MeasuredCommunity confidence only after confirming no local data exists. The iteration over cached_preset_labels at line 3145 and loading via cached_leaderboard_for_preset at line 3149 occurs solely as a secondary estimation path.

Key Source Files and Functions

The precedence logic spans multiple files in the llmfit-core crate:

Summary

  • Local benchmarks take absolute precedence: When measured_tps contains data from a user-run llmfit bench command, the system assigns EstimateConfidence::Measured and ignores community data.
  • Community presets serve as fallback: The calibration loop activating at fit.rs#L3145 only executes when local measurements are absent, assigning EstimateConfidence::MeasuredCommunity.
  • Formula estimates rank lowest: Pure theoretical calculations via EstimateConfidence::Formula apply only when neither local nor community data exists.
  • Hierarchy is enforced in ModelFit::analyze: The core logic resides in llmfit-core/src/fit.rs with checks at lines 3049-3068 determining which path executes.

Frequently Asked Questions

What happens when both local and community benchmarks exist for the same model?

The local measured benchmarks always win. As implemented in ModelFit::analyze at lines 3049-3068 of fit.rs, the presence of measured_tps immediately sets EstimateConfidence::Measured, causing the fitting routine to skip the community calibration loop entirely regardless of what community data exists on disk.

How does the system determine which confidence level to assign during estimation?

The assignment follows a strict cascade defined in the EstimateConfidence enum. First, the code checks for local measurements at fit.rs#L3049. If absent, it falls back to community calibration at line 3145, assigning MeasuredCommunity. Only if both sources fail does it default to Formula confidence.

Where does llmfit store the community benchmark data used for fallback?

Community benchmarks reside in the cache accessed through llmfit-core/src/benchmarks.rs. The cached_preset_labels() function retrieves available presets, while cached_leaderboard_for_preset() loads the specific leaderboard data for each label, as referenced in the calibration loop at fit.rs#L3149.

Can users force the system to use community presets instead of local benchmarks?

According to the source code in fit.rs, there is no override mechanism exposed in ModelFit::analyze. The precedence logic hardcodes the check for self.measured_tps before considering community data. To use community presets, users must ensure no local benchmark file exists for the target model.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →