How Local, Community, and Measured Benchmark Results Override Throughput Estimates in llmfit

llmfit prioritizes throughput data in a strict hierarchy—user-local measurements override community benchmarks, which in turn override formulaic hardware estimates—ensuring the most trustworthy real-world data always takes precedence.

When estimating model inference performance, llmfit starts with theoretical hardware calculations but immediately defers to empirical evidence when available. This open-source tool implements a three-tier confidence system that merges local, community, and measured benchmark results to override throughput estimates with the most accurate data possible.

The Three-Tier Confidence Hierarchy

The merge order follows the EstimateConfidence enum defined in llmfit-core/src/fit.rs, creating a cascading fallback mechanism that prioritizes real-world measurements over theoretical predictions.

Tier 1: User-Local Measurements (Measured)

Local benchmarks represent the highest-confidence source and always supersede calculated estimates. When you run a benchmark on your specific machine, llmfit stores the result as a measured entry tagged with your exact hardware profile.

According to the source code in llmfit-core/src/analysis.rs, the system treats these user-generated results as "most trustworthy first" and uses them to overwrite any formula-derived throughput numbers. This ensures that your actual observed performance takes precedence over generic hardware calculations.

Tier 2: Community-Submitted Benchmarks (MeasuredCommunity)

When no local measurement exists, llmfit queries the embedded community benchmark JSON (generated from llmfit-core/data/community) for entries matching your specific hardware configuration—CPU, GPU, RAM, and other relevant specs.

The system calculates the median of matching community submissions and assigns the EstimateConfidence::MeasuredCommunity confidence level, defined in llmfit-core/src/fit.rs (lines 260-287). This community median supersedes the default formulaic estimate but yields priority to any user-local measurements.

Tier 3: Formulaic Hardware Estimates (Estimated)

If neither local nor community data exists for your hardware profile, llmfit falls back to calculated throughput based on your hardware's peak fp16 matmul throughput and memory bandwidth. This baseline "estimated" confidence provides a theoretical prediction when empirical data is unavailable.

Implementation in the Source Code

The actual merging logic resides in three core modules that orchestrate the selection process.

The EstimateConfidence Enum

In llmfit-core/src/fit.rs, the EstimateConfidence enum establishes the precedence order:

// Simplified representation of the hierarchy
Measured -> MeasuredCommunity -> Estimated

The estimate_basis field is selected based on the provenance of available measurements, with the enum variants explicitly defining which source takes priority.

Benchmark Lookup and Analysis

The llmfit-core/src/benchmarks.rs module (lines 399-421) supplies the lookup functions that retrieve community results for a given hardware specification. Meanwhile, llmfit-core/src/analysis.rs orchestrates the final selection, implementing the "most trustworthy first" logic that drives the override behavior.

When processing a fit request, the analysis code checks for measured data in order of confidence, ensuring that empirical results always dominate over theoretical predictions.

Practical Usage Examples

You can control which data sources llmfit considers when running performance estimates:


# Forces use of local benchmark data (highest priority)

cargo run -- fit --model "meta-llama/Meta-Llama-3.1-8B-Instruct" --local-bench

# Explicitly requests community median for matching hardware

cargo run -- fit --model "meta-llama/Meta-Llama-3.1-8B-Instruct" --community

# Default behavior: checks local, then community, then formulaic

cargo run -- fit --model "meta-llama/Meta-Llama-3.1-8B-Instruct"

Summary

  • Local measurements always win: User-run benchmarks on the exact machine receive Measured confidence and override all other sources.
  • Community data fills gaps: When local data is absent, matching community submissions provide MeasuredCommunity confidence as a reliable secondary source.
  • Formulaic estimates are the fallback: Hardware-based calculations only activate when no empirical data exists for the specific hardware profile.
  • Source implementation: The hierarchy is enforced by the EstimateConfidence enum in llmfit-core/src/fit.rs and orchestrated through llmfit-core/src/analysis.rs.

Frequently Asked Questions

How does llmfit determine which benchmark source to use?

llmfit checks sources in strict order of trustworthiness. It first looks for user-local measurements stored in the system. If none exist, it queries the embedded community benchmark JSON for hardware-matching entries. Only if both sources return empty does it fall back to calculated hardware estimates based on memory bandwidth and fp16 throughput.

What is the EstimateConfidence enum and where is it defined?

The EstimateConfidence enum is defined in llmfit-core/src/fit.rs (lines 260-287) and establishes three confidence levels: Measured (local benchmarks), MeasuredCommunity (median of community submissions), and Estimated (formulaic calculations). This enum drives the merge logic by explicitly ranking data source reliability.

Can I force llmfit to ignore my local benchmarks and use community data?

Yes. While llmfit automatically prioritizes local measurements as the highest-confidence source, you can override this behavior using CLI flags. Running with --community explicitly requests the community median for your hardware profile, bypassing local measurements if you suspect they are anomalous or want to compare against aggregate community performance.

Where does llmfit store community benchmark data?

Community benchmarks are embedded as JSON within the llmfit-core/data/community directory and compiled into the binary. The llmfit-core/src/benchmarks.rs module handles loading this data and providing lookup functions that match your current hardware specifications against the community dataset, retrieving median throughput values for matching configurations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →