# How benchmarks.rs Creates a Measured Throughput Index Overriding the Memory-Bandwidth Formula in llmfit

> Discover how benchmarks.rs crafts a measured throughput index, replacing memory bandwidth calculations in llmfit. Access real-world TPS benchmarks for accurate performance insights.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: deep-dive
- Published: 2026-09-12

---

**The `MeasuredTpsIndex` struct in [`llmfit-core/src/benchmarks.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/benchmarks.rs) loads real-world token-per-second (TPS) benchmarks into a hardware-keyed `HashMap`, and [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) queries this index to override theoretical memory-bandwidth estimates with observed performance data.**

In the `AlexsJones/llmfit` repository, the fitting engine determines whether an LLM will execute efficiently on a target system by estimating throughput. While [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) implements a generic memory-bandwidth formula to calculate theoretical TPS limits, [`benchmarks.rs`](https://github.com/AlexsJones/llmfit/blob/main/benchmarks.rs) provides a `MeasuredTpsIndex` that supersedes these calculations with empirical benchmarks when available.

## The Memory-Bandwidth Baseline in fit.rs

The default throughput estimation in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) relies on a memory-bandwidth formula that calculates theoretical token-per-second limits based on available RAM and bandwidth constants. This approach provides a safe, generic fallback for unknown hardware configurations. However, theoretical calculations often diverge from real-world performance due to driver overhead, quantization inefficiencies, and hardware-specific optimizations.

## Building the Measured Throughput Index in benchmarks.rs

The override mechanism centers on the `MeasuredTpsIndex` struct defined in [`llmfit-core/src/benchmarks.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/benchmarks.rs), which caches empirical benchmark results and exposes a lookup interface.

### Defining the Index Structure (Line 443)

At line 443, the code defines the `MeasuredTpsIndex` struct containing a `HashMap<(String, String), f64>` that maps hardware-quantization pairs to measured TPS values:

```rust
// llmfit-core/src/benchmarks.rs#L443
pub struct MeasuredTpsIndex {
    inner: HashMap<(String, String), f64>,
}

```

### Loading Benchmark Data via from_rows

The implementation block starting at line 449 provides `from_rows`, which parses the embedded benchmark cache (typically JSON) and populates the hash map. This method transforms raw benchmark rows into the typed `(hardware, quantization) → tps` mapping stored in the `inner` field.

### Singleton Pattern with for_specs and OnceLock

To avoid reloading data for every estimation, lines 485-502 implement `for_specs` using `std::sync::OnceLock`. This method initializes a global singleton `MeasuredTpsIndex` instance based on the detected `SystemSpecs`, ensuring the benchmark data loads exactly once per process:

```rust
// llmfit-core/src/benchmarks.rs#L485-L502
pub fn for_specs(specs: &SystemSpecs) -> Option<&'static MeasuredTpsIndex> {
    static INDEX: OnceLock<Option<MeasuredTpsIndex>> = OnceLock::new();
    INDEX.get_or_init(|| {
        // Load and parse benchmark data...
        Some(MeasuredTpsIndex::from_rows(/* ... */))
    }).as_ref()
}

```

### The lookup Method

Lines 511-520 define the `lookup` method, which accepts a `model_hf_id` and `quant` string and returns `Option<f64>`. This method performs the hash map retrieval, returning `Some(tps)` when a matching hardware-quantization entry exists, or `None` to signal that no measured data is available:

```rust
// llmfit-core/src/benchmarks.rs#L511-L520
pub fn lookup(&self, model_hf_id: &str, quant: &str) -> Option<f64> {
    self.inner.get(&(model_hf_id.to_string(), quant.to_string())).copied()
}

```

## How the Override Works in fit.rs

At approximately line 530 in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs), the throughput estimation logic queries the measured index before applying the memory-bandwidth formula. The code attempts to retrieve a measured TPS value via `MeasuredTpsIndex::for_specs(specs)?.lookup(model_hf_id, quantization)`. If the lookup returns `Some(tps)`, the function uses this empirical value directly. Only when the lookup returns `None` does the system fall back to the theoretical `estimate_tps_from_bandwidth` calculation.

This priority ensures that real-world benchmark data always supersedes theoretical estimates, providing more accurate fit predictions for hardware configurations present in the benchmark cache.

## Practical Implementation Example

The following pattern demonstrates how the two estimation strategies interact:

```rust
// llmfit-core/src/fit.rs (conceptual)
let estimated_tps = if let Some(index) = MeasuredTpsIndex::for_specs(&system_specs) {
    if let Some(measured) = index.lookup(&model.id, &quantization) {
        measured // Use empirical benchmark
    } else {
        estimate_tps_from_bandwidth(&system_specs, &model) // Fallback to formula
    }
} else {
    estimate_tps_from_bandwidth(&system_specs, &model) // No index available
};

```

## Summary

- **`MeasuredTpsIndex`** in [`benchmarks.rs`](https://github.com/AlexsJones/llmfit/blob/main/benchmarks.rs) stores a `HashMap` of hardware-specific, measured TPS values indexed by quantization method.
- **Construction** occurs via `from_rows` (parsing embedded data) and `for_specs` (managing a `OnceLock` singleton).
- **Override logic** in [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) (around line 530) queries the index first; a `Some` result bypasses the memory-bandwidth formula entirely.
- **Fallback** to the theoretical memory-bandwidth estimate only occurs when no matching benchmark exists for the current hardware-quantization pair.

## Frequently Asked Questions

### What is the MeasuredTpsIndex in llmfit?

The `MeasuredTpsIndex` is a struct defined in [`llmfit-core/src/benchmarks.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/benchmarks.rs) at line 443 that caches real-world token-per-second benchmarks. It maps tuples of hardware identifiers and quantization methods to observed TPS values, allowing the fitting engine to use empirical data instead of theoretical calculations.

### How does benchmarks.rs load benchmark data?

The `from_rows` method (starting at line 449 in [`benchmarks.rs`](https://github.com/AlexsJones/llmfit/blob/main/benchmarks.rs)) parses the embedded benchmark cache—typically a JSON dataset compiled into the binary—and populates the `HashMap` inside `MeasuredTpsIndex`. The `for_specs` method then wraps this in a `OnceLock` singleton to ensure efficient, thread-safe access across the application.

### When does fit.rs use the memory-bandwidth formula instead of measured data?

[`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) uses the memory-bandwidth formula as a fallback when `MeasuredTpsIndex::lookup` returns `None`, indicating no benchmark exists for the specific combination of hardware model and quantization level. This occurs around line 530 in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs), where the code branches between measured and estimated throughput.

### Can I extend the measured index with custom benchmarks?

While the current implementation loads from an embedded cache at compile time, the `from_rows` method accepts iterable data, meaning you could modify the source to ingest custom benchmark files at runtime. However, the default `for_specs` singleton pattern expects the standard hardware detection logic; extending it would require modifying the initialization sequence in [`benchmarks.rs`](https://github.com/AlexsJones/llmfit/blob/main/benchmarks.rs).