How benchmarks.rs Creates a Measured Throughput Index Overriding the Memory-Bandwidth Formula in llmfit
The MeasuredTpsIndex struct in llmfit-core/src/benchmarks.rs loads real-world token-per-second (TPS) benchmarks into a hardware-keyed HashMap, and llmfit-core/src/fit.rs queries this index to override theoretical memory-bandwidth estimates with observed performance data.
In the AlexsJones/llmfit repository, the fitting engine determines whether an LLM will execute efficiently on a target system by estimating throughput. While fit.rs implements a generic memory-bandwidth formula to calculate theoretical TPS limits, benchmarks.rs provides a MeasuredTpsIndex that supersedes these calculations with empirical benchmarks when available.
The Memory-Bandwidth Baseline in fit.rs
The default throughput estimation in llmfit-core/src/fit.rs relies on a memory-bandwidth formula that calculates theoretical token-per-second limits based on available RAM and bandwidth constants. This approach provides a safe, generic fallback for unknown hardware configurations. However, theoretical calculations often diverge from real-world performance due to driver overhead, quantization inefficiencies, and hardware-specific optimizations.
Building the Measured Throughput Index in benchmarks.rs
The override mechanism centers on the MeasuredTpsIndex struct defined in llmfit-core/src/benchmarks.rs, which caches empirical benchmark results and exposes a lookup interface.
Defining the Index Structure (Line 443)
At line 443, the code defines the MeasuredTpsIndex struct containing a HashMap<(String, String), f64> that maps hardware-quantization pairs to measured TPS values:
// llmfit-core/src/benchmarks.rs#L443
pub struct MeasuredTpsIndex {
inner: HashMap<(String, String), f64>,
}
Loading Benchmark Data via from_rows
The implementation block starting at line 449 provides from_rows, which parses the embedded benchmark cache (typically JSON) and populates the hash map. This method transforms raw benchmark rows into the typed (hardware, quantization) → tps mapping stored in the inner field.
Singleton Pattern with for_specs and OnceLock
To avoid reloading data for every estimation, lines 485-502 implement for_specs using std::sync::OnceLock. This method initializes a global singleton MeasuredTpsIndex instance based on the detected SystemSpecs, ensuring the benchmark data loads exactly once per process:
// llmfit-core/src/benchmarks.rs#L485-L502
pub fn for_specs(specs: &SystemSpecs) -> Option<&'static MeasuredTpsIndex> {
static INDEX: OnceLock<Option<MeasuredTpsIndex>> = OnceLock::new();
INDEX.get_or_init(|| {
// Load and parse benchmark data...
Some(MeasuredTpsIndex::from_rows(/* ... */))
}).as_ref()
}
The lookup Method
Lines 511-520 define the lookup method, which accepts a model_hf_id and quant string and returns Option<f64>. This method performs the hash map retrieval, returning Some(tps) when a matching hardware-quantization entry exists, or None to signal that no measured data is available:
// llmfit-core/src/benchmarks.rs#L511-L520
pub fn lookup(&self, model_hf_id: &str, quant: &str) -> Option<f64> {
self.inner.get(&(model_hf_id.to_string(), quant.to_string())).copied()
}
How the Override Works in fit.rs
At approximately line 530 in llmfit-core/src/fit.rs, the throughput estimation logic queries the measured index before applying the memory-bandwidth formula. The code attempts to retrieve a measured TPS value via MeasuredTpsIndex::for_specs(specs)?.lookup(model_hf_id, quantization). If the lookup returns Some(tps), the function uses this empirical value directly. Only when the lookup returns None does the system fall back to the theoretical estimate_tps_from_bandwidth calculation.
This priority ensures that real-world benchmark data always supersedes theoretical estimates, providing more accurate fit predictions for hardware configurations present in the benchmark cache.
Practical Implementation Example
The following pattern demonstrates how the two estimation strategies interact:
// llmfit-core/src/fit.rs (conceptual)
let estimated_tps = if let Some(index) = MeasuredTpsIndex::for_specs(&system_specs) {
if let Some(measured) = index.lookup(&model.id, &quantization) {
measured // Use empirical benchmark
} else {
estimate_tps_from_bandwidth(&system_specs, &model) // Fallback to formula
}
} else {
estimate_tps_from_bandwidth(&system_specs, &model) // No index available
};
Summary
MeasuredTpsIndexinbenchmarks.rsstores aHashMapof hardware-specific, measured TPS values indexed by quantization method.- Construction occurs via
from_rows(parsing embedded data) andfor_specs(managing aOnceLocksingleton). - Override logic in
fit.rs(around line 530) queries the index first; aSomeresult bypasses the memory-bandwidth formula entirely. - Fallback to the theoretical memory-bandwidth estimate only occurs when no matching benchmark exists for the current hardware-quantization pair.
Frequently Asked Questions
What is the MeasuredTpsIndex in llmfit?
The MeasuredTpsIndex is a struct defined in llmfit-core/src/benchmarks.rs at line 443 that caches real-world token-per-second benchmarks. It maps tuples of hardware identifiers and quantization methods to observed TPS values, allowing the fitting engine to use empirical data instead of theoretical calculations.
How does benchmarks.rs load benchmark data?
The from_rows method (starting at line 449 in benchmarks.rs) parses the embedded benchmark cache—typically a JSON dataset compiled into the binary—and populates the HashMap inside MeasuredTpsIndex. The for_specs method then wraps this in a OnceLock singleton to ensure efficient, thread-safe access across the application.
When does fit.rs use the memory-bandwidth formula instead of measured data?
fit.rs uses the memory-bandwidth formula as a fallback when MeasuredTpsIndex::lookup returns None, indicating no benchmark exists for the specific combination of hardware model and quantization level. This occurs around line 530 in llmfit-core/src/fit.rs, where the code branches between measured and estimated throughput.
Can I extend the measured index with custom benchmarks?
While the current implementation loads from an embedded cache at compile time, the from_rows method accepts iterable data, meaning you could modify the source to ingest custom benchmark files at runtime. However, the default for_specs singleton pattern expects the standard hardware detection logic; extending it would require modifying the initialization sequence in benchmarks.rs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →