# How llmfit Estimates LLM Throughput: Roofline Modeling and Efficiency Factors

> Discover how llmfit estimates LLM throughput using roofline modeling and efficiency factors. Learn about TPS prediction, GPU memory bandwidth, and runtime mode adjustments for optimal performance.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: deep-dive
- Published: 2026-09-13

---

**llmfit predicts token-per-second (TPS) throughput by applying a roofline bandwidth model to GPU memory bandwidth, scaling by a hardware efficiency factor (default 0.55), and adjusting for runtime modes like Tensor Parallel or CPU offloading.**

`llmfit` is an open-source Rust tool from the [AlexsJones/llmfit](https://github.com/AlexsJones/llmfit) repository that forecasts how many tokens per second a large language model will generate on specific hardware. The **llmfit estimate LLM throughput** algorithm combines memory-bandwidth physics with empirical efficiency adjustments to deliver transparent, reproducible performance predictions without executing the model.

## The Roofline Bandwidth Foundation

The core estimation logic lives in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs). The system implements a **roofline model** that treats memory bandwidth as the primary bottleneck for autoregressive token generation.

When `estimate_tps` is invoked (around line 7780), it constructs an `EstimateBasis` (defined at lines 24-44) that records the estimation path:

- `"gpu_bandwidth_roofline"` for known GPUs with cataloged memory bandwidth
- `"backend_constant"` for unrecognized GPUs using per-backend defaults
- `"cpu_constant"` for CPU-only inference scenarios

The function `resolve_gpu_bandwidth` queries detected hardware specs. If the GPU exists in the internal lookup table, llmfit uses its measured memory bandwidth. Users can override this value via `CalcConfig::gpu_bandwidth_gbps_override` to test hypothetical hardware or overclocked configurations.

## Hardware Efficiency and Calibration

Raw bandwidth figures require adjustment for real-world inefficiencies. The estimator applies an **efficiency factor** defined in `CalcConfig` (lines 20-23).

The default value is **0.55**, accounting for:
- Kernel launch overhead
- KV-cache read latency during generation
- Memory controller inefficiencies

An optional **local calibration** factor (`EstimateBasis::local_calibration`) can further refine the estimate based on actual benchmark runs from the user's machine. The `EstimateConfidence` enum (lines 50-66) expresses the resulting trust level—ranging from community-measured data to pure theoretical estimation.

## Runtime Mode Speed Multipliers

After computing base TPS, llmfit applies **run-mode speed multipliers** defined in `CalcConfig::run_mode_factors` (lines 25-27). The path-selection logic in `analyze_inner` (lines 5130-5180) determines which mode applies:

- **`gpu`** (1.0): Baseline single-GPU inference
- **`tensor_parallel`** (~0.9): Multi-GPU tensor parallelism overhead
- **`moe_offload`** (~0.8): Mixture-of-Experts with partial offloading
- **`cpu_offload`** (~0.5): Weights partially resident in system RAM
- **`cpu_only`** (~0.3): Entire model running on CPU

These factors multiply against the raw TPS to produce the final `estimated_tps` value attached to each `ModelFit`.

## Inspecting and Customizing Estimates

Every `ModelFit` result exposes its `estimate_basis` (line 754), making the calculation fully transparent and auditable.

**CLI Examples:**

```bash

# Display TPS estimates for all compatible models

cargo run -- --cli --tps

# Export detailed JSON for programmatic analysis

cargo run -- --cli --json | jq '.models[] | {name: .model.name, tps: .estimated_tps}'

# Override efficiency for high-performance memory subsystems

cargo run -- --cli --tps --efficiency 0.65

```

**Rust API:**

```rust
use llmfit_core::fit::{ModelFit, CalcConfig};

let system = llmfit_core::hardware::SystemSpecs::detect().unwrap();
let model = llmfit_core::models::load_model("meta-llama/Meta-Llama-3-8B-Instruct").unwrap();

let config = CalcConfig {
    efficiency: 0.70,  // Adjust for optimized kernels
    ..Default::default()
};

let fit = ModelFit::analyze_with_config(&model, &system, config);
println!("Estimated TPS: {:.1}", fit.estimated_tps);

```

## Summary

- llmfit uses a **roofline bandwidth model** in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) to calculate theoretical maximum throughput based on GPU memory bandwidth from `resolve_gpu_bandwidth`.
- The **efficiency factor** (default 0.55) and optional local calibration adjust raw bandwidth for real-world overheads like KV-cache latency.
- **Run-mode multipliers** (1.0 for GPU down to 0.3 for CPU-only) scale the estimate based on the inference configuration selected in `analyze_inner`.
- Every estimate stores an `EstimateBasis` and `EstimateConfidence` level, ensuring traceable, debuggable predictions via `ModelFit::estimate_basis`.

## Frequently Asked Questions

### What is the default efficiency factor in llmfit?

The default **efficiency factor is 0.55**, defined in `CalcConfig` at lines 20-23 of [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs). This constant accounts for kernel launch overhead, KV-cache reads, and memory controller inefficiencies that prevent GPUs from reaching theoretical peak bandwidth during autoregressive generation.

### How does llmfit handle unknown GPUs?

When `resolve_gpu_bandwidth` cannot identify the GPU in its internal lookup table, llmfit falls back to a **per-backend constant** stored in the `EstimateBasis` as `"backend_constant"`. Users can force a specific bandwidth value using `CalcConfig::gpu_bandwidth_gbps_override` to improve accuracy on exotic or unreleased hardware.

### Can I calibrate llmfit with local benchmarks?

Yes. The `EstimateBasis` struct supports a `local_calibration` field that adjusts the TPS prediction based on actual measured throughput from your specific machine. When calibration data is present, the `EstimateConfidence` enum upgrades the confidence level from theoretical estimation to locally measured or calibrated status.

### Where does llmfit store the estimation methodology?

The complete estimation pipeline, including `estimate_tps`, `analyze_inner`, and the `ModelFit` struct definition, resides in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs). The `ModelFit::estimate_basis` method (line 754) provides programmatic access to the calculation parameters, allowing users to audit exactly which bandwidth value, efficiency factor, and run-mode multiplier produced a given TPS figure.