How llmfit Estimates LLM Throughput: Roofline Modeling and Efficiency Factors

llmfit predicts token-per-second (TPS) throughput by applying a roofline bandwidth model to GPU memory bandwidth, scaling by a hardware efficiency factor (default 0.55), and adjusting for runtime modes like Tensor Parallel or CPU offloading.

llmfit is an open-source Rust tool from the AlexsJones/llmfit repository that forecasts how many tokens per second a large language model will generate on specific hardware. The llmfit estimate LLM throughput algorithm combines memory-bandwidth physics with empirical efficiency adjustments to deliver transparent, reproducible performance predictions without executing the model.

The Roofline Bandwidth Foundation

The core estimation logic lives in llmfit-core/src/fit.rs. The system implements a roofline model that treats memory bandwidth as the primary bottleneck for autoregressive token generation.

When estimate_tps is invoked (around line 7780), it constructs an EstimateBasis (defined at lines 24-44) that records the estimation path:

  • "gpu_bandwidth_roofline" for known GPUs with cataloged memory bandwidth
  • "backend_constant" for unrecognized GPUs using per-backend defaults
  • "cpu_constant" for CPU-only inference scenarios

The function resolve_gpu_bandwidth queries detected hardware specs. If the GPU exists in the internal lookup table, llmfit uses its measured memory bandwidth. Users can override this value via CalcConfig::gpu_bandwidth_gbps_override to test hypothetical hardware or overclocked configurations.

Hardware Efficiency and Calibration

Raw bandwidth figures require adjustment for real-world inefficiencies. The estimator applies an efficiency factor defined in CalcConfig (lines 20-23).

The default value is 0.55, accounting for:

  • Kernel launch overhead
  • KV-cache read latency during generation
  • Memory controller inefficiencies

An optional local calibration factor (EstimateBasis::local_calibration) can further refine the estimate based on actual benchmark runs from the user's machine. The EstimateConfidence enum (lines 50-66) expresses the resulting trust level—ranging from community-measured data to pure theoretical estimation.

Runtime Mode Speed Multipliers

After computing base TPS, llmfit applies run-mode speed multipliers defined in CalcConfig::run_mode_factors (lines 25-27). The path-selection logic in analyze_inner (lines 5130-5180) determines which mode applies:

  • gpu (1.0): Baseline single-GPU inference
  • tensor_parallel (~0.9): Multi-GPU tensor parallelism overhead
  • moe_offload (~0.8): Mixture-of-Experts with partial offloading
  • cpu_offload (~0.5): Weights partially resident in system RAM
  • cpu_only (~0.3): Entire model running on CPU

These factors multiply against the raw TPS to produce the final estimated_tps value attached to each ModelFit.

Inspecting and Customizing Estimates

Every ModelFit result exposes its estimate_basis (line 754), making the calculation fully transparent and auditable.

CLI Examples:


# Display TPS estimates for all compatible models

cargo run -- --cli --tps

# Export detailed JSON for programmatic analysis

cargo run -- --cli --json | jq '.models[] | {name: .model.name, tps: .estimated_tps}'

# Override efficiency for high-performance memory subsystems

cargo run -- --cli --tps --efficiency 0.65

Rust API:

use llmfit_core::fit::{ModelFit, CalcConfig};

let system = llmfit_core::hardware::SystemSpecs::detect().unwrap();
let model = llmfit_core::models::load_model("meta-llama/Meta-Llama-3-8B-Instruct").unwrap();

let config = CalcConfig {
    efficiency: 0.70,  // Adjust for optimized kernels
    ..Default::default()
};

let fit = ModelFit::analyze_with_config(&model, &system, config);
println!("Estimated TPS: {:.1}", fit.estimated_tps);

Summary

  • llmfit uses a roofline bandwidth model in llmfit-core/src/fit.rs to calculate theoretical maximum throughput based on GPU memory bandwidth from resolve_gpu_bandwidth.
  • The efficiency factor (default 0.55) and optional local calibration adjust raw bandwidth for real-world overheads like KV-cache latency.
  • Run-mode multipliers (1.0 for GPU down to 0.3 for CPU-only) scale the estimate based on the inference configuration selected in analyze_inner.
  • Every estimate stores an EstimateBasis and EstimateConfidence level, ensuring traceable, debuggable predictions via ModelFit::estimate_basis.

Frequently Asked Questions

What is the default efficiency factor in llmfit?

The default efficiency factor is 0.55, defined in CalcConfig at lines 20-23 of llmfit-core/src/fit.rs. This constant accounts for kernel launch overhead, KV-cache reads, and memory controller inefficiencies that prevent GPUs from reaching theoretical peak bandwidth during autoregressive generation.

How does llmfit handle unknown GPUs?

When resolve_gpu_bandwidth cannot identify the GPU in its internal lookup table, llmfit falls back to a per-backend constant stored in the EstimateBasis as "backend_constant". Users can force a specific bandwidth value using CalcConfig::gpu_bandwidth_gbps_override to improve accuracy on exotic or unreleased hardware.

Can I calibrate llmfit with local benchmarks?

Yes. The EstimateBasis struct supports a local_calibration field that adjusts the TPS prediction based on actual measured throughput from your specific machine. When calibration data is present, the EstimateConfidence enum upgrades the confidence level from theoretical estimation to locally measured or calibrated status.

Where does llmfit store the estimation methodology?

The complete estimation pipeline, including estimate_tps, analyze_inner, and the ModelFit struct definition, resides in llmfit-core/src/fit.rs. The ModelFit::estimate_basis method (line 754) provides programmatic access to the calculation parameters, allowing users to audit exactly which bandwidth value, efficiency factor, and run-mode multiplier produced a given TPS figure.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →