# The Memory-Bandwidth Roof-Line Model Used by llmfit for Estimating LLM Inference Speed

> Discover how llmfit uses a memory-bandwidth roof-line model to estimate LLM inference speed. Learn about its calculation of theoretical throughput and efficiency factors for real-world performance.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: deep-dive
- Published: 2026-08-22

---

**llmfit employs a memory-bandwidth-limited roof-line model that calculates theoretical throughput by dividing GPU memory bandwidth by quantized model size, then applies configurable efficiency factors to account for real-world kernel overhead and runtime constraints.**

The `AlexsJones/llmfit` repository provides a predictive performance framework that forecasts LLM inference throughput without executing empirical benchmarks. Rather than measuring actual token generation, the underlying model used by llmfit for estimating LLM inference speed derives baseline performance from hardware memory bandwidth limits and model architecture parameters, storing the calculation basis in the `EstimateBasis` struct for full reproducibility.

## Roof-Line Formula and Theoretical Ceiling

The estimation engine in [`src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/src/fit.rs) implements a bandwidth-bound roof-line calculation where memory throughput serves as the primary constraint. The theoretical ceiling follows the formula `max_tps = memory_bandwidth_GB / model_size_GB`, establishing the maximum possible tokens per second given the hardware's memory subsystem. In practice, the `estimate_tps` function applies an **efficiency factor** defaulting to **0.55** to account for kernel-launch overhead, KV-cache read operations, and GPU controller inefficiencies, as implemented in lines 1088-1104 of [`src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/src/fit.rs).

### Efficiency and Run-Mode Factors

Beyond the baseline 0.55 efficiency multiplier, llmfit incorporates distinct **run-mode factors** that further calibrate throughput estimates for specific execution paths. These factors adjust the calculation for GPU-only inference, MoE-offload scenarios, and CPU-only execution, ensuring the final `estimated_tps` value reflects the practical constraints of the target deployment environment.

## Quantization-Aware Model Size Calculation

The accuracy of the roof-line model depends on precise model size determination, which varies by quantization precision. The function `quant_bytes_per_param()` in [`src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/src/fit.rs) (lines 1224-1226) maps quantization identifiers to bytes-per-parameter ratios—for example, Q4 quantization consumes approximately 0.5 bytes per parameter. This value feeds directly into the denominator of the bandwidth division, ensuring the memory footprint calculation matches the actual deployed model weights stored in `ModelDatabase`.

## GPU Bandwidth Detection and Fallbacks

Hardware capability discovery resides in [`src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/src/hardware.rs), which exports `gpu_memory_bandwidth_gbps` containing empirical bandwidth measurements for known accelerators. The table includes entries such as **RTX 4090 at approximately 1008 GB/s** and **Apple M1 Max at roughly 400 GB/s**. When encountering unlisted GPUs, the system gracefully falls back to conservative per-backend constants, ensuring continuous estimation capability without manual hardware profiling.

## Mixture-of-Experts (MoE) Handling

For sparse Mixture-of-Experts architectures, llmfit differentiates between total parameters and active parameters. When configured for `MoeOffload` run-mode, the estimator in [`src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/src/fit.rs) (lines 1191-1198) uses only the **active-expert parameter count**, reflecting the memory traffic reduction achieved through expert offloading. For standard run-modes, the calculation conservatively assumes the full model size, as most inference runtimes read the complete weight matrix regardless of expert activation patterns.

## Practical Implementation Example

The following Rust code demonstrates how to invoke the roof-line estimation model using the `ModelFit` API, `SystemSpecs` detection, and custom `CalcConfig` parameters:

```rust
use llmfit_core::fit::{ModelFit, CalcConfig};
use llmfit_core::hardware::SystemSpecs;
use llmfit_core::models::ModelDatabase;

// Load the embedded model catalog
let db = ModelDatabase::new();

// Detect host hardware including GPU bandwidth
let specs = SystemSpecs::detect()?;

// Analyze a specific model configuration
let model = db.get("llama-2-7b-chat").unwrap();
let fit = ModelFit::analyze(&model, &specs);

// Access the estimated throughput
println!("Estimated speed: {:.1} tok/s", fit.estimated_tps);

// Customize efficiency assumptions for newer hardware
let mut cfg = CalcConfig::default();
cfg.efficiency = 0.60; // 60% bandwidth efficiency
let custom_fit = ModelFit::analyze_with_config(&model, &specs, cfg);
println!("Custom estimate: {:.1} tok/s", custom_fit.estimated_tps);

```

## Summary

- **Memory-bandwidth roof-line**: The core estimation divides GPU memory bandwidth by quantized model size to establish theoretical throughput ceilings in `estimate_tps`.
- **Configurable efficiency**: A default 0.55 efficiency factor in `CalcConfig` accounts for real-world overhead, adjustable for optimized runtimes or newer hardware.
- **Quantization awareness**: Dynamic model sizing via `quant_bytes_per_param()` ensures accurate memory calculations across Q4, Q8, and other precision formats.
- **Hardware detection**: The `gpu_memory_bandwidth_gbps` table in [`src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/src/hardware.rs) provides measured bandwidths for known GPUs with automatic fallback constants.
- **MoE optimization**: Sparse model handling uses active-expert parameter counts for `MoeOffload` modes, improving accuracy for large sparse architectures.

## Frequently Asked Questions

### What hardware specifications does llmfit need to estimate inference speed?

llmfit automatically detects GPU models through `SystemSpecs::detect()` and references an internal bandwidth table in [`src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/src/hardware.rs). For unknown accelerators, the tool falls back to conservative per-backend constants, allowing estimation to proceed without manual hardware configuration or benchmarking.

### How does llmfit handle different quantization formats like Q4 or Q8?

The tool maps quantization strings to bytes-per-parameter ratios through the `quant_bytes_per_param()` function in [`src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/src/fit.rs). For example, Q4 quantization uses approximately 0.5 bytes per parameter, directly reducing the model size denominator in the roof-line throughput calculation to reflect compressed memory footprints.

### Why does llmfit use a 0.55 efficiency factor by default?

This factor accounts for real-world constraints absent from raw bandwidth specifications, including CUDA kernel launch overhead, KV-cache memory traffic during autoregressive generation, and GPU controller inefficiencies. Users can override this default via `CalcConfig` to match highly optimized inference engines or next-generation hardware characteristics.

### Can llmfit estimate speed for Mixture-of-Experts models accurately?

Yes, specifically for `MoeOffload` run-modes, the estimator uses only active-expert parameters rather than the full model size, accurately reflecting the reduced memory bandwidth requirements of sparse computation. For non-offloaded MoE execution, the model conservatively assumes full weight matrix reads per token.