# How the Memory-Bandwidth Roofline Model Estimates Throughput in LLMFIT

> Understand how the memory-bandwidth roofline model estimates LLMFIT throughput. Discover how it calculates theoretical ceilings and applies efficiency factors for accurate token-generation speed predictions.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: deep-dive
- Published: 2026-08-21

---

**LLMFIT predicts token-generation speeds by treating transformer inference as a memory-bandwidth-bound operation, calculating the theoretical ceiling as GPU memory bandwidth divided by model size, then applying efficiency factors to account for real-world overhead.**

The [AlexsJones/llmfit](https://github.com/AlexsJones/llmfit) repository implements a physics-based **memory-bandwidth roofline model** to forecast LLM performance without relying solely on empirical benchmarks. This approach recognizes that generating each token requires reading the entire model weight set from GPU memory, making inference throughput primarily constrained by memory bandwidth rather than compute capacity. The implementation centers on the `estimate_tps` function in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs), which orchestrates hardware detection, model analysis, and calibration adjustments.

## The Six-Step Roofline Calculation

The `estimate_tps` function (starting at line 1111 of [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs)) executes the roofline estimation through six sequential phases:

1. **Hardware Bandwidth Lookup.** The system invokes `crate::hardware::gpu_memory_bandwidth_gbps` from [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs) to retrieve the GPU's memory bandwidth in GB/s. If the specific GPU model is unknown, the implementation falls back to a per-backend constant.

2. **Model Size Determination.** For dense architectures, the calculation uses `params × bytes_per_param(quant)`. For Mixture-of-Experts (MoE) models, the implementation preferentially uses `model.active_parameters` to count only the active expert weights, converting the final value to gigabytes (see line 1194).

3. **Raw Throughput Ceiling.** The theoretical maximum is computed as `raw_tps = bandwidth_GB_s / model_size_GB`. This division appears explicitly at line 1226: `let raw_tps = bandwidth / active_gb;`.

4. **Efficiency Factor Application.** The model accounts for kernel launch overhead, KV-cache reads, and controller inefficiency through `config.efficiency`, which defaults to **0.55** (line 1229).

5. **Run-Mode Adjustment.** Different execution paths (GPU, MoE-offload, CPU-only) apply specific scaling factors via `config.run_mode_factors`. The final estimate equals `raw_tps × efficiency × run_mode_factor`.

6. **Basis Recording.** All assumptions—including the method name (`"gpu_bandwidth_roofline"`), bandwidth, efficiency, and assumed context length—are stored in the `EstimateBasis` struct (line 1080) to ensure reproducible predictions.

## Handling Mixture-of-Experts Architectures

The roofline implementation distinguishes between two MoE execution strategies based on memory access patterns:

**GPU-Only MoE Execution.** Even when only active experts are required for computation, most inference runtimes load the entire weight set into VRAM for each token. Consequently, the implementation uses the full GPU memory bandwidth for throughput calculations rather than scaling proportionally to active parameters.

**MoE-Offload Scenarios.** When active experts reside in system RAM rather than VRAM, the model substitutes `ddr_bandwidth_gbps` for GPU bandwidth. This reflects the physics of reading weights across the system memory bus or PCIe interface, which typically operates at significantly lower bandwidth than GPU VRAM.

These architectural distinctions are documented inline at lines 1240-1250 of [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs).

## Calibration and Validation

The roofline constants are validated against empirical benchmarks from NVIDIA RTX 4090, T4, and Apple M1 Max GPUs (lines 1215-1218). When the system detects a known GPU with established bandwidth characteristics, it selects the roofline estimation path; otherwise, it automatically falls back to the historic `backend_constant` method to ensure functional predictions across all hardware configurations.

## Practical Usage Examples

Run a fit with GPU roofline estimation and inspect the underlying basis:

```bash

# Run a fit and display the estimated tokens-per-second

cargo run -- fit --model "Llama-2-7b-chat-q4_0" --max-context 8192

# Example output showing roofline methodology

# Model: Llama-2-7b-chat-q4_0

# Run mode: GPU

# Baseline estimated speed: 58.3 tok/s

# Estimate basis:

#   method: "gpu_bandwidth_roofline"

#   gpu_bandwidth_gbps: Some(400.0)

#   efficiency: 0.55

#   assumed_context: 8192

```

Access the `EstimateBasis` programmatically from Rust to audit the calculation parameters:

```rust
use llmfit_core::fit::ModelFit;

fn print_estimate(fit: &ModelFit) {
    println!("Method: {}", fit.estimate_basis.method);
    if let Some(bw) = fit.estimate_basis.gpu_bandwidth_gbps {
        println!("GPU bandwidth (GB/s): {:.1}", bw);
    }
    println!("Efficiency factor: {:.2}", fit.estimate_basis.efficiency);
    println!("Assumed context tokens: {}", fit.estimate_basis.assumed_context);
    println!("Estimated TPS: {:.1}", fit.estimated_tps);
}

```

Force a CPU-only estimate to observe the fallback behavior when GPU bandwidth is unavailable:

```bash

# Force CPU-only estimation (triggers backend_constant fallback)

cargo run -- fit --model "Llama-2-7b-chat-q4_0" --cpu-only

# Output indicates method = "cpu_constant"

```

## Summary

- The **memory-bandwidth roofline model** calculates theoretical throughput by dividing GPU memory bandwidth (GB/s) by model size (GB), establishing a physical ceiling based on memory access requirements.
- Implementation resides in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) within the `estimate_tps` function, with hardware specifications supplied by [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs).
- Real-world efficiency defaults to **0.55** to account for kernel overhead and system latency, adjustable via `config.efficiency`.
- **MoE architectures** switch between GPU bandwidth and DDR bandwidth depending on whether experts reside in VRAM or system RAM.
- All estimation parameters are captured in the `EstimateBasis` struct, enabling full reproducibility and auditing of predictions.

## Frequently Asked Questions

### What hardware information does LLMFIT require for the roofline model?

The system primarily requires the GPU identifier to lookup memory bandwidth via `gpu_memory_bandwidth_gbps` in [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs). If the GPU is unrecognized, the model automatically falls back to a generic per-backend constant rather than failing, ensuring predictions remain available for novel hardware configurations.

### Why does the efficiency factor default to 0.55?

The default 0.55 efficiency factor accounts for real-world overheads including CUDA kernel launch latency, KV-cache memory traffic, and PCIe controller inefficiencies that prevent applications from achieving the theoretical memory bandwidth ceiling. This value is calibrated against measured performance from RTX 4090, T4, and Apple M1 Max GPUs.

### How does LLMFIT handle quantization in model size calculations?

The model translates quantization strings (e.g., `"q4_0"`) into bytes-per-parameter values using helper functions from [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs), specifically `quant_bytes_per_param` and `quant_bpp`. These utilities calculate the effective model size by multiplying the raw parameter count by the compression ratio of the specified quantization scheme.

### What distinguishes GPU-only MoE from MoE-offload in the calculations?

GPU-only MoE uses the full GPU memory bandwidth because most runtimes load the entire expert set into VRAM regardless of active selection, while MoE-offload substitutes `ddr_bandwidth_gbps` to reflect reading active experts from system RAM. This distinction is critical because DDR bandwidth is typically an order of magnitude lower than GPU VRAM bandwidth.