How the Memory-Bandwidth Roofline Model Estimates Throughput in LLMFIT

LLMFIT predicts token-generation speeds by treating transformer inference as a memory-bandwidth-bound operation, calculating the theoretical ceiling as GPU memory bandwidth divided by model size, then applying efficiency factors to account for real-world overhead.

The AlexsJones/llmfit repository implements a physics-based memory-bandwidth roofline model to forecast LLM performance without relying solely on empirical benchmarks. This approach recognizes that generating each token requires reading the entire model weight set from GPU memory, making inference throughput primarily constrained by memory bandwidth rather than compute capacity. The implementation centers on the estimate_tps function in llmfit-core/src/fit.rs, which orchestrates hardware detection, model analysis, and calibration adjustments.

The Six-Step Roofline Calculation

The estimate_tps function (starting at line 1111 of llmfit-core/src/fit.rs) executes the roofline estimation through six sequential phases:

  1. Hardware Bandwidth Lookup. The system invokes crate::hardware::gpu_memory_bandwidth_gbps from llmfit-core/src/hardware.rs to retrieve the GPU's memory bandwidth in GB/s. If the specific GPU model is unknown, the implementation falls back to a per-backend constant.

  2. Model Size Determination. For dense architectures, the calculation uses params × bytes_per_param(quant). For Mixture-of-Experts (MoE) models, the implementation preferentially uses model.active_parameters to count only the active expert weights, converting the final value to gigabytes (see line 1194).

  3. Raw Throughput Ceiling. The theoretical maximum is computed as raw_tps = bandwidth_GB_s / model_size_GB. This division appears explicitly at line 1226: let raw_tps = bandwidth / active_gb;.

  4. Efficiency Factor Application. The model accounts for kernel launch overhead, KV-cache reads, and controller inefficiency through config.efficiency, which defaults to 0.55 (line 1229).

  5. Run-Mode Adjustment. Different execution paths (GPU, MoE-offload, CPU-only) apply specific scaling factors via config.run_mode_factors. The final estimate equals raw_tps × efficiency × run_mode_factor.

  6. Basis Recording. All assumptions—including the method name ("gpu_bandwidth_roofline"), bandwidth, efficiency, and assumed context length—are stored in the EstimateBasis struct (line 1080) to ensure reproducible predictions.

Handling Mixture-of-Experts Architectures

The roofline implementation distinguishes between two MoE execution strategies based on memory access patterns:

GPU-Only MoE Execution. Even when only active experts are required for computation, most inference runtimes load the entire weight set into VRAM for each token. Consequently, the implementation uses the full GPU memory bandwidth for throughput calculations rather than scaling proportionally to active parameters.

MoE-Offload Scenarios. When active experts reside in system RAM rather than VRAM, the model substitutes ddr_bandwidth_gbps for GPU bandwidth. This reflects the physics of reading weights across the system memory bus or PCIe interface, which typically operates at significantly lower bandwidth than GPU VRAM.

These architectural distinctions are documented inline at lines 1240-1250 of fit.rs.

Calibration and Validation

The roofline constants are validated against empirical benchmarks from NVIDIA RTX 4090, T4, and Apple M1 Max GPUs (lines 1215-1218). When the system detects a known GPU with established bandwidth characteristics, it selects the roofline estimation path; otherwise, it automatically falls back to the historic backend_constant method to ensure functional predictions across all hardware configurations.

Practical Usage Examples

Run a fit with GPU roofline estimation and inspect the underlying basis:


# Run a fit and display the estimated tokens-per-second

cargo run -- fit --model "Llama-2-7b-chat-q4_0" --max-context 8192

# Example output showing roofline methodology

# Model: Llama-2-7b-chat-q4_0

# Run mode: GPU

# Baseline estimated speed: 58.3 tok/s

# Estimate basis:

#   method: "gpu_bandwidth_roofline"

#   gpu_bandwidth_gbps: Some(400.0)

#   efficiency: 0.55

#   assumed_context: 8192

Access the EstimateBasis programmatically from Rust to audit the calculation parameters:

use llmfit_core::fit::ModelFit;

fn print_estimate(fit: &ModelFit) {
    println!("Method: {}", fit.estimate_basis.method);
    if let Some(bw) = fit.estimate_basis.gpu_bandwidth_gbps {
        println!("GPU bandwidth (GB/s): {:.1}", bw);
    }
    println!("Efficiency factor: {:.2}", fit.estimate_basis.efficiency);
    println!("Assumed context tokens: {}", fit.estimate_basis.assumed_context);
    println!("Estimated TPS: {:.1}", fit.estimated_tps);
}

Force a CPU-only estimate to observe the fallback behavior when GPU bandwidth is unavailable:


# Force CPU-only estimation (triggers backend_constant fallback)

cargo run -- fit --model "Llama-2-7b-chat-q4_0" --cpu-only

# Output indicates method = "cpu_constant"

Summary

  • The memory-bandwidth roofline model calculates theoretical throughput by dividing GPU memory bandwidth (GB/s) by model size (GB), establishing a physical ceiling based on memory access requirements.
  • Implementation resides in llmfit-core/src/fit.rs within the estimate_tps function, with hardware specifications supplied by llmfit-core/src/hardware.rs.
  • Real-world efficiency defaults to 0.55 to account for kernel overhead and system latency, adjustable via config.efficiency.
  • MoE architectures switch between GPU bandwidth and DDR bandwidth depending on whether experts reside in VRAM or system RAM.
  • All estimation parameters are captured in the EstimateBasis struct, enabling full reproducibility and auditing of predictions.

Frequently Asked Questions

What hardware information does LLMFIT require for the roofline model?

The system primarily requires the GPU identifier to lookup memory bandwidth via gpu_memory_bandwidth_gbps in llmfit-core/src/hardware.rs. If the GPU is unrecognized, the model automatically falls back to a generic per-backend constant rather than failing, ensuring predictions remain available for novel hardware configurations.

Why does the efficiency factor default to 0.55?

The default 0.55 efficiency factor accounts for real-world overheads including CUDA kernel launch latency, KV-cache memory traffic, and PCIe controller inefficiencies that prevent applications from achieving the theoretical memory bandwidth ceiling. This value is calibrated against measured performance from RTX 4090, T4, and Apple M1 Max GPUs.

How does LLMFIT handle quantization in model size calculations?

The model translates quantization strings (e.g., "q4_0") into bytes-per-parameter values using helper functions from llmfit-core/src/models.rs, specifically quant_bytes_per_param and quant_bpp. These utilities calculate the effective model size by multiplying the raw parameter count by the compression ratio of the specified quantization scheme.

What distinguishes GPU-only MoE from MoE-offload in the calculations?

GPU-only MoE uses the full GPU memory bandwidth because most runtimes load the entire expert set into VRAM regardless of active selection, while MoE-offload substitutes ddr_bandwidth_gbps to reflect reading active experts from system RAM. This distinction is critical because DDR bandwidth is typically an order of magnitude lower than GPU VRAM bandwidth.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →