How llmfit Calculates Theoretical Maximum Tokens Per Second from GPU Memory Bandwidth
llmfit derives the theoretical maximum tokens per second by dividing the GPU's peak memory bandwidth by the model's active parameter size in gigabytes, then applying an efficiency factor of approximately 0.55 to account for kernel overhead and KV-cache operations.
LLM inference throughput is fundamentally constrained by memory bandwidth rather than raw computational power, requiring accurate physics-based modeling to predict realistic performance ceilings. The llmfit crate implements this calculation in the estimate_tps function within llmfit-core/src/fit.rs, providing users with data-driven expectations for token generation speed based on hardware specifications and model quantization.
The Memory-Bandwidth Bottleneck in LLM Inference
Autoregressive token generation is memory-bandwidth-bound because each new token requires reading the entire model weight matrix from VRAM. Unlike training or batched inference where matrix multiplication dominates, single-token generation (batch size 1) spends most of its time waiting for weights to transfer across the memory bus rather than performing compute operations.
This physical limitation means that knowing a GPU's peak memory bandwidth and the model's effective size allows precise calculation of the theoretical throughput ceiling. The llmfit codebase explicitly models this relationship through a deterministic formula implemented in the core estimation pipeline.
How llmfit Calculates Theoretical Maximum TPS
The calculation follows a three-stage pipeline that converts hardware specifications and model metadata into a practical TPS estimate.
Step 1: GPU Memory Bandwidth Detection
The system first identifies the available hardware through hardware::gpu_memory_bandwidth_gbps() located in llmfit-core/src/hardware.rs. This function maintains a hard-coded lookup table mapping known GPU model names to their manufacturer-specified peak memory bandwidths.
For example, the function returns 1008 GB/s for an RTX 4090 and corresponding values for other data-center and consumer cards. This lookup occurs during the speed estimation routine and provides the bw variable used in subsequent calculations.
Step 2: Model Size Normalization with Quantization
In llmfit-core/src/fit.rs, the estimate_tps function calculates the effective model size in gigabytes through the model.params_b() method. For dense models, this uses the full parameter count, while MoE (Mixture of Experts) configurations may use only the active-expert parameter count depending on the runtime mode.
The per-parameter byte size is determined by models::quant_bytes_per_param, which converts quantization levels (such as q4_k_m or q8_0) into their corresponding storage sizes. This yields the active_gb value representing the actual memory footprint that must traverse the GPU bus during each generation step.
Step 3: The Physics-Based Formula and Efficiency Adjustment
With bandwidth (bw) and model size (active_gb) established, estimate_tps computes the raw theoretical ceiling as:
let raw_tps = bw / active_gb;
This represents the absolute physical limit if 100% of memory bandwidth were available for weight transfer. However, the implementation applies an efficiency factor (defaulting to approximately 0.55) to account for:
- Kernel launch overhead and scheduling latency
- KV-cache reads and writes during generation
- Memory controller contention and other fixed costs
The final calculation becomes raw_tps * efficiency, producing a realistic estimate rather than an unattainable theoretical peak. The source code comments in fit.rs explicitly document this 0.55 adjustment as accounting for these operational overheads.
Implementation Deep Dive
The actual estimation logic resides in llmfit-core/src/fit.rs and integrates with the hardware detection layer to produce TPS predictions. Here is how developers interact with the calculation programmatically:
use llmfit_core::{
hardware::SystemSpecs,
models::LlmModel,
fit::{estimate_tps, CalcConfig},
providers::InferenceRuntime,
plan::RunMode,
};
fn main() {
// Example: a 7B model quantized to Q4_K_M on an RTX 4090
let system = SystemSpecs::detect().unwrap();
let model = LlmModel::new("facebook/opt-7b").unwrap();
let config = CalcConfig::default();
let tps = estimate_tps(
&model,
"q4_k_m",
&system,
RunMode::Gpu,
InferenceRuntime::Ollama,
&config,
);
println!("Theoretical max TPS ≈ {:.1}", tps);
}
When executed on hardware matching the detected specifications, this snippet returns values consistent with the benchmark comments found in the source code. The estimate_tps function handles the internal conversion of quantization tags to byte counts and manages the division by the bandwidth value retrieved from hardware.rs.
Run-Mode Considerations and Fallbacks
The TPS calculation logic branches based on the RunMode enum passed to estimate_tps. For RunMode::Gpu and RunMode::MoeOffload, the system uses the memory-bandwidth-based formula described above.
For CPU-only execution or unsupported configurations, the implementation falls back to constant values rather than attempting bandwidth-based calculation, as CPU memory subsystems exhibit different bottleneck characteristics than discrete GPUs. This conditional logic appears in fit.rs around lines 1221-1230, ensuring appropriate estimation strategies for heterogeneous hardware environments.
Summary
- Memory bandwidth is the primary constraint for autoregressive LLM inference, making
bandwidth / model_sizethe correct asymptotic formula for theoretical TPS. - Hardware detection occurs in
hardware.rsthroughgpu_memory_bandwidth_gbps(), which maps GPU names to their rated bandwidths. - Model size calculations in
fit.rsaccount for quantization viaquant_bytes_per_param, converting billions of parameters into actual gigabytes transferred. - Efficiency adjustments default to 0.55 in
estimate_tpsto reflect real-world kernel overhead and KV-cache operations. - Run-mode handling switches between bandwidth-based estimates for GPU inference and fallback constants for CPU-only execution.
Frequently Asked Questions
Why does llmfit use memory bandwidth instead of FLOPS to calculate theoretical TPS?
LLM token generation at batch size 1 is memory-bandwidth-bound because the computation consists primarily of matrix-vector multiplications where the cost of loading weights from VRAM dominates the cost of performing the arithmetic. The source code in llmfit-core/src/fit.rs implements this "memory wall" observation by dividing memory bandwidth by model size rather than using peak compute metrics, which would overestimate throughput for single-token generation scenarios.
How does the efficiency factor of 0.55 account for real-world performance?
The efficiency factor hardcoded in estimate_tps accounts for kernel launch overhead, KV-cache memory traffic, and other fixed costs that consume bandwidth without contributing directly to weight matrix reads. The comments in llmfit-core/src/fit.rs explicitly describe this 0.55 multiplier as adjusting for operational realities including PCIe bottlenecks and GPU kernel scheduling latency, bringing theoretical peaks down to achievable sustained throughput.
Does llmfit adjust TPS calculations for MoE (Mixture of Experts) models?
Yes, the estimate_tps function in fit.rs adjusts the model size calculation for MoE architectures by potentially using only the active-expert parameter count rather than the full model size, depending on the specified RunMode. This reflects the sparse activation patterns in MoE inference where only a subset of parameters participates in each forward pass, significantly altering the memory bandwidth requirements compared to dense models.
Where does llmfit store the GPU memory bandwidth specifications?
The bandwidth lookup table resides in llmfit-core/src/hardware.rs within the gpu_memory_bandwidth_gbps() function. This hard-coded table maps specific GPU model strings (such as "RTX 4090") to their manufacturer-rated peak memory bandwidths in gigabytes per second, enabling the system to inject accurate hardware constraints into the TPS estimation formula without requiring runtime benchmarks.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →