How llmfit Estimates Model Speed (Throughput): The Physics-Based TPS Calculator Explained
TLDR: llmfit estimates model throughput in tokens per second (TPS) by dividing the host's memory bandwidth by the model's size in gigabytes, then applying quantization scaling, an efficiency factor, and run-mode multipliers — all implemented in llmfit-core/src/fit.rs.
The AlexsJones/llmfit repository is an open-source tool that analyzes whether a given large language model will run comfortably on your hardware. A critical output of its analysis pipeline is the estimated tokens per second (TPS) — a prediction of generation speed. Instead of relying on arbitrary heuristics, llmfit uses a deterministic, hardware-aware formula rooted in the physics of memory bandwidth. Let's walk through exactly how llmfit estimates model speed from your system's specs.
The Fundamental Bandwidth Model
At its core, token generation during inference is memory-bandwidth-bound. Each generated token requires the model weights to be read once from the accelerator's memory (VRAM or system RAM). This means the theoretical maximum throughput is simply:
max_tps = memory_bandwidth_GB_s / model_size_GB
The dividing line — where memory traffic dominates compute — is often called the roofline model, and it is exactly what llmfit implements. The function comment in fit.rs at line 1084–1094 lays out this premise: if you know how fast the memory can stream data and how big the model weights are, you know the absolute ceiling for generation speed.
Calculating Model Size with Quantization
The model size used in the denominator depends on which quantization is active. In llmfit-core/src/fit.rs, the helper models::quant_bytes_per_param returns the bytes-per-parameter for the selected precision (Q4, Q8, FP16, etc.). The "effective" parameter count is then scaled by that value:
let bytes_per_param = models::quant_bytes_per_param(quant);
let active_gb = params * bytes_per_param;
This snippet, from lines 1224–1227, shows the model size conversion. For Mixture-of-Experts (MoE) models, the "active" parameters (the experts actually used per token) are used instead of the full model count, giving a more realistic size for throughput math.
Where the Bandwidth Value Comes From
The numerator — memory bandwidth — has two sources, as implemented by the SystemSpecs detection layer:
- GPU detected:
hardware::gpu_memory_bandwidth_gbpslooks up a table of real-world measured bandwidths. For example, an RTX 4090 is rated at approximately 1008 GB/s, while an Apple M1 Max sits near 400 GB/s. - GPU unknown: a per-backend constant is used as fallback, so the calculation always completes even on unusual hardware.
The lookup itself appears at lines 1219–1221 in the estimate_tps function.
The Efficiency Factor: Correcting the Optimistic Ceiling
Raw bandwidth divided by model size yields an ideal value that no real kernel can reach. Kernel launch overhead, KV-cache reads, and other fixed costs consume a meaningful slice of that bandwidth. llmfit therefore applies a configurable efficiency factor, defaulting to 0.55, to produce a realistic estimate:
let base_est = efficiency * (bandwidth / active_gb);
The efficiency default and override mechanism are defined at lines 1229–1232 of fit.rs. Users can raise or lower the value through the TUI’s “Advanced Configuration” panel to match their own observed hardware behavior.
Run-Mode Multipliers
After the bandwidth-based raw TPS is computed, a run-mode-specific factor scales the number again. The run_mode_factors defined in CalcConfig capture the relative performance of each execution path:
RunModeFactors { gpu: 1.0, tensor_parallel: 0.9, moe_offload: 0.8,
cpu_offload: 0.5, cpu_only: 0.3 }
These defaults — visible at lines 25–31 — reflect real-world penalties: CPU offloading loses half throughput, and pure CPU execution drops to 30%. Tensor parallelism retains 90% of the GPU rate, while MoE offload sits at 80%.
MoE-Specific Bandwidth Handling
A key nuance lies in how llmfit treats Mixture-of-Experts models. The conditional logic in estimate_tps (lines 1231–1249) distinguishes two scenarios:
- MoE-offload: when only active experts reside in VRAM and the remaining experts stream from system RAM, the bottleneck shifts to DDR bandwidth. The function uses
ddr_bandwidth_gbps— the system's RAM bandwidth — rather than GPU bandwidth. - GPU-only MoE: current runtimes load all layers uniformly into VRAM, so the estimation still reads the full model size from the GPU's memory.
This distinction is critical because using GPU bandwidth to calculate a CPU-offloaded MoE model would produce a wildly inflated TPS estimate.
Returning the Final TPS Value
The final scaled TPS is returned and stored as ModelFit::estimated_tps. The function's return statement sits at lines 1268–1270. From there, the value flows into:
- The CLI/TUI result tables, in the
SPEED (tok/s)column. - The
ModelFitstruct, for downstream scoring like the "Perfect" or "Marginal" fit ratings.
Actually Using the TPS Estimate in Practice
Run the CLI to see a token-per-second prediction with your current machine:
# Show fits for the 7B phi-2 model on the current machine
llmfit fit --model phi-2 --perfect
The output includes a tok/s column:
MODEL RUN MODE FIT SPEED (tok/s)
phi-2 GPU Perfect 61.2
If you prefer to call the estimator from Rust, the API is small and direct:
use llmfit_core::{hardware::SystemSpecs, models::LlmModel, fit::ModelFit};
fn main() {
// Detect the host system
let system = SystemSpecs::detect().expect("system detection failed");
// Load a model from the embedded catalog
let model = LlmModel::from_name("phi-2")
.expect("model not found");
// Run the analysis
let fit = ModelFit::analyze(&model, &system);
// Estimated tokens per second
println!("Estimated TPS: {:.1}", fit.estimated_tps);
}
Running this snippet prints the same 61.2 value shown in the CLI table.
Key Files That Implement the Estimation
| File | Purpose |
|---|---|
[llmfit-core/src/fit.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) |
Core pipeline with estimate_tps function; quantization handling, efficiency, run-mode factors, MoE branching |
[llmfit-core/src/hardware.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs) |
GPU/CPU bandwidth tables and host-detection helpers |
[llmfit-core/src/models.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs) |
Quantization metadata, quant_bytes_per_param, quant_bpp |
[llmfit-core/src/benchmarks.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/benchmarks.rs) |
Optional community or locally measured throughput overrides |
Calibration with Real-World Benchmarks
The deterministic, physics-based estimate is not the final word. When community or locally measured benchmarks exist for the model and hardware pairing, llmfit overrides the formula in benchmarks.rs. This hybrid approach gives you a confident estimate for unknown hardware and a calibrated number for well-tested hardware.
Summary
Here are the key takeaways for how llmfit calculates model speed (throughput):
- The theoretical TPS is memory bandwidth divided by model size, based on a roofline model of bandwidth-bound generation.
- Model size honors quantization (Q4, Q8, FP16) via
quant_bytes_per_param, and uses active parameters for MoE models. - Bandwidth comes from hardware tables in
hardware.rs, with per-backend defaults as fallback. - A default efficiency factor of 0.55 and run-mode multipliers (from
run_mode_factors) correct the theoretical ceiling. - MoE-offloaded models switch bandwidth sources to system DDR, so the estimate stays realistic.
Frequently Asked Questions
How does llmfit calculate tokens per second?
llmfit divides the host's memory bandwidth (GB/s) by the model's memory footprint (GB) — including quantization — then multiplies the result by an efficiency factor and a run-mode multiplier. The final value is stored as ModelFit::estimated_tps in llmfit-core/src/fit.rs.
What is the default efficiency factor in llmfit's speed estimate?
The default efficiency factor is 0.55. It accounts for kernel overhead and other operational costs. Users can adjust it through the TUI's “Advanced Configuration” panel.
Does llmfit's speed estimate work differently for MoE models?
Yes. If a Mixture-of-Experts model is offloaded (experts in VRAM while others stream from memory), the estimate uses the system's DDR bandwidth rather than GPU bandwidth. For GPU-only MoE, full model size from VRAM is used, per the code in estimate_tps.
Where can I find the speed estimation logic in the source code?
The estimation lives in llmfit-core/src/fit.rs in the estimate_tps function. Return the value into ModelFit::estimated_tps, with supporting data in hardware.rs (bandwidth) and models.rs (quantization bytes).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →