# How llmfit Estimates Model Speed (Throughput): The Physics-Based TPS Calculator Explained

> Discover how llmfit estimates model speed using its physics-based TPS calculator. Learn the formula and factors influencing token per second calculations for LLMs.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: internals
- Published: 2026-08-23

---

**TLDR:** `llmfit` estimates model throughput in tokens per second (TPS) by dividing the host's memory bandwidth by the model's size in gigabytes, then applying quantization scaling, an efficiency factor, and run-mode multipliers — all implemented in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs).

The [AlexsJones/llmfit](https://github.com/AlexsJones/llmfit) repository is an open-source tool that analyzes whether a given large language model will run comfortably on your hardware. A critical output of its analysis pipeline is the **estimated tokens per second (TPS)** — a prediction of generation speed. Instead of relying on arbitrary heuristics, `llmfit` uses a deterministic, hardware-aware formula rooted in the physics of memory bandwidth. Let's walk through exactly how **llmfit estimates model speed** from your system's specs.

## The Fundamental Bandwidth Model

At its core, token generation during inference is **memory-bandwidth-bound**. Each generated token requires the model weights to be read once from the accelerator's memory (VRAM or system RAM). This means the theoretical maximum throughput is simply:

```text
max_tps = memory_bandwidth_GB_s / model_size_GB

```

The dividing line — where memory traffic dominates compute — is often called the **roofline model**, and it is exactly what llmfit implements. The function comment in [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) at [line 1084–1094](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) lays out this premise: if you know how fast the memory can stream data and how big the model weights are, you know the absolute ceiling for generation speed.

## Calculating Model Size with Quantization

The model size used in the denominator depends on which **quantization** is active. In [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs), the helper `models::quant_bytes_per_param` returns the bytes-per-parameter for the selected precision (Q4, Q8, FP16, etc.). The "effective" parameter count is then scaled by that value:

```rust
let bytes_per_param = models::quant_bytes_per_param(quant);
let active_gb = params * bytes_per_param;

```

This snippet, from [lines 1224–1227](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs), shows the **model size conversion**. For Mixture-of-Experts (MoE) models, the "active" parameters (the experts actually used per token) are used instead of the full model count, giving a more realistic size for throughput math.

## Where the Bandwidth Value Comes From

The numerator — memory bandwidth — has two sources, as implemented by the `SystemSpecs` detection layer:

* **GPU detected:** `hardware::gpu_memory_bandwidth_gbps` looks up a table of real-world measured bandwidths. For example, an RTX 4090 is rated at approximately 1008 GB/s, while an Apple M1 Max sits near 400 GB/s.
* **GPU unknown:** a per-backend constant is used as fallback, so the calculation always completes even on unusual hardware.

The lookup itself appears at [lines 1219–1221](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) in the `estimate_tps` function.

## The Efficiency Factor: Correcting the Optimistic Ceiling

Raw bandwidth divided by model size yields an ideal value that no real kernel can reach. **Kernel launch overhead**, **KV-cache reads**, and other fixed costs consume a meaningful slice of that bandwidth. `llmfit` therefore applies a configurable **efficiency factor**, defaulting to `0.55`, to produce a realistic estimate:

```rust
let base_est = efficiency * (bandwidth / active_gb);

```

The efficiency default and override mechanism are defined at [lines 1229–1232](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) of [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs). Users can raise or lower the value through the TUI’s “Advanced Configuration” panel to match their own observed hardware behavior.

## Run-Mode Multipliers

After the bandwidth-based raw TPS is computed, a run-mode-specific factor scales the number again. The `run_mode_factors` defined in `CalcConfig` capture the relative performance of each execution path:

```rust
RunModeFactors { gpu: 1.0, tensor_parallel: 0.9, moe_offload: 0.8,
                 cpu_offload: 0.5, cpu_only: 0.3 }

```

These defaults — visible at [lines 25–31](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) — reflect real-world penalties: CPU offloading loses half throughput, and pure CPU execution drops to 30%. Tensor parallelism retains 90% of the GPU rate, while MoE offload sits at 80%.

## MoE-Specific Bandwidth Handling

A key nuance lies in how llmfit treats **Mixture-of-Experts** models. The conditional logic in `estimate_tps` (lines [1231–1249](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs)) distinguishes two scenarios:

* **MoE-offload:** when only active experts reside in VRAM and the remaining experts stream from system RAM, the bottleneck shifts to DDR bandwidth. The function uses `ddr_bandwidth_gbps` — the system's RAM bandwidth — rather than GPU bandwidth.
* **GPU-only MoE:** current runtimes load all layers uniformly into VRAM, so the estimation still reads the full model size from the GPU's memory.

This distinction is critical because using GPU bandwidth to calculate a CPU-offloaded MoE model would produce a wildly inflated TPS estimate.

## Returning the Final TPS Value

The final scaled TPS is returned and stored as `ModelFit::estimated_tps`. The function's return statement sits at [lines 1268–1270](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs). From there, the value flows into:

* The CLI/TUI result tables, in the `SPEED (tok/s)` column.
* The `ModelFit` struct, for downstream scoring like the "Perfect" or "Marginal" fit ratings.

## Actually Using the TPS Estimate in Practice

Run the CLI to see a token-per-second prediction with your current machine:

```bash

# Show fits for the 7B phi-2 model on the current machine

llmfit fit --model phi-2 --perfect

```

The output includes a `tok/s` column:

```text
MODEL      RUN MODE  FIT     SPEED (tok/s)
phi-2      GPU      Perfect   61.2

```

If you prefer to call the estimator from Rust, the API is small and direct:

```rust
use llmfit_core::{hardware::SystemSpecs, models::LlmModel, fit::ModelFit};

fn main() {
    // Detect the host system
    let system = SystemSpecs::detect().expect("system detection failed");

    // Load a model from the embedded catalog
    let model = LlmModel::from_name("phi-2")
        .expect("model not found");

    // Run the analysis
    let fit = ModelFit::analyze(&model, &system);

    // Estimated tokens per second
    println!("Estimated TPS: {:.1}", fit.estimated_tps);
}

```

Running this snippet prints the same `61.2` value shown in the CLI table.

## Key Files That Implement the Estimation

| File | Purpose |
|------|---------|
| [[`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs)](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) | Core pipeline with `estimate_tps` function; quantization handling, efficiency, run-mode factors, MoE branching |
| [[`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs)](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs) | GPU/CPU bandwidth tables and host-detection helpers |
| [[`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs)](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs) | Quantization metadata, `quant_bytes_per_param`, `quant_bpp` |
| [[`llmfit-core/src/benchmarks.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/benchmarks.rs)](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/benchmarks.rs) | Optional community or locally measured throughput overrides |

## Calibration with Real-World Benchmarks

The deterministic, physics-based estimate is not the final word. When community or locally measured benchmarks exist for the model and hardware pairing, `llmfit` overrides the formula in [`benchmarks.rs`](https://github.com/AlexsJones/llmfit/blob/main/benchmarks.rs). This hybrid approach gives you a confident estimate for unknown hardware and a calibrated number for well-tested hardware.

## Summary

Here are the key takeaways for how `llmfit` calculates **model speed (throughput)**:

- The theoretical TPS is **memory bandwidth divided by model size**, based on a roofline model of bandwidth-bound generation.
- Model size honors quantization (Q4, Q8, FP16) via `quant_bytes_per_param`, and uses active parameters for MoE models.
- Bandwidth comes from hardware tables in [`hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/hardware.rs), with per-backend defaults as fallback.
- A default **efficiency factor of 0.55** and run-mode multipliers (from `run_mode_factors`) correct the theoretical ceiling.
- MoE-offloaded models switch bandwidth sources to system DDR, so the estimate stays realistic.

## Frequently Asked Questions

### How does llmfit calculate tokens per second?

llmfit divides the host's memory bandwidth (GB/s) by the model's memory footprint (GB) — including quantization — then multiplies the result by an efficiency factor and a run-mode multiplier. The final value is stored as `ModelFit::estimated_tps` in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs).

### What is the default efficiency factor in llmfit's speed estimate?

The default efficiency factor is **0.55**. It accounts for kernel overhead and other operational costs. Users can adjust it through the TUI's “Advanced Configuration” panel.

### Does llmfit's speed estimate work differently for MoE models?

Yes. If a Mixture-of-Experts model is offloaded (experts in VRAM while others stream from memory), the estimate uses the system's DDR bandwidth rather than GPU bandwidth. For GPU-only MoE, full model size from VRAM is used, per the code in `estimate_tps`.

### Where can I find the speed estimation logic in the source code?

The estimation lives in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) in the `estimate_tps` function. Return the value into `ModelFit::estimated_tps`, with supporting data in [`hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/hardware.rs) (bandwidth) and [`models.rs`](https://github.com/AlexsJones/llmfit/blob/main/models.rs) (quantization bytes).