# How to Configure llmfit Efficiency and Run Mode Factors for TPS Calibration

> Calibrate llmfit token per second TPS predictions using efficiency and run mode factors in CalcConfig. Optimize your model's performance against real hardware.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: how-to-guide
- Published: 2026-08-20

---

**Set the `efficiency` scalar and `run_mode_factors` multipliers in `CalcConfig` to calibrate llmfit's token-per-second predictions against real hardware measurements.**

The `llmfit` tool estimates LLM throughput using a bandwidth-based model where accuracy depends on two tunable parameters: an **efficiency factor** that accounts for kernel overhead and memory inefficiencies, and **run-mode factors** that adjust for different execution paths (GPU, CPU offload, Tensor Parallel, etc.). This guide shows how to configure both for precise TPS calibration on your hardware.

---

## Understanding the TPS Calculation Model

`llmfit` computes estimated throughput as:

```

TPS ≈ bandwidth × efficiency × run_mode_factor / model_size_GB

```

Three inputs drive this formula:

- **Bandwidth** — theoretical or measured memory bandwidth (GB/s)
- **Efficiency factor** — scalar (default ≈ 0.55) capturing kernel launch overhead, KV-cache reads, and memory controller inefficiencies
- **Run-mode factor** — multiplier specific to the execution path (GPU, CPU-offload, MoE-offload, Tensor-Parallel, CPU-only)

Both the **efficiency** and **run-mode factors** are user-tunable. Adjusting them aligns `llmfit` predictions with your actual benchmarked performance.

---

## Where Configuration Parameters Live

The core configuration resides in **[`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs)**:

```rust
// CalcConfig struct (fit.rs)
pub struct CalcConfig {
    pub efficiency: f64,                 // default 0.55
    pub run_mode_factors: RunModeFactors,
    // ... additional fields
}

// RunModeFactors struct with per-mode multipliers
pub struct RunModeFactors {
    pub gpu: f64,
    pub cpu_offload: f64,
    pub moe_offload: f64,
    pub tensor_parallel: f64,
    pub cpu_only: f64,
}

```

The calculation logic applies these values in two locations:

1. **[`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs)** — `raw_tps` computation multiplies base bandwidth by `config.efficiency` (around line 1207)
2. **[`plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/plan.rs)** — final plan multiplies by `config.run_mode_factors.for_run_mode(path.run_mode())` (lines 1277-1412)

---

## Configuration Methods

### CLI: Fast Parameter Overrides

Use command-line flags for quick adjustments:

```bash

# Set efficiency to 80% (0.80)

llmfit --efficiency 80

# Override specific run-mode factors

llmfit --efficiency 75 --run-mode-factor gpu=1.10 --run-mode-factor cpu_offload=0.85

```

The `--efficiency` flag accepts a percentage (55 = 0.55 default). The `--run-mode-factor` flag accepts `<MODE>=<VALUE>` pairs for any of the five run modes.

---

### TUI: Interactive Configuration

Press **`A`** in the `llmfit` terminal interface to open the **Advanced Configuration** panel. This exposes:

- **Efficiency** field — editable percentage value
- **Five Run-Mode fields** — GPU, CPU-offload, MoE-offload, Tensor-Parallel, CPU-only

The UI writes values directly back into `calc_config`. Implementation details are in **[`llmfit-tui/src/tui_app.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_app.rs)** (lines 3822-3841 for field updates, line 4146 for config serialization).

---

### HTTP API: JSON Payload

When calling `/api/v1/fit`, include the parameters in the request body:

```json
{
  "model_specs": [...],
  "efficiency": 0.85,
  "runModeFactors": {
    "gpu": 1.0,
    "cpuOffload": 0.90,
    "moeOffload": 0.85,
    "tensorParallel": 0.75,
    "cpuOnly": 0.60
  }
}

```

See **[`API.md`](https://github.com/AlexsJones/llmfit/blob/main/API.md)** for complete request/response schemas.

---

### Programmatic Rust API

Construct `CalcConfig` directly for embedded use:

```rust
use llmfit_core::fit::{CalcConfig, RunModeFactors};

let cfg = CalcConfig {
    efficiency: 0.80,
    run_mode_factors: RunModeFactors {
        gpu: 1.00,
        cpu_offload: 0.90,
        moe_offload: 0.85,
        tensor_parallel: 0.75,
        cpu_only: 0.60,
    },
    ..Default::default()  // fills remaining fields with defaults
};

let results = llmfit_core::analysis::build_model_fits(&specs, &db, &cfg);

```

The `..Default::default()` syntax preserves default values for scoring weights and other configuration fields.

---

## When to Tune These Parameters

| Scenario | Recommended Action |
|----------|------------------|
| **Hardware-specific variance** | Apple Silicon unified memory, AMD GPUs, and older NVIDIA cards show different effective bandwidths—adjust `efficiency` to match measured performance |
| **Benchmark alignment** | Back-solve required efficiency/run-mode factor from measured TPS, then set values so `llmfit` reproduces those numbers for similar models |
| **Experimental runtimes** | Custom containers or new inference engines may have different overhead—tweak run-mode factors to let the planner favor appropriate paths |

---

## Source Code Reference

| File | Purpose | Key Location |
|------|---------|--------------|
| [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) | `CalcConfig`, `RunModeFactors`, default values, `raw_tps` calculation | Lines 1207 (efficiency application), 1277-1412 (run-mode factor selection) |
| [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs) | Applies run-mode factor to planning pipeline | `config.run_mode_factors.for_run_mode()` call site |
| [`llmfit-tui/src/tui_app.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_app.rs) | UI state for Advanced Configuration popup | Lines 3822-3841 (field bindings), 4146 (config write-back) |
| [`docs/tui.md`](https://github.com/AlexsJones/llmfit/blob/main/docs/tui.md) | TUI interaction documentation | `A` key binding for advanced settings |
| [`docs/how-it-works.md`](https://github.com/AlexsJones/llmfit/blob/main/docs/how-it-works.md) | TPS formula explanation | Efficiency and run-mode factor theory |
| [`API.md`](https://github.com/AlexsJones/llmfit/blob/main/API.md) | JSON API specification | `efficiency` and `runModeFactors` payload fields |

---

## Summary

- **`efficiency`** and **`run_mode_factors`** in `CalcConfig` control `llmfit` TPS calibration
- Default efficiency is **0.55** (55%); run-mode factors have mode-specific defaults defined in `RunModeFactors::default()`
- Configure via **CLI flags**, **TUI (`A` key)**, **HTTP API JSON**, or **direct Rust API**
- Apply adjustments when hardware differs from reference platforms or when aligning against your own benchmarks
- Core implementation spans [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) (calculation), [`plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/plan.rs) (planning integration), and [`tui_app.rs`](https://github.com/AlexsJones/llmfit/blob/main/tui_app.rs) (interactive UI)

---

## Frequently Asked Questions

### How do I find the right efficiency value for my GPU?

Run a benchmark on a known model, then solve backward: `efficiency = measured_TPS × model_size_GB / bandwidth`. Set this value and verify against other models of similar size. According to the `llmfit` source code in [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) (lines 1209-1212), the default 0.55 accounts for kernel launch overhead, KV-cache reads, and memory controller inefficiencies, but your specific hardware may differ.

### Can I set different run-mode factors for different models?

No—run-mode factors are global configuration values in `CalcConfig`, not per-model parameters. However, you can invoke `llmfit` with different `--run-mode-factor` flags per invocation, or maintain separate configuration objects when using the programmatic API. The planner in [`plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/plan.rs) applies the same `RunModeFactors` across all candidate paths for a single analysis run.

### Why does the TUI use percentage for efficiency but decimals in the API?

The TUI presents efficiency as a percentage (55) for user familiarity, while the underlying `CalcConfig.efficiency` field stores it as a decimal (0.55). The CLI `--efficiency` flag also accepts percentages. The HTTP API and Rust API use the native decimal representation. Conversion happens at the interface layer in [`tui_app.rs`](https://github.com/AlexsJones/llmfit/blob/main/tui_app.rs) and the CLI argument parser.