How to Configure llmfit Efficiency and Run Mode Factors for TPS Calibration

Set the efficiency scalar and run_mode_factors multipliers in CalcConfig to calibrate llmfit's token-per-second predictions against real hardware measurements.

The llmfit tool estimates LLM throughput using a bandwidth-based model where accuracy depends on two tunable parameters: an efficiency factor that accounts for kernel overhead and memory inefficiencies, and run-mode factors that adjust for different execution paths (GPU, CPU offload, Tensor Parallel, etc.). This guide shows how to configure both for precise TPS calibration on your hardware.


Understanding the TPS Calculation Model

llmfit computes estimated throughput as:


TPS ≈ bandwidth × efficiency × run_mode_factor / model_size_GB

Three inputs drive this formula:

  • Bandwidth — theoretical or measured memory bandwidth (GB/s)
  • Efficiency factor — scalar (default ≈ 0.55) capturing kernel launch overhead, KV-cache reads, and memory controller inefficiencies
  • Run-mode factor — multiplier specific to the execution path (GPU, CPU-offload, MoE-offload, Tensor-Parallel, CPU-only)

Both the efficiency and run-mode factors are user-tunable. Adjusting them aligns llmfit predictions with your actual benchmarked performance.


Where Configuration Parameters Live

The core configuration resides in llmfit-core/src/fit.rs:

// CalcConfig struct (fit.rs)
pub struct CalcConfig {
    pub efficiency: f64,                 // default 0.55
    pub run_mode_factors: RunModeFactors,
    // ... additional fields
}

// RunModeFactors struct with per-mode multipliers
pub struct RunModeFactors {
    pub gpu: f64,
    pub cpu_offload: f64,
    pub moe_offload: f64,
    pub tensor_parallel: f64,
    pub cpu_only: f64,
}

The calculation logic applies these values in two locations:

  1. fit.rs — raw_tps computation multiplies base bandwidth by config.efficiency (around line 1207)
  2. plan.rs — final plan multiplies by config.run_mode_factors.for_run_mode(path.run_mode()) (lines 1277-1412)

Configuration Methods

CLI: Fast Parameter Overrides

Use command-line flags for quick adjustments:


# Set efficiency to 80% (0.80)

llmfit --efficiency 80

# Override specific run-mode factors

llmfit --efficiency 75 --run-mode-factor gpu=1.10 --run-mode-factor cpu_offload=0.85

The --efficiency flag accepts a percentage (55 = 0.55 default). The --run-mode-factor flag accepts <MODE>=<VALUE> pairs for any of the five run modes.


TUI: Interactive Configuration

Press A in the llmfit terminal interface to open the Advanced Configuration panel. This exposes:

  • Efficiency field — editable percentage value
  • Five Run-Mode fields — GPU, CPU-offload, MoE-offload, Tensor-Parallel, CPU-only

The UI writes values directly back into calc_config. Implementation details are in llmfit-tui/src/tui_app.rs (lines 3822-3841 for field updates, line 4146 for config serialization).


HTTP API: JSON Payload

When calling /api/v1/fit, include the parameters in the request body:

{
  "model_specs": [...],
  "efficiency": 0.85,
  "runModeFactors": {
    "gpu": 1.0,
    "cpuOffload": 0.90,
    "moeOffload": 0.85,
    "tensorParallel": 0.75,
    "cpuOnly": 0.60
  }
}

See API.md for complete request/response schemas.


Programmatic Rust API

Construct CalcConfig directly for embedded use:

use llmfit_core::fit::{CalcConfig, RunModeFactors};

let cfg = CalcConfig {
    efficiency: 0.80,
    run_mode_factors: RunModeFactors {
        gpu: 1.00,
        cpu_offload: 0.90,
        moe_offload: 0.85,
        tensor_parallel: 0.75,
        cpu_only: 0.60,
    },
    ..Default::default()  // fills remaining fields with defaults
};

let results = llmfit_core::analysis::build_model_fits(&specs, &db, &cfg);

The ..Default::default() syntax preserves default values for scoring weights and other configuration fields.


When to Tune These Parameters

Scenario Recommended Action
Hardware-specific variance Apple Silicon unified memory, AMD GPUs, and older NVIDIA cards show different effective bandwidths—adjust efficiency to match measured performance
Benchmark alignment Back-solve required efficiency/run-mode factor from measured TPS, then set values so llmfit reproduces those numbers for similar models
Experimental runtimes Custom containers or new inference engines may have different overhead—tweak run-mode factors to let the planner favor appropriate paths

Source Code Reference

File Purpose Key Location
llmfit-core/src/fit.rs CalcConfig, RunModeFactors, default values, raw_tps calculation Lines 1207 (efficiency application), 1277-1412 (run-mode factor selection)
llmfit-core/src/plan.rs Applies run-mode factor to planning pipeline config.run_mode_factors.for_run_mode() call site
llmfit-tui/src/tui_app.rs UI state for Advanced Configuration popup Lines 3822-3841 (field bindings), 4146 (config write-back)
docs/tui.md TUI interaction documentation A key binding for advanced settings
docs/how-it-works.md TPS formula explanation Efficiency and run-mode factor theory
API.md JSON API specification efficiency and runModeFactors payload fields

Summary

  • efficiency and run_mode_factors in CalcConfig control llmfit TPS calibration
  • Default efficiency is 0.55 (55%); run-mode factors have mode-specific defaults defined in RunModeFactors::default()
  • Configure via CLI flags, TUI (A key), HTTP API JSON, or direct Rust API
  • Apply adjustments when hardware differs from reference platforms or when aligning against your own benchmarks
  • Core implementation spans fit.rs (calculation), plan.rs (planning integration), and tui_app.rs (interactive UI)

Frequently Asked Questions

How do I find the right efficiency value for my GPU?

Run a benchmark on a known model, then solve backward: efficiency = measured_TPS × model_size_GB / bandwidth. Set this value and verify against other models of similar size. According to the llmfit source code in fit.rs (lines 1209-1212), the default 0.55 accounts for kernel launch overhead, KV-cache reads, and memory controller inefficiencies, but your specific hardware may differ.

Can I set different run-mode factors for different models?

No—run-mode factors are global configuration values in CalcConfig, not per-model parameters. However, you can invoke llmfit with different --run-mode-factor flags per invocation, or maintain separate configuration objects when using the programmatic API. The planner in plan.rs applies the same RunModeFactors across all candidate paths for a single analysis run.

Why does the TUI use percentage for efficiency but decimals in the API?

The TUI presents efficiency as a percentage (55) for user familiarity, while the underlying CalcConfig.efficiency field stores it as a decimal (0.55). The CLI --efficiency flag also accepts percentages. The HTTP API and Rust API use the native decimal representation. Conversion happens at the interface layer in tui_app.rs and the CLI argument parser.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →