How to Configure llmfit Efficiency and Run Mode Factors for TPS Calibration
Set the efficiency scalar and run_mode_factors multipliers in CalcConfig to calibrate llmfit's token-per-second predictions against real hardware measurements.
The llmfit tool estimates LLM throughput using a bandwidth-based model where accuracy depends on two tunable parameters: an efficiency factor that accounts for kernel overhead and memory inefficiencies, and run-mode factors that adjust for different execution paths (GPU, CPU offload, Tensor Parallel, etc.). This guide shows how to configure both for precise TPS calibration on your hardware.
Understanding the TPS Calculation Model
llmfit computes estimated throughput as:
TPS ≈ bandwidth × efficiency × run_mode_factor / model_size_GB
Three inputs drive this formula:
- Bandwidth — theoretical or measured memory bandwidth (GB/s)
- Efficiency factor — scalar (default ≈ 0.55) capturing kernel launch overhead, KV-cache reads, and memory controller inefficiencies
- Run-mode factor — multiplier specific to the execution path (GPU, CPU-offload, MoE-offload, Tensor-Parallel, CPU-only)
Both the efficiency and run-mode factors are user-tunable. Adjusting them aligns llmfit predictions with your actual benchmarked performance.
Where Configuration Parameters Live
The core configuration resides in llmfit-core/src/fit.rs:
// CalcConfig struct (fit.rs)
pub struct CalcConfig {
pub efficiency: f64, // default 0.55
pub run_mode_factors: RunModeFactors,
// ... additional fields
}
// RunModeFactors struct with per-mode multipliers
pub struct RunModeFactors {
pub gpu: f64,
pub cpu_offload: f64,
pub moe_offload: f64,
pub tensor_parallel: f64,
pub cpu_only: f64,
}
The calculation logic applies these values in two locations:
fit.rs—raw_tpscomputation multiplies base bandwidth byconfig.efficiency(around line 1207)plan.rs— final plan multiplies byconfig.run_mode_factors.for_run_mode(path.run_mode())(lines 1277-1412)
Configuration Methods
CLI: Fast Parameter Overrides
Use command-line flags for quick adjustments:
# Set efficiency to 80% (0.80)
llmfit --efficiency 80
# Override specific run-mode factors
llmfit --efficiency 75 --run-mode-factor gpu=1.10 --run-mode-factor cpu_offload=0.85
The --efficiency flag accepts a percentage (55 = 0.55 default). The --run-mode-factor flag accepts <MODE>=<VALUE> pairs for any of the five run modes.
TUI: Interactive Configuration
Press A in the llmfit terminal interface to open the Advanced Configuration panel. This exposes:
- Efficiency field — editable percentage value
- Five Run-Mode fields — GPU, CPU-offload, MoE-offload, Tensor-Parallel, CPU-only
The UI writes values directly back into calc_config. Implementation details are in llmfit-tui/src/tui_app.rs (lines 3822-3841 for field updates, line 4146 for config serialization).
HTTP API: JSON Payload
When calling /api/v1/fit, include the parameters in the request body:
{
"model_specs": [...],
"efficiency": 0.85,
"runModeFactors": {
"gpu": 1.0,
"cpuOffload": 0.90,
"moeOffload": 0.85,
"tensorParallel": 0.75,
"cpuOnly": 0.60
}
}
See API.md for complete request/response schemas.
Programmatic Rust API
Construct CalcConfig directly for embedded use:
use llmfit_core::fit::{CalcConfig, RunModeFactors};
let cfg = CalcConfig {
efficiency: 0.80,
run_mode_factors: RunModeFactors {
gpu: 1.00,
cpu_offload: 0.90,
moe_offload: 0.85,
tensor_parallel: 0.75,
cpu_only: 0.60,
},
..Default::default() // fills remaining fields with defaults
};
let results = llmfit_core::analysis::build_model_fits(&specs, &db, &cfg);
The ..Default::default() syntax preserves default values for scoring weights and other configuration fields.
When to Tune These Parameters
| Scenario | Recommended Action |
|---|---|
| Hardware-specific variance | Apple Silicon unified memory, AMD GPUs, and older NVIDIA cards show different effective bandwidths—adjust efficiency to match measured performance |
| Benchmark alignment | Back-solve required efficiency/run-mode factor from measured TPS, then set values so llmfit reproduces those numbers for similar models |
| Experimental runtimes | Custom containers or new inference engines may have different overhead—tweak run-mode factors to let the planner favor appropriate paths |
Source Code Reference
| File | Purpose | Key Location |
|---|---|---|
llmfit-core/src/fit.rs |
CalcConfig, RunModeFactors, default values, raw_tps calculation |
Lines 1207 (efficiency application), 1277-1412 (run-mode factor selection) |
llmfit-core/src/plan.rs |
Applies run-mode factor to planning pipeline | config.run_mode_factors.for_run_mode() call site |
llmfit-tui/src/tui_app.rs |
UI state for Advanced Configuration popup | Lines 3822-3841 (field bindings), 4146 (config write-back) |
docs/tui.md |
TUI interaction documentation | A key binding for advanced settings |
docs/how-it-works.md |
TPS formula explanation | Efficiency and run-mode factor theory |
API.md |
JSON API specification | efficiency and runModeFactors payload fields |
Summary
efficiencyandrun_mode_factorsinCalcConfigcontrolllmfitTPS calibration- Default efficiency is 0.55 (55%); run-mode factors have mode-specific defaults defined in
RunModeFactors::default() - Configure via CLI flags, TUI (
Akey), HTTP API JSON, or direct Rust API - Apply adjustments when hardware differs from reference platforms or when aligning against your own benchmarks
- Core implementation spans
fit.rs(calculation),plan.rs(planning integration), andtui_app.rs(interactive UI)
Frequently Asked Questions
How do I find the right efficiency value for my GPU?
Run a benchmark on a known model, then solve backward: efficiency = measured_TPS × model_size_GB / bandwidth. Set this value and verify against other models of similar size. According to the llmfit source code in fit.rs (lines 1209-1212), the default 0.55 accounts for kernel launch overhead, KV-cache reads, and memory controller inefficiencies, but your specific hardware may differ.
Can I set different run-mode factors for different models?
No—run-mode factors are global configuration values in CalcConfig, not per-model parameters. However, you can invoke llmfit with different --run-mode-factor flags per invocation, or maintain separate configuration objects when using the programmatic API. The planner in plan.rs applies the same RunModeFactors across all candidate paths for a single analysis run.
Why does the TUI use percentage for efficiency but decimals in the API?
The TUI presents efficiency as a percentage (55) for user familiarity, while the underlying CalcConfig.efficiency field stores it as a decimal (0.55). The CLI --efficiency flag also accepts percentages. The HTTP API and Rust API use the native decimal representation. Conversion happens at the interface layer in tui_app.rs and the CLI argument parser.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →