llmfit Run Modes: Speed Multipliers and Performance Factors Explained
llmfit applies hardware-specific speed multipliers ranging from 1.0 for GPU mode down to 0.3 for CPU-only execution, adjusting baseline token-per-second estimates based on deployment topology.
The llmfit library by AlexsJones estimates large language model throughput by combining hardware bandwidth calculations with deployment-specific efficiency factors. These speed multipliers scale baseline performance predictions to reflect real-world overhead from distributed inference, memory offloading, and compute-device selection.
Default Speed Multipliers by Run Mode
The RunModeFactors struct in llmfit-core/src/fit.rs (lines 65‑71) defines the default scaling values used across the analysis pipeline.
| llmfit Run Mode | Speed Multiplier | Performance Characteristic |
|---|---|---|
| GPU | 1.0 |
native device execution |
| Tensor Parallel | 0.9 |
multi-device communication overhead |
| MoE Offload | 0.8 |
expert parameter streaming to RAM |
| CPU Offload | 0.5 |
mixed GPU/DDR bandwidth limitations |
| CPU Only | 0.3 |
pure system memory inference |
These constants represent the relative efficiency of each llmfit run mode compared to unconstrained GPU execution.
Architecture Rationale Behind the Multipliers
The library selects these values based on the data-path constraints inherent to each deployment strategy.
-
GPU (1.0) serves as the baseline, representing maximum throughput when the entire model resides in accelerator VRAM with no cross-device traffic.
-
Tensor Parallel (0.9) accounts for the latency penalty of splitting layers across multiple GPUs. The slight 10 % reduction reflects inter-GPU synchronization overhead mentioned in the source comments.
-
MoE Offload (0.8) handles Mixture-of-Experts models where inactive expert parameters stream from system RAM. The 20 % penalty captures the effective bandwidth reduction compared to resident GPU storage.
-
CPU Offload (0.5) applies to hybrid execution contexts where activations or weights shuttle between GPU and CPU memory. This heavier 50 % penalty reflects DDR bandwidth bottlenecks relative to HBM/VRAM.
-
CPU Only (0.3) represents the most constrained path, running the full forward pass on system RAM without accelerator assistance, limited by CPU compute throughput and memory latency.
Implementation in the llmfit Core
The multiplier system centers on two key components in the Rust source: the factor definition struct and the mode-selection method.
RunModeFactors Struct Definition
In llmfit-core/src/fit.rs, the struct implements Default to supply the baseline values:
// Lines 65-71 in llmfit-core/src/fit.rs
impl Default for RunModeFactors {
fn default() -> Self {
RunModeFactors {
gpu: 1.0,
tensor_parallel: 0.9,
moe_offload: 0.8,
cpu_offload: 0.5,
cpu_only: 0.3,
}
}
}
Selecting the Active Multiplier
The for_run_mode method (lines 84‑91 in the same file) returns the appropriate factor for a given enumeration variant:
pub fn for_run_mode(&self, mode: RunMode) -> f64 {
match mode {
RunMode::Gpu => self.gpu,
RunMode::TensorParallel => self.tensor_parallel,
RunMode::MoeOffload => self.moe_offload,
RunMode::CpuOffload => self.cpu_offload,
RunMode::CpuOnly => self.cpu_only,
}
}
Application in the Analysis Pipeline
The multiplier applies at two critical points in the estimation workflow:
- Model Analysis – Inside
ModelFit::analyze_inner, the factor scales raw TPS estimates derived from hardware bandwidth metrics. - Execution Planning – In
llmfit-core/src/plan.rs(line 288), the multiplier adjusts predicted throughput when building execution plans for specific deployment scenarios.
Configuring Custom Speed Multipliers
Users can override the defaults programmatically via CalcConfig or through the TUI advanced settings panel (llmfit-tui/src/tui_app.rs, lines 3823‑3831).
Retrieving Baseline Factors
use llmfit_core::fit::{RunMode, RunModeFactors};
// Initialize with library defaults
let factors = RunModeFactors::default();
println!("GPU baseline: {}", factors.gpu);
println!("CPU-only baseline: {}", factors.cpu_only);
// Query specific mode
let mode = RunMode::MoeOffload;
let multiplier = factors.for_run_mode(mode);
println!("Multiplier for {:?}: {}", mode, multiplier);
Applying Multipliers to Throughput Estimates
// Baseline calculated from memory bandwidth (GB/s) * efficiency
let baseline_tps: f64 = 120.0;
let adjusted_tps = baseline_tps * multiplier;
println!("Raw TPS: {:.1}, Adjusted TPS: {:.1}", baseline_tps, adjusted_tps);
Custom CalcConfig for Non-Standard Hardware
For clusters with atypical interconnects or memory hierarchies, instantiate a custom configuration:
use llmfit_core::fit::{CalcConfig, RunModeFactors, ModelFit};
let mut cfg = CalcConfig::default();
cfg.run_mode_factors = RunModeFactors {
gpu: 1.0,
tensor_parallel: 0.85, // tighter NVLink penalty
moe_offload: 0.75, // slower DDR5 channel
cpu_offload: 0.45,
cpu_only: 0.25,
};
let fit = ModelFit::analyze_with_config(&model, &system, cfg);
println!("Estimated GPU throughput: {} tokens/sec", fit.estimated_tps);
Summary
- llmfit run modes use defined speed multipliers in
RunModeFactorsto adjust raw hardware bandwidth estimates. - Default values span
1.0(GPU) to0.3(CPU Only), reflecting topology-specific overhead. - The selection logic lives in
llmfit-core/src/fit.rs(for_run_modemethod), with application in the analysis pipeline and planner. - Override defaults via
CalcConfigfor custom hardware or advanced tuning scenarios.
Frequently Asked Questions
How does llmfit calculate the final tokens-per-second estimate?
llmfit first computes a theoretical baseline from hardware specifications (memory bandwidth, compute TFLOPS, and model parameter count), then multiplies that value by the speed multiplier corresponding to the selected run mode. This adjusted figure appears as estimated_tps in the ModelFit output struct.
Can I adjust speed multipliers without recompiling llmfit?
Yes. While the Rust source defines defaults in RunModeFactors::default(), the library exposes these values through the CalcConfig API at runtime. Additionally, the TUI interface in llmfit-tui/src/tui_app.rs provides input fields to tweak multipliers interactively before running analysis.
Why is the Tensor Parallel multiplier only 0.9 instead of a larger penalty?
The 0.9 value in llmfit run modes reflects efficient all-reduce collectives over high-bandwidth interconnects (NVLink/InfiniBand). The 10 % overhead accounts for synchronization latency and gradient bucketing, but assumes optimal device placement. Heavily network-constrained clusters may require lowering this value via custom configuration.
Where does the multiplier apply during model planning?
The factor is applied in llmfit-core/src/plan.rs (approximately line 288) when the planner converts theoretical throughput into layer execution schedules. This ensures that pipeline partitioning and batch-size decisions account for the reduced effective bandwidth of offloaded or parallelized execution paths.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →