llmfit Run Modes: Speed Multipliers and Performance Factors Explained

llmfit applies hardware-specific speed multipliers ranging from 1.0 for GPU mode down to 0.3 for CPU-only execution, adjusting baseline token-per-second estimates based on deployment topology.

The llmfit library by AlexsJones estimates large language model throughput by combining hardware bandwidth calculations with deployment-specific efficiency factors. These speed multipliers scale baseline performance predictions to reflect real-world overhead from distributed inference, memory offloading, and compute-device selection.

Default Speed Multipliers by Run Mode

The RunModeFactors struct in llmfit-core/src/fit.rs (lines 65‑71) defines the default scaling values used across the analysis pipeline.

llmfit Run Mode Speed Multiplier Performance Characteristic
GPU 1.0 native device execution
Tensor Parallel 0.9 multi-device communication overhead
MoE Offload 0.8 expert parameter streaming to RAM
CPU Offload 0.5 mixed GPU/DDR bandwidth limitations
CPU Only 0.3 pure system memory inference

These constants represent the relative efficiency of each llmfit run mode compared to unconstrained GPU execution.

Architecture Rationale Behind the Multipliers

The library selects these values based on the data-path constraints inherent to each deployment strategy.

  • GPU (1.0) serves as the baseline, representing maximum throughput when the entire model resides in accelerator VRAM with no cross-device traffic.

  • Tensor Parallel (0.9) accounts for the latency penalty of splitting layers across multiple GPUs. The slight 10 % reduction reflects inter-GPU synchronization overhead mentioned in the source comments.

  • MoE Offload (0.8) handles Mixture-of-Experts models where inactive expert parameters stream from system RAM. The 20 % penalty captures the effective bandwidth reduction compared to resident GPU storage.

  • CPU Offload (0.5) applies to hybrid execution contexts where activations or weights shuttle between GPU and CPU memory. This heavier 50 % penalty reflects DDR bandwidth bottlenecks relative to HBM/VRAM.

  • CPU Only (0.3) represents the most constrained path, running the full forward pass on system RAM without accelerator assistance, limited by CPU compute throughput and memory latency.

Implementation in the llmfit Core

The multiplier system centers on two key components in the Rust source: the factor definition struct and the mode-selection method.

RunModeFactors Struct Definition

In llmfit-core/src/fit.rs, the struct implements Default to supply the baseline values:

// Lines 65-71 in llmfit-core/src/fit.rs
impl Default for RunModeFactors {
    fn default() -> Self {
        RunModeFactors {
            gpu: 1.0,
            tensor_parallel: 0.9,
            moe_offload: 0.8,
            cpu_offload: 0.5,
            cpu_only: 0.3,
        }
    }
}

Selecting the Active Multiplier

The for_run_mode method (lines 84‑91 in the same file) returns the appropriate factor for a given enumeration variant:

pub fn for_run_mode(&self, mode: RunMode) -> f64 {
    match mode {
        RunMode::Gpu => self.gpu,
        RunMode::TensorParallel => self.tensor_parallel,
        RunMode::MoeOffload => self.moe_offload,
        RunMode::CpuOffload => self.cpu_offload,
        RunMode::CpuOnly => self.cpu_only,
    }
}

Application in the Analysis Pipeline

The multiplier applies at two critical points in the estimation workflow:

  1. Model Analysis – Inside ModelFit::analyze_inner, the factor scales raw TPS estimates derived from hardware bandwidth metrics.
  2. Execution Planning – In llmfit-core/src/plan.rs (line 288), the multiplier adjusts predicted throughput when building execution plans for specific deployment scenarios.

Configuring Custom Speed Multipliers

Users can override the defaults programmatically via CalcConfig or through the TUI advanced settings panel (llmfit-tui/src/tui_app.rs, lines 3823‑3831).

Retrieving Baseline Factors

use llmfit_core::fit::{RunMode, RunModeFactors};

// Initialize with library defaults
let factors = RunModeFactors::default();
println!("GPU baseline: {}", factors.gpu);
println!("CPU-only baseline: {}", factors.cpu_only);

// Query specific mode
let mode = RunMode::MoeOffload;
let multiplier = factors.for_run_mode(mode);
println!("Multiplier for {:?}: {}", mode, multiplier);

Applying Multipliers to Throughput Estimates

// Baseline calculated from memory bandwidth (GB/s) * efficiency
let baseline_tps: f64 = 120.0; 
let adjusted_tps = baseline_tps * multiplier;

println!("Raw TPS: {:.1}, Adjusted TPS: {:.1}", baseline_tps, adjusted_tps);

Custom CalcConfig for Non-Standard Hardware

For clusters with atypical interconnects or memory hierarchies, instantiate a custom configuration:

use llmfit_core::fit::{CalcConfig, RunModeFactors, ModelFit};

let mut cfg = CalcConfig::default();
cfg.run_mode_factors = RunModeFactors {
    gpu: 1.0,
    tensor_parallel: 0.85,  // tighter NVLink penalty
    moe_offload: 0.75,      // slower DDR5 channel
    cpu_offload: 0.45,
    cpu_only: 0.25,
};

let fit = ModelFit::analyze_with_config(&model, &system, cfg);
println!("Estimated GPU throughput: {} tokens/sec", fit.estimated_tps);

Summary

  • llmfit run modes use defined speed multipliers in RunModeFactors to adjust raw hardware bandwidth estimates.
  • Default values span 1.0 (GPU) to 0.3 (CPU Only), reflecting topology-specific overhead.
  • The selection logic lives in llmfit-core/src/fit.rs (for_run_mode method), with application in the analysis pipeline and planner.
  • Override defaults via CalcConfig for custom hardware or advanced tuning scenarios.

Frequently Asked Questions

How does llmfit calculate the final tokens-per-second estimate?

llmfit first computes a theoretical baseline from hardware specifications (memory bandwidth, compute TFLOPS, and model parameter count), then multiplies that value by the speed multiplier corresponding to the selected run mode. This adjusted figure appears as estimated_tps in the ModelFit output struct.

Can I adjust speed multipliers without recompiling llmfit?

Yes. While the Rust source defines defaults in RunModeFactors::default(), the library exposes these values through the CalcConfig API at runtime. Additionally, the TUI interface in llmfit-tui/src/tui_app.rs provides input fields to tweak multipliers interactively before running analysis.

Why is the Tensor Parallel multiplier only 0.9 instead of a larger penalty?

The 0.9 value in llmfit run modes reflects efficient all-reduce collectives over high-bandwidth interconnects (NVLink/InfiniBand). The 10 % overhead accounts for synchronization latency and gradient bucketing, but assumes optimal device placement. Heavily network-constrained clusters may require lowering this value via custom configuration.

Where does the multiplier apply during model planning?

The factor is applied in llmfit-core/src/plan.rs (approximately line 288) when the planner converts theoretical throughput into layer execution schedules. This ensures that pipeline partitioning and batch-size decisions account for the reduced effective bandwidth of offloaded or parallelized execution paths.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →