# llmfit Run Modes: Speed Multipliers and Performance Factors Explained

> Discover llmfit run modes and their speed multipliers from 1.0 on GPU to 0.3 on CPU. Understand performance factors impacting token generation.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: deep-dive
- Published: 2026-08-22

---

**llmfit applies hardware-specific speed multipliers ranging from 1.0 for GPU mode down to 0.3 for CPU-only execution, adjusting baseline token-per-second estimates based on deployment topology.**

The `llmfit` library by AlexsJones estimates large language model throughput by combining hardware bandwidth calculations with deployment-specific efficiency factors. These **speed multipliers** scale baseline performance predictions to reflect real-world overhead from distributed inference, memory offloading, and compute-device selection.

## Default Speed Multipliers by Run Mode

The `RunModeFactors` struct in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) (lines 65‑71) defines the default scaling values used across the analysis pipeline.

| llmfit Run Mode | Speed Multiplier | Performance Characteristic |
|-----------------|------------------|---------------------------|
| **GPU** | `1.0` | native device execution |
| **Tensor Parallel** | `0.9` | multi-device communication overhead |
| **MoE Offload** | `0.8` | expert parameter streaming to RAM |
| **CPU Offload** | `0.5` | mixed GPU/DDR bandwidth limitations |
| **CPU Only** | `0.3` | pure system memory inference |

These constants represent the relative efficiency of each **llmfit run mode** compared to unconstrained GPU execution.

## Architecture Rationale Behind the Multipliers

The library selects these values based on the data-path constraints inherent to each deployment strategy.

- **GPU (1.0)** serves as the baseline, representing maximum throughput when the entire model resides in accelerator VRAM with no cross-device traffic.

- **Tensor Parallel (0.9)** accounts for the latency penalty of splitting layers across multiple GPUs. The slight 10 % reduction reflects inter-GPU synchronization overhead mentioned in the source comments.

- **MoE Offload (0.8)** handles Mixture-of-Experts models where inactive expert parameters stream from system RAM. The 20 % penalty captures the effective bandwidth reduction compared to resident GPU storage.

- **CPU Offload (0.5)** applies to hybrid execution contexts where activations or weights shuttle between GPU and CPU memory. This heavier 50 % penalty reflects DDR bandwidth bottlenecks relative to HBM/VRAM.

- **CPU Only (0.3)** represents the most constrained path, running the full forward pass on system RAM without accelerator assistance, limited by CPU compute throughput and memory latency.

## Implementation in the llmfit Core

The multiplier system centers on two key components in the Rust source: the factor definition struct and the mode-selection method.

### RunModeFactors Struct Definition

In [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs), the struct implements `Default` to supply the baseline values:

```rust
// Lines 65-71 in llmfit-core/src/fit.rs
impl Default for RunModeFactors {
    fn default() -> Self {
        RunModeFactors {
            gpu: 1.0,
            tensor_parallel: 0.9,
            moe_offload: 0.8,
            cpu_offload: 0.5,
            cpu_only: 0.3,
        }
    }
}

```

### Selecting the Active Multiplier

The `for_run_mode` method (lines 84‑91 in the same file) returns the appropriate factor for a given enumeration variant:

```rust
pub fn for_run_mode(&self, mode: RunMode) -> f64 {
    match mode {
        RunMode::Gpu => self.gpu,
        RunMode::TensorParallel => self.tensor_parallel,
        RunMode::MoeOffload => self.moe_offload,
        RunMode::CpuOffload => self.cpu_offload,
        RunMode::CpuOnly => self.cpu_only,
    }
}

```

### Application in the Analysis Pipeline

The multiplier applies at two critical points in the estimation workflow:

1. **Model Analysis** – Inside `ModelFit::analyze_inner`, the factor scales raw TPS estimates derived from hardware bandwidth metrics.
2. **Execution Planning** – In [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs) (line 288), the multiplier adjusts predicted throughput when building execution plans for specific deployment scenarios.

## Configuring Custom Speed Multipliers

Users can override the defaults programmatically via `CalcConfig` or through the TUI advanced settings panel ([`llmfit-tui/src/tui_app.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_app.rs), lines 3823‑3831).

### Retrieving Baseline Factors

```rust
use llmfit_core::fit::{RunMode, RunModeFactors};

// Initialize with library defaults
let factors = RunModeFactors::default();
println!("GPU baseline: {}", factors.gpu);
println!("CPU-only baseline: {}", factors.cpu_only);

// Query specific mode
let mode = RunMode::MoeOffload;
let multiplier = factors.for_run_mode(mode);
println!("Multiplier for {:?}: {}", mode, multiplier);

```

### Applying Multipliers to Throughput Estimates

```rust
// Baseline calculated from memory bandwidth (GB/s) * efficiency
let baseline_tps: f64 = 120.0; 
let adjusted_tps = baseline_tps * multiplier;

println!("Raw TPS: {:.1}, Adjusted TPS: {:.1}", baseline_tps, adjusted_tps);

```

### Custom CalcConfig for Non-Standard Hardware

For clusters with atypical interconnects or memory hierarchies, instantiate a custom configuration:

```rust
use llmfit_core::fit::{CalcConfig, RunModeFactors, ModelFit};

let mut cfg = CalcConfig::default();
cfg.run_mode_factors = RunModeFactors {
    gpu: 1.0,
    tensor_parallel: 0.85,  // tighter NVLink penalty
    moe_offload: 0.75,      // slower DDR5 channel
    cpu_offload: 0.45,
    cpu_only: 0.25,
};

let fit = ModelFit::analyze_with_config(&model, &system, cfg);
println!("Estimated GPU throughput: {} tokens/sec", fit.estimated_tps);

```

## Summary

- **llmfit run modes** use defined speed multipliers in `RunModeFactors` to adjust raw hardware bandwidth estimates.
- Default values span `1.0` (GPU) to `0.3` (CPU Only), reflecting topology-specific overhead.
- The selection logic lives in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) (`for_run_mode` method), with application in the analysis pipeline and planner.
- Override defaults via `CalcConfig` for custom hardware or advanced tuning scenarios.

## Frequently Asked Questions

### How does llmfit calculate the final tokens-per-second estimate?

llmfit first computes a theoretical baseline from hardware specifications (memory bandwidth, compute TFLOPS, and model parameter count), then multiplies that value by the **speed multiplier** corresponding to the selected run mode. This adjusted figure appears as `estimated_tps` in the `ModelFit` output struct.

### Can I adjust speed multipliers without recompiling llmfit?

Yes. While the Rust source defines defaults in `RunModeFactors::default()`, the library exposes these values through the `CalcConfig` API at runtime. Additionally, the TUI interface in [`llmfit-tui/src/tui_app.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_app.rs) provides input fields to tweak multipliers interactively before running analysis.

### Why is the Tensor Parallel multiplier only 0.9 instead of a larger penalty?

The `0.9` value in **llmfit run modes** reflects efficient all-reduce collectives over high-bandwidth interconnects (NVLink/InfiniBand). The 10 % overhead accounts for synchronization latency and gradient bucketing, but assumes optimal device placement. Heavily network-constrained clusters may require lowering this value via custom configuration.

### Where does the multiplier apply during model planning?

The factor is applied in [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs) (approximately line 288) when the planner converts theoretical throughput into layer execution schedules. This ensures that pipeline partitioning and batch-size decisions account for the reduced effective bandwidth of offloaded or parallelized execution paths.