How llmfit Applies Bandwidth and Efficiency Overrides from Hardware Profiles to CalcConfig

The HardwareProfile::apply_to_config method in llmfit-core/src/hwprofile.rs translates JSON hardware specifications into runtime estimator configuration by selectively overriding fields in CalcConfig, including GPU bandwidth, system DDR bandwidth, compute throughput, efficiency factors, and run-mode speed multipliers.

The llmfit estimator separates raw hardware capacity from performance estimation logic. When a user provides a custom hardware profile, the system must bridge these layers by applying specific overrides to the calculation state. This article examines how the apply_to_config implementation maps profile data onto CalcConfig to influence throughput predictions.

Hardware Profile Structure in llmfit

A valid hardware profile contains two distinct sections that drive the override behavior. The hardware section defines physical capabilities, while the estimation section contains tuning parameters for the mathematical model.

Hardware specifications may include:

  • total_ram_gb and unified_memory flags
  • gpu_memory_bandwidth_gbps — overrides the built-in GPU bandwidth lookup table
  • ddr_bandwidth_gbps — sets system memory transfer speed for MoE offload calculations
  • gpu_compute_tflops_fp16 — switches prefill estimation from bandwidth-bound to compute-bound models

Estimation tuning parameters include:

  • efficiency — a scaling factor (default 0.55) accounting for kernel overhead and cache effects
  • run_mode_factors — per-execution-mode speed multipliers for GPU, tensor parallelism, MoE offload, CPU offload, and CPU-only paths

How apply_to_config Overrides CalcConfig

The apply_to_config(&mut self, config: &mut CalcConfig) method, defined at lines 27-63 of llmfit-core/src/hwprofile.rs, performs conditional assignments for each optional field. Absent entries preserve existing defaults, enabling partial profile configurations.

GPU Memory Bandwidth Override

When the profile specifies gpu_memory_bandwidth_gbps, the method bypasses the default resolve_gpu_bandwidth lookup and forces a specific ceiling:

if let Some(bandwidth) = hw.gpu_memory_bandwidth_gbps {
    config.gpu_bandwidth_gbps_override = Some(bandwidth);
}

This value directly constrains the roof-line model during token generation estimation.

System DDR Bandwidth Override

For machines with known DDR throughput limits, the profile can specify ddr_bandwidth_gbps. The method stores this in config.ddr_bandwidth_gbps, which the estimator later retrieves via ddr_bandwidth_gbps(&config) when calculating MoE offload transfer rates.

GPU Compute Throughput Override

The optional gpu_compute_tflops_fp16 field triggers compute-bound prefill modeling rather than pure bandwidth calculations:

if let Some(tflops) = hw.gpu_compute_tflops_fp16 {
    config.gpu_compute_tflops_fp16 = Some(tflops);
}

When present, the estimator uses this value to calculate arithmetic intensity limits during the prompt processing phase.

Efficiency Factor Calibration

The efficiency scalar models real-world overhead. The default value is 0.55, but profiles may specify higher or lower values based on observed hardware behavior:

if let Some(efficiency) = est.efficiency {
    config.efficiency = efficiency;
}

This factor scales raw bandwidth figures to account for kernel launch latency, cache coherency traffic, and PCIe overhead.

Run-Mode Speed Multipliers

The run_mode_factors struct contains optional overrides for each execution strategy. The method iterates through available factors and updates the corresponding entries in config.run_mode_factors:

if let Some(v) = factors.gpu { target.gpu = v; }
if let Some(v) = factors.tensor_parallel { target.tensor_parallel = v; }
if let Some(v) = factors.moe_offload { target.moe_offload = v; }
if let Some(v) = factors.cpu_offload { target.cpu_offload = v; }
if let Some(v) = factors.cpu_only { target.cpu_only = v; }

These multipliers are consumed by estimate_tps in llmfit-core/src/fit.rs (lines 78-84) to adjust baseline tokens-per-second estimates for each execution path.

Practical Implementation Example

The following Rust example demonstrates loading a JSON profile and applying its overrides to a fresh CalcConfig:

use llmfit_core::hwprofile::{HardwareProfile, ProfileEstimation};
use llmfit_core::fit::CalcConfig;

// Load a JSON profile with specific bandwidth and efficiency values
let json = r#"
{
  "schema_version": 1,
  "name": "high-end-workstation",
  "hardware": {
    "total_ram_gb": 128,
    "unified_memory": false,
    "gpu_memory_bandwidth_gbps": 900,
    "ddr_bandwidth_gbps": 50,
    "gpu_compute_tflops_fp16": 35
  },
  "estimation": {
    "efficiency": 0.62,
    "run_mode_factors": { "cpu_only": 0.28 }
  }
}
"#;

let profile = HardwareProfile::parse(json).expect("valid profile");
let mut cfg = CalcConfig::default();

// Apply overrides from the profile
profile.apply_to_config(&mut cfg);

// Verify the configuration state
assert_eq!(cfg.gpu_bandwidth_gbps_override, Some(900.0));
assert_eq!(cfg.ddr_bandwidth_gbps, Some(50.0));
assert_eq!(cfg.gpu_compute_tflops_fp16, Some(35.0));
assert_eq!(cfg.efficiency, 0.62);
assert_eq!(cfg.run_mode_factors.cpu_only, 0.28);

Summary

  • The HardwareProfile::apply_to_config method at lines 27-63 of llmfit-core/src/hwprofile.rs provides the bridge between static JSON configuration and runtime estimation state.
  • Bandwidth overrides (gpu_memory_bandwidth_gbps, ddr_bandwidth_gbps) replace built-in lookup values and directly influence roof-line calculations.
  • Compute overrides (gpu_compute_tflops_fp16) switch prefill estimation from bandwidth-bound to compute-bound models.
  • Efficiency factors scale raw hardware specifications to match observed real-world performance, defaulting to 0.55 when unspecified.
  • Run-mode factors allow per-execution-path tuning, consumed during the final estimate_tps calculation in fit.rs.
  • All override fields are optional; partial profiles modify only specific CalcConfig values while preserving defaults for others.

Frequently Asked Questions

What is the default efficiency factor in llmfit?

According to the source code in llmfit-core/src/hwprofile.rs, the default efficiency factor is 0.55 when no override is provided in the hardware profile's estimation section. This value accounts for kernel launch overhead, cache effects, and other real-world bandwidth inefficiencies not captured by theoretical hardware specifications.

Can I provide a partial hardware profile with only specific overrides?

Yes. The apply_to_config method uses conditional assignment for every field. If a hardware profile omits gpu_memory_bandwidth_gbps, efficiency, or any other parameter, the corresponding CalcConfig field retains its default value. This design allows users to specify only the parameters that differ from standard assumptions without defining a complete hardware specification.

How does the run_mode_factors override affect throughput estimates?

The run_mode_factors struct in the hardware profile contains optional multipliers for each execution strategy: GPU, tensor parallelism, MoE offload, CPU offload, and CPU-only modes. When present, these values overwrite the corresponding fields in config.run_mode_factors. During estimation, the estimate_tps function in llmfit-core/src/fit.rs (lines 78-84) applies these multipliers to the baseline tokens-per-second calculation, allowing users to calibrate predictions for specific deployment scenarios like cloud GPU instances or heterogeneous compute environments.

Where is the overridden bandwidth data consumed in the estimation pipeline?

Overridden bandwidth values are consumed in llmfit-core/src/fit.rs. The gpu_bandwidth_gbps_override field bypasses the standard GPU lookup table during roof-line calculations, while ddr_bandwidth_gbps is retrieved via the ddr_bandwidth_gbps(&config) helper when estimating MoE offload transfer speeds. The gpu_compute_tflops_fp16 override shifts the prefill phase calculation from bandwidth-bound to compute-bound logic based on arithmetic intensity models.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →