# How llmfit Estimates Speed on Unrecognized GPUs: 7-Step Fallback Method Explained

> Discover how llmfit estimates speed on unrecognized GPUs using a 7-step fallback method. Learn about its deterministic path and fallback techniques.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: deep-dive
- Published: 2026-08-22

---

**When llmfit encounters an unrecognized GPU, it falls back to a deterministic speed estimation path using a conservative 50 GB/s bandwidth constant scaled by a 0.55 efficiency factor, combined with model quantization data to calculate tokens-per-second (TPS).**

The **llmfit** project (AlexsJones/llmfit) provides deterministic throughput estimates for large language models even when running on hardware absent from its internal GPU database. When vendor-specific detection mechanisms—such as **nvidia-smi**, **rocm-smi**, or Vulkan—fail to identify the device, the system activates a hardware-agnostic fallback method that ensures users always receive a conservative, reproducible performance prediction.

## Step 1: Detection Failure in hardware.rs

The fallback process begins when the GPU detection code in [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs) exhausts all vendor-specific data sources. The system attempts to query **nvidia-smi** for NVIDIA cards, **rocm-smi** for AMD GPUs, and relevant **sysfs** entries before attempting a Vulkan fallback. When none of these sources return a measurable bandwidth value, the runtime proceeds to the constant-bandwidth estimation path without a measured figure.

## Step 2: Conservative Bandwidth Baseline

In the throughput-estimation routine located in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) at approximately line 1160, the code defines a hard-coded bandwidth of **50 GB/s**. This value represents the approximate bandwidth of DDR4-3200 dual-channel memory and serves as the conservative baseline whenever the GPU’s memory bandwidth cannot be derived from hardware detection.

## Step 3: Efficiency Factor Application

The constant bandwidth figure is scaled by an efficiency factor of **0.55** to account for real-world overheads including memory-controller inefficiency and address-translation penalties. This factor is defined in [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) at approximately line 1448:

```rust
let fallback_efficiency = 0.55;

```

## Step 4: K-Factor and Quantization Integration

To compute the effective GPU bandwidth, the system combines the fallback constants with model-specific data. Using a **K-factor** constant derived from known GPU profiles and the model’s quantization bytes per parameter, the calculation at line ~1450 of [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) follows this formula:

```rust
let estimated_gpu_bw = k * models::quant_bytes_per_param(quant) / fallback_efficiency;

```

This deterministic approach ensures that even without hardware-specific profiling, the tool derives a plausible bandwidth figure based on the quantization format (e.g., Q4_K_M, Q5_K_S) and the architectural K-factor.

## Step 5: Compute Time and TPS Calculation

Using the estimated bandwidth, the routine at lines ~1455-1457 calculates the GPU compute time for active parameters and derives the final **tokens-per-second** (TPS) figure. The system divides the model’s active parameter memory footprint (in GB) by the effective bandwidth to determine processing time, then converts this to throughput:

```rust
let active_params = model.active_parameters.unwrap_or(model.parameters_raw);
let active_gb = (active_params as f64 * models::quant_bpp(quant)) / 1e9;
let gpu_compute_time = active_gb / (estimated_gpu_bw * fallback_efficiency);
let tps = model.context_len as f64 / gpu_compute_time;

```

## Step 6: MoE-Specific Fallback Path

For **Mixture-of-Experts (MoE)** models, the fallback logic between lines ~884-893 in [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) implements a specialized path. If the algorithm determines that offloading cannot be satisfied for an unrecognized GPU, it falls back to treating the entire model as **RAM-only**, explicitly noting a "significantly reduced" performance estimate in the output.

## Step 7: Validation via Unit Tests

The fallback behavior is rigorously validated through unit tests in [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs) (around line 1280) and [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) (around line 3063). The test `test_bandwidth_estimation_unknown_gpu_uses_fallback` deliberately feeds unknown GPU identifiers to verify that the system correctly triggers the constant-K path and returns deterministic results.

## Practical Implementation Example

When running on a machine with an unknown GPU (for example, a brand-new NVIDIA RTX 4090 not yet in the detection database), llmfit applies the fallback automatically. The core logic from [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) demonstrates how the constants combine with model quantization:

```rust
/// Compute a throughput estimate for an unknown GPU.
fn estimate_tps_with_gpu(
    model: &LlmModel,
    system: &SystemSpecs,
    quant: &str,
) -> f64 {
    // `k` is the constant-K factor derived from known GPUs.
    const K_FACTOR: f64 = 1.0; // (illustrative – actual value is calculated elsewhere)

    // Conservative bandwidth fallback (50 GB/s) + efficiency factor.
    let fallback_efficiency = 0.55;
    let estimated_gpu_bw = K_FACTOR
        * models::quant_bytes_per_param(quant)
        / fallback_efficiency;                // → effective BW in GB/s

    // Compute the time to process active parameters.
    let active_params = model.active_parameters.unwrap_or(model.parameters_raw);
    let active_gb = (active_params as f64 * models::quant_bpp(quant)) / 1e9;
    let gpu_compute_time = active_gb / (estimated_gpu_bw * fallback_efficiency);

    // Convert compute time to tokens-per-second.
    let tps = model.context_len as f64 / gpu_compute_time;
    tps
}

```

Running the CLI on such hardware produces output similar to:

```bash
$ llmfit fit --model mistralai/Mistral-7B-Instruct
Model: Mistral-7B-Instruct
Run mode: Gpu
Estimated throughput: 45 TPS
Notes:
 • Using conservative 50 GB/s memory-bandwidth fallback
 • Applied efficiency factor of 0.55

```

## Summary

- **Detection Failure**: When [`hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/hardware.rs) cannot query **nvidia-smi**, **rocm-smi**, or Vulkan, it signals the fallback path.
- **Fixed Bandwidth**: The system assumes **50 GB/s** (DDR4-3200 dual-channel equivalent) as a conservative baseline in [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs).
- **Efficiency Scaling**: A hard-coded **0.55** efficiency factor accounts for memory controller overhead.
- **Deterministic Formula**: The effective bandwidth combines the K-factor, quantization bytes per parameter, and fallback efficiency.
- **MoE Handling**: Mixture-of-Experts models fall back to RAM-only estimation with performance warnings.
- **Verified Behavior**: Unit tests in [`plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/plan.rs) and [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) ensure the fallback activates correctly for unknown GPU identifiers.

## Frequently Asked Questions

### What specific bandwidth value does llmfit use for unrecognized GPUs?

The system uses a conservative constant of **50 GB/s**, which approximates the bandwidth of DDR4-3200 dual-channel memory. This value is hard-coded in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) at approximately line 1160 and activates whenever hardware detection fails to return a vendor-specific measurement.

### Why does llmfit apply a 0.55 efficiency factor to the fallback bandwidth?

The **0.55** factor accounts for real-world memory subsystem inefficiencies including memory-controller overhead, address-translation costs, and bus contention. Defined at line ~1448 in [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs), this multiplier ensures the 50 GB/s theoretical bandwidth reflects practical achievable throughput on unknown hardware.

### How does llmfit handle MoE models when the GPU is unknown?

For **Mixture-of-Experts** architectures on unrecognized GPUs, the logic in [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) (lines ~884-893) falls back to a RAM-only path if offloading requirements cannot be met. This approach treats the entire model as residing in system memory and explicitly flags the resulting throughput estimate as significantly reduced compared to GPU-accelerated inference.

### Where is the fallback logic located in the llmfit source code?

The primary implementation resides in **[`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs)**, containing the constant-bandwidth definition, efficiency factor, and TPS calculation. The detection trigger logic is in **[`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs)**, while validation tests are located in **[`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs)** (line ~1280) and **[`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs)** (line ~3063) under the test name `test_bandwidth_estimation_unknown_gpu_uses_fallback`.