# How TurboQuant KV Compression Enables TooTight Models to Fit on CUDA Systems in LLMFIT

> Discover how TurboQuant KV compression slashes key-value cache size, allowing TooTight LLM models to fit and run on CUDA systems with LLMFIT and vLLM.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: deep-dive
- Published: 2026-08-21

---

**TurboQuant KV compression reduces the key-value cache footprint from 2 bytes per element (fp16) to 0.34 bytes, enabling models marked as TooTight to fit within available CUDA VRAM when using the vLLM backend.**

LLMFIT is a Rust-based hardware compatibility evaluator that determines whether large language models can run on specific GPU configurations. When a model's estimated KV-cache size exceeds available GPU memory, the system classifies it as **TooTight** and excludes it from runnable results. The TurboQuant compression option, implemented in the `AlexsJones/llmfit` repository, specifically targets this bottleneck by compressing the attention cache by approximately 9.8×, but only on CUDA-enabled systems.

## What Makes a Model "TooTight" in LLMFIT

LLMFIT calculates memory requirements by estimating the size of the KV (key-value) cache, which stores intermediate attention states during inference. When this calculation exceeds the available VRAM on the target GPU, the model receives a **TooTight** classification and is filtered from the default result set.

Without compression, a standard fp16 KV cache consumes **2 bytes per element**. For large context windows and batch sizes, this accumulates rapidly, causing high-parameter models to fail the fit analysis even on modern hardware.

## How TurboQuant KV Compression Reduces VRAM Usage

### The Compression Mechanism

In [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs), the `KvQuant` enum defines `TurboQuant` as a specialized quantization mode. The implementation specifies a custom `bytes_per_element()` value of **0.34 bytes** (lines 615–645), compared to the standard 2.0 bytes for fp16 caches.

This represents approximately **9.8× compression**, allowing models with massive KV-cache requirements to fit into significantly smaller memory footprints. When enabled, LLMFIT recalculates the total memory budget using this reduced per-element cost, potentially reclassifying a **TooTight** model as runnable.

## CUDA and vLLM Backend Requirements

TurboQuant is not universally available. According to the validation logic in [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs) (lines 652–659), the system explicitly checks that TurboQuant is only applied when:

- The backend is **vLLM**
- The hardware is **CUDA**-capable

If these conditions are not met, the planner returns an error, preventing incompatible configurations from proceeding.

## Source Code Implementation

### Defining the Compression Scheme in models.rs

The quantization options are centralized in [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs). The `TurboQuant` variant is defined alongside other `KvQuant` options, with its distinctive memory characteristics hardcoded:

```rust
// Located in llmfit-core/src/models.rs, lines 615-645
pub enum KvQuant {
    Fp16,
    TurboQuant,  // 0.34 bytes per element
}

impl KvQuant {
    pub fn bytes_per_element(&self) -> f32 {
        match self {
            KvQuant::Fp16 => 2.0,
            KvQuant::TurboQuant => 0.34,  // ~9.8x compression
        }
    }
}

```

### Runtime Validation in plan.rs

Before execution, [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs) validates hardware compatibility. Lines 652–659 enforce the CUDA constraint:

```rust
// Located in llmfit-core/src/plan.rs, lines 652-659
if kv_quant == KvQuant::TurboQuant && !backend.is_cuda() {
    return Err(PlanError::TurboQuantRequiresCuda);
}

```

### Fit Analysis Logic in fit.rs

The core recalculation happens in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) (lines 652–660). When a model exceeds memory limits under standard fp16 quantization but fits within the compressed budget, the system sets the `fits_with_turboquant` flag:

```rust
// Located in llmfit-core/src/fit.rs, lines 652-660
let fits_with_turboquant = if model_kv_size_fp16 > vram_available {
    let kv_size_turbo = model_kv_size_fp16 * 0.17; // 0.34/2.0
    kv_size_turbo < vram_available
} else {
    false
};

```

This flag indicates that while the model is **TooTight** under normal conditions, it becomes runnable with TurboQuant enabled.

### TUI Filter Integration in tui_app.rs

The terminal interface exposes this functionality through the `TurboQuantFit` filter. In [`llmfit-tui/src/tui_app.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_app.rs) (lines 641–678), the TUI maps the `fits_with_turboquant` flag to a user-visible category, displaying models that only fit when compression is applied.

Additionally, [`llmfit-tui/src/tui_ui.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_ui.rs) (line 2294) renders the specific UI label: *"TurboQuant+: Would fit with 9.8x KV compression"*.

## Practical Usage Examples

### Programmatic Configuration

```rust
// Request a model plan with TurboQuant KV compression
use llmfit_core::models::KvQuant;
use llmfit_core::hardware::GpuBackend;

let plan = llmfit_core::plan::Plan::new()
    .with_kv_quant(KvQuant::TurboQuant)
    .with_backend(GpuBackend::Cuda);

match plan.build() {
    Ok(p) if p.fits_with_turboquant => {
        println!("Model fits thanks to TurboQuant compression!");
    }
    Ok(_) => println!("Model fits without compression."),
    Err(e) => eprintln!("Planning error: {}", e),
}

```

### Command-Line Interface

```bash

# Request TurboQuant mode using the 'tq' shorthand

llmfit fit --model "Llama-2-70B" --kv-quant tq --backend cuda

# Output categorizes the model under "TQ+ Fit" if it only fits with compression

```

## Summary

- **TurboQuant** compresses the KV cache from 2 bytes/element (fp16) to **0.34 bytes/element**, achieving **9.8× memory reduction**.
- Models exceeding VRAM limits are marked **TooTight**, but can be reclassified as runnable when `fits_with_turboquant` is calculated in [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs).
- The feature is **restricted to CUDA systems** using the vLLM backend, enforced by validation logic in [`plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/plan.rs).
- The TUI exposes this through the **TurboQuantFit** filter, allowing users to identify models that require compression to run.

## Frequently Asked Questions

### What is the exact compression ratio of TurboQuant?

TurboQuant achieves approximately **9.8× compression** by reducing the KV cache from 2 bytes per element (fp16) to 0.34 bytes per element, as defined in [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs) (lines 615–645).

### Can I use TurboQuant on AMD or CPU-only systems?

No. According to [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs) (lines 652–659), TurboQuant is explicitly gated to **CUDA-only** configurations. Attempting to use it with AMD GPUs or CPU backends returns a validation error.

### How do I identify models that only fit because of TurboQuant?

The `fits_with_turboquant` flag in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) (lines 652–660) marks these models. In the TUI, they appear under the **TurboQuantFit** filter with the label *"TurboQuant+: Would fit with 9.8x KV compression"* (defined in [`llmfit-tui/src/tui_ui.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_ui.rs), line 2294).

### Does TurboQuant affect inference accuracy?

The source code analysis indicates TurboQuant is a memory optimization technique, but specific accuracy implications depend on the underlying vLLM implementation rather than the LLMFIT orchestration layer. The repository treats it as a transparent memory reduction layer applied to the KV cache.