How TurboQuant KV Compression Enables TooTight Models to Fit on CUDA Systems in LLMFIT

TurboQuant KV compression reduces the key-value cache footprint from 2 bytes per element (fp16) to 0.34 bytes, enabling models marked as TooTight to fit within available CUDA VRAM when using the vLLM backend.

LLMFIT is a Rust-based hardware compatibility evaluator that determines whether large language models can run on specific GPU configurations. When a model's estimated KV-cache size exceeds available GPU memory, the system classifies it as TooTight and excludes it from runnable results. The TurboQuant compression option, implemented in the AlexsJones/llmfit repository, specifically targets this bottleneck by compressing the attention cache by approximately 9.8×, but only on CUDA-enabled systems.

What Makes a Model "TooTight" in LLMFIT

LLMFIT calculates memory requirements by estimating the size of the KV (key-value) cache, which stores intermediate attention states during inference. When this calculation exceeds the available VRAM on the target GPU, the model receives a TooTight classification and is filtered from the default result set.

Without compression, a standard fp16 KV cache consumes 2 bytes per element. For large context windows and batch sizes, this accumulates rapidly, causing high-parameter models to fail the fit analysis even on modern hardware.

How TurboQuant KV Compression Reduces VRAM Usage

The Compression Mechanism

In llmfit-core/src/models.rs, the KvQuant enum defines TurboQuant as a specialized quantization mode. The implementation specifies a custom bytes_per_element() value of 0.34 bytes (lines 615–645), compared to the standard 2.0 bytes for fp16 caches.

This represents approximately 9.8× compression, allowing models with massive KV-cache requirements to fit into significantly smaller memory footprints. When enabled, LLMFIT recalculates the total memory budget using this reduced per-element cost, potentially reclassifying a TooTight model as runnable.

CUDA and vLLM Backend Requirements

TurboQuant is not universally available. According to the validation logic in llmfit-core/src/plan.rs (lines 652–659), the system explicitly checks that TurboQuant is only applied when:

  • The backend is vLLM
  • The hardware is CUDA-capable

If these conditions are not met, the planner returns an error, preventing incompatible configurations from proceeding.

Source Code Implementation

Defining the Compression Scheme in models.rs

The quantization options are centralized in llmfit-core/src/models.rs. The TurboQuant variant is defined alongside other KvQuant options, with its distinctive memory characteristics hardcoded:

// Located in llmfit-core/src/models.rs, lines 615-645
pub enum KvQuant {
    Fp16,
    TurboQuant,  // 0.34 bytes per element
}

impl KvQuant {
    pub fn bytes_per_element(&self) -> f32 {
        match self {
            KvQuant::Fp16 => 2.0,
            KvQuant::TurboQuant => 0.34,  // ~9.8x compression
        }
    }
}

Runtime Validation in plan.rs

Before execution, llmfit-core/src/plan.rs validates hardware compatibility. Lines 652–659 enforce the CUDA constraint:

// Located in llmfit-core/src/plan.rs, lines 652-659
if kv_quant == KvQuant::TurboQuant && !backend.is_cuda() {
    return Err(PlanError::TurboQuantRequiresCuda);
}

Fit Analysis Logic in fit.rs

The core recalculation happens in llmfit-core/src/fit.rs (lines 652–660). When a model exceeds memory limits under standard fp16 quantization but fits within the compressed budget, the system sets the fits_with_turboquant flag:

// Located in llmfit-core/src/fit.rs, lines 652-660
let fits_with_turboquant = if model_kv_size_fp16 > vram_available {
    let kv_size_turbo = model_kv_size_fp16 * 0.17; // 0.34/2.0
    kv_size_turbo < vram_available
} else {
    false
};

This flag indicates that while the model is TooTight under normal conditions, it becomes runnable with TurboQuant enabled.

TUI Filter Integration in tui_app.rs

The terminal interface exposes this functionality through the TurboQuantFit filter. In llmfit-tui/src/tui_app.rs (lines 641–678), the TUI maps the fits_with_turboquant flag to a user-visible category, displaying models that only fit when compression is applied.

Additionally, llmfit-tui/src/tui_ui.rs (line 2294) renders the specific UI label: "TurboQuant+: Would fit with 9.8x KV compression".

Practical Usage Examples

Programmatic Configuration

// Request a model plan with TurboQuant KV compression
use llmfit_core::models::KvQuant;
use llmfit_core::hardware::GpuBackend;

let plan = llmfit_core::plan::Plan::new()
    .with_kv_quant(KvQuant::TurboQuant)
    .with_backend(GpuBackend::Cuda);

match plan.build() {
    Ok(p) if p.fits_with_turboquant => {
        println!("Model fits thanks to TurboQuant compression!");
    }
    Ok(_) => println!("Model fits without compression."),
    Err(e) => eprintln!("Planning error: {}", e),
}

Command-Line Interface


# Request TurboQuant mode using the 'tq' shorthand

llmfit fit --model "Llama-2-70B" --kv-quant tq --backend cuda

# Output categorizes the model under "TQ+ Fit" if it only fits with compression

Summary

  • TurboQuant compresses the KV cache from 2 bytes/element (fp16) to 0.34 bytes/element, achieving 9.8× memory reduction.
  • Models exceeding VRAM limits are marked TooTight, but can be reclassified as runnable when fits_with_turboquant is calculated in fit.rs.
  • The feature is restricted to CUDA systems using the vLLM backend, enforced by validation logic in plan.rs.
  • The TUI exposes this through the TurboQuantFit filter, allowing users to identify models that require compression to run.

Frequently Asked Questions

What is the exact compression ratio of TurboQuant?

TurboQuant achieves approximately 9.8× compression by reducing the KV cache from 2 bytes per element (fp16) to 0.34 bytes per element, as defined in llmfit-core/src/models.rs (lines 615–645).

Can I use TurboQuant on AMD or CPU-only systems?

No. According to llmfit-core/src/plan.rs (lines 652–659), TurboQuant is explicitly gated to CUDA-only configurations. Attempting to use it with AMD GPUs or CPU backends returns a validation error.

How do I identify models that only fit because of TurboQuant?

The fits_with_turboquant flag in llmfit-core/src/fit.rs (lines 652–660) marks these models. In the TUI, they appear under the TurboQuantFit filter with the label "TurboQuant+: Would fit with 9.8x KV compression" (defined in llmfit-tui/src/tui_ui.rs, line 2294).

Does TurboQuant affect inference accuracy?

The source code analysis indicates TurboQuant is a memory optimization technique, but specific accuracy implications depend on the underlying vLLM implementation rather than the LLMFIT orchestration layer. The repository treats it as a transparent memory reduction layer applied to the KV cache.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →