# How llmfit Uses TurboQuant for KV Cache Optimization: Implementation Guide

> Discover how llmfit leverages TurboQuant for efficient KV cache optimization. Learn about 3-bit and 2-bit quantization for memory reduction on vLLM systems.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: implementation-guide
- Published: 2026-08-22

---

**TurboQuant reduces the KV cache memory footprint to 0.34 bytes per element using 3-bit quantization for keys and 2-bit quantization for values, and llmfit implements this compression through backend-gated memory estimation that selectively applies mixed-precision quantization only to full-attention layers on CUDA-enabled vLLM systems.**

The AlexsJones/llmfit repository provides memory planning tools for large language model deployment, and its TurboQuant KV cache optimization feature allows models to run in constrained GPU memory by compressing attention caches beyond standard fp16 precision. This implementation selectively targets architectural components while maintaining computational precision for linear and state-space layers.

## What is TurboQuant KV Cache Compression?

TurboQuant is a KV cache compression scheme that achieves aggressive memory reduction by storing **keys at 3-bit precision** and **values at 2-bit precision**. This results in approximately **0.34 bytes per element**, compared to 2 bytes per element in standard fp16 storage. The compression applies primarily to transformer attention mechanisms while preserving higher precision for other layer types.

## Architectural Integration in llmfit

llmfit implements TurboQuant across three distinct architectural layers, combining quantization definitions, memory calculations, and backend validation.

### Quantization Schema Definition

In [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs) (lines 15-20), the `KvQuant` enum defines supported quantization formats, including the `TurboQuant` variant labeled as `"tq"`. The implementation assigns **0.34 bytes per element** to this compression scheme, distinguishing it from standard fp16 or int8 alternatives.

### Mixed-Precision Memory Estimation

The `Model::kv_cache_gb()` method in [`models.rs`](https://github.com/AlexsJones/llmfit/blob/main/models.rs) (lines 94-101) computes KV cache sizes using a mixed-precision approach. For TurboQuant, the function compresses **only the full-attention slice**—layers participating in full self-attention—while maintaining fp16 precision for linear and state-space layers. This selective compression prevents accuracy degradation in non-attention components.

### CUDA Backend Validation

TurboQuant requires specific hardware acceleration available only through vLLM on CUDA. In [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs) (lines 658-660), the `Plan::new()` constructor validates the system configuration, checking `system.backend == GpuBackend::Cuda` and returning an error for unsupported configurations such as CPU or ROCm execution.

### Fit Detection and Flagging

The `Fit` struct in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) (lines 252-260) contains the `fits_with_turboquant: bool` flag. When a model exceeds memory capacity under fp16 KV caching but fits within constraints after TurboQuant compression, this flag enables the "TurboQuant+" fit status surfaced in [`llmfit-tui/src/tui_ui.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_ui.rs) (line 2294).

## Practical Usage Examples

### Enabling TurboQuant via CLI

To request TurboQuant compression from the command line, use the `--kv-quant` flag with the `tq` identifier:

```bash
llmfit fit --kv-quant tq llama-2-7b

```

When the system supports CUDA and the model fits with compression enabled, the output displays the TurboQuant status:

```

TurboQuant+: Would fit with 9.8x KV compression
Model               KV cache (GB)   Memory (GB)   Fit
----------------------------------------------------
llama-2-7b          9.4 (tq)        12.1          TurboQuant+

```

If the backend lacks CUDA support, [`llmfit-tui/src/main.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/main.rs) (lines 2030-2032) triggers a warning: "TurboQuant is experimental, not in upstream vLLM yet."

### Programmatic Memory Calculation

Calculate KV cache sizes programmatically using the core library:

```rust
use llmfit_core::models::{Model, KvQuant};

fn main() -> anyhow::Result<()> {
    // Load model metadata without initializing weights
    let model = Model::load("meta-llama/Llama-2-7b-chat-hf")?;
    
    // Calculate TurboQuant size at 8k context window
    let kv_gb = model.kv_cache_gb(8192, KvQuant::TurboQuant);
    println!("TurboQuant KV cache size: {:.2} GB", kv_gb);
    
    Ok(())
}

```

This returns the compressed memory footprint, accounting for the mixed-precision layer handling.

### Checking Model Compatibility

Verify whether a model fits within hardware constraints using the planning interface:

```rust
use llmfit_core::{plan::Plan, hardware::SystemSpecs, models::KvQuant};

fn main() -> anyhow::Result<()> {
    // Detect hardware specifications
    let specs = SystemSpecs::detect()?;
    
    // Load model definition
    let model = Model::load("meta-llama/Llama-2-7b-chat-hf")?;
    
    // Create execution plan with TurboQuant enabled
    let plan = Plan::new(&model, &specs, KvQuant::TurboQuant)?;
    
    if plan.fits_with_turboquant {
        println!("Model fits with TurboQuant compression");
    } else {
        println!("Model requires additional optimization");
    }
    
    Ok(())
}

```

## Summary

- TurboQuant achieves **0.34 bytes per element** through 3-bit key and 2-bit value quantization.
- llmfit applies compression selectively to **full-attention layers only**, preserving fp16 precision for linear and state-space components as implemented in [`models.rs`](https://github.com/AlexsJones/llmfit/blob/main/models.rs).
- The implementation validates **CUDA backend availability** in [`plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/plan.rs) before permitting TurboQuant usage.
- The **`fits_with_turboquant`** flag in [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) enables UI differentiation between standard fits and compressed fits.
- The TUI surface in [`tui_ui.rs`](https://github.com/AlexsJones/llmfit/blob/main/tui_ui.rs) displays compression ratios when TurboQuant enables model deployment.

## Frequently Asked Questions

### What compression ratio does TurboQuant achieve?

TurboQuant reduces KV cache memory requirements by approximately **9.8x** compared to fp16 storage, compressing keys to 3 bits and values to 2 bits for an effective rate of 0.34 bytes per element.

### Why does TurboQuant only work with CUDA backends?

The quantization kernels and vLLM integration for TurboQuant are currently implemented exclusively for NVIDIA CUDA architectures. The `Plan::new()` constructor in [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs) explicitly checks for `GpuBackend::Cuda` and rejects other backends because the experimental compression operators are not yet available in upstream vLLM for CPU or ROCm execution.

### Which model layers are compressed by TurboQuant?

Only layers participating in **full self-attention** receive TurboQuant compression. The `kv_cache_gb()` implementation in [`models.rs`](https://github.com/AlexsJones/llmfit/blob/main/models.rs) maintains fp16 precision for linear projections and state-space model layers to prevent accuracy degradation in non-attention computations.

### How do I know if a model will fit with TurboQuant enabled?

The `Fit` struct's `fits_with_turboquant` boolean indicates whether a model that exceeds memory limits under fp16 KV caching will fit within available GPU memory when TurboQuant compression is applied. The CLI and TUI display "TurboQuant+" when this condition is met, as implemented in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) and surfaced through [`llmfit-tui/src/tui_ui.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_ui.rs).