# KV Cache Quantization Options in llmfit: Byte Sizes and Memory Optimization Guide

> Explore KV cache quantization in llmfit. Discover options from FP16 (2.0 bytes) to TurboQuant (~0.34 bytes) and achieve up to 83% VRAM reduction for efficient LLM inference.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: deep-dive
- Published: 2026-08-22

---

**llmfit supports five distinct KV cache quantization schemes ranging from 2.0 bytes per element (FP16) down to approximately 0.34 bytes per element (TurboQuant), enabling VRAM reduction of up to 83% during inference.**

The **KV cache quantization options** available in the [llmfit](https://github.com/AlexsJones/llmfit) repository allow developers to trade numerical precision for memory efficiency when running large language models. These compression schemes are defined in the core model logic and directly impact how much GPU memory is required for long-context inference.

## Understanding KV Cache Quantization in llmfit

In transformer architectures, the key-value (KV) cache stores intermediate attention states to avoid recomputing them during autoregressive generation. As context length increases, this cache dominates memory consumption. The `KvQuant` enum in [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs) (lines 601-618) implements the quantization schemes that compress these tensors, with the `bytes_per_element` method (lines 634-647) providing the precise memory cost for each variant.

## Available KV Cache Quantization Options

The `KvQuant` enum defines five quantization variants, each optimized for different hardware capabilities and memory constraints.

### FP16 (Default) — 2.0 Bytes per Element

**`KvQuant::Fp16`** uses standard 16-bit floating-point (or BF16) representation. This is the baseline option that preserves full precision but consumes the most memory at **2.0 bytes per KV element**. It is the default when no quantization is specified.

### FP8 — 1.0 Byte per Element

**`KvQuant::Fp8`** implements 8-bit floating-point quantization, reducing memory usage by 50% compared to FP16. This format is supported by modern inference engines like vLLM and recent llama.cpp builds (via the `--cache-type-k fp8` flag) and costs **1.0 byte per element**.

### Q8_0 — 1.0 Byte per Element

**`KvQuant::Q8_0`** provides 8-bit integer quantization, offering the same **1.0 byte per element** footprint as FP8 but using integer arithmetic. This corresponds to llama.cpp's `q8_0` format and vLLM's int8 implementation, suitable for hardware without native FP8 support.

### Q4_0 — 0.5 Bytes per Element

**`KvQuant::Q4_0`** compresses the cache to 4-bit integers, achieving **0.5 bytes per element** (a 75% reduction from FP16). This aggressive quantization is compatible with llama.cpp's `q4_0` and vLLM's int4 backends, though it may impact generation quality on sensitive tasks.

### TurboQuant — ~0.34 Bytes per Element

**`KvQuant::TurboQuant`** is an experimental scheme achieving approximately **0.34 bytes per element** (roughly 2.7 bits). It uses a hybrid 3-bit key plus 2-bit value approach via the [turboquant](https://github.com/0xSero/turboquant) research project. Note that this is CUDA-only and works only on full-attention layers.

## Calculating Memory Usage with bytes_per_element

The `bytes_per_element` method in [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs) returns the exact memory cost for each variant:

```rust
pub fn bytes_per_element(&self) -> f64 {
    match self {
        KvQuant::Fp16 => 2.0,
        KvQuant::Fp8 => 1.0,
        KvQuant::Q8_0 => 1.0,
        KvQuant::Q4_0 => 0.5,
        // TurboQuant uses ~2.7 bits per element on the compressible slice.
        KvQuant::TurboQuant => 0.34,
    }
}

```

These values feed into the memory-estimation logic (`model.kv_cache_gb`) to calculate total VRAM requirements based on context length and quantization choice.

## Parsing and Selecting Quantization Schemes

You can instantiate quantization options from CLI arguments or configuration strings using the `parse` method:

```rust
use llmfit_core::models::KvQuant;

// Parse user input (e.g., from --kv-quant flag)
let kv = KvQuant::parse("q4_0").expect("unsupported quantization scheme");

// Retrieve precise byte cost
println!(
    "KV cache uses {:.2} bytes per element for {}",
    kv.bytes_per_element(),
    kv.label()
);

// Enumerate all supported options
for q in KvQuant::all() {
    println!("{} → {:.2} bytes/elem", q.label(), q.bytes_per_element());
}

```

This outputs:

```text
KV cache uses 0.50 bytes per element for q4_0
fp16 → 2.00 bytes/elem
fp8 → 1.00 bytes/elem
q8_0 → 1.00 bytes/elem
q4_0 → 0.50 bytes/elem
tq → 0.34 bytes/elem

```

## Integration with Memory Estimation

The quantization selection directly impacts memory planning in [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs). When generating execution plans, the system uses the selected `KvQuant` variant to compute `kv_cache_gb`, determining whether a model fitting job will run within available VRAM constraints. The TUI layer in [`llmfit-tui/src/main.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/main.rs) exposes this via the `--kv-quant` command-line flag, converting user input into the appropriate enum variant before invoking the planner.

## Summary

- **Five quantization levels**: FP16 (2.0B), FP8 (1.0B), Q8_0 (1.0B), Q4_0 (0.5B), and TurboQuant (~0.34B)
- **Memory savings**: Range from 0% compression (FP16) to ~83% compression (TurboQuant) compared to baseline
- **Implementation location**: `KvQuant` enum and `bytes_per_element` method defined in [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs)
- **Usage**: Parse from strings via `KvQuant::parse()`, retrieve bytes with `bytes_per_element()`, integrate into memory estimation via [`plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/plan.rs)
- **Hardware constraints**: TurboQuant requires CUDA and full-attention layers; other formats work broadly with vLLM and llama.cpp backends

## Frequently Asked Questions

### What is the default KV cache quantization in llmfit?

**FP16 is the default** when no quantization is specified. This corresponds to `KvQuant::Fp16` at 2.0 bytes per element, providing maximum accuracy at the cost of higher VRAM usage. You can override this by passing a specific variant to the `--kv-quant` flag.

### How much memory does TurboQuant save compared to FP16?

**TurboQuant reduces KV cache memory by approximately 83%** compared to FP16. While FP16 consumes 2.0 bytes per element, TurboQuant uses roughly 0.34 bytes per element (about 2.7 bits). For a model with a 32k context window, this difference can save several gigabytes of VRAM.

### Can I use TurboQuant on any hardware?

**No, TurboQuant requires CUDA-compatible NVIDIA GPUs** and only works on full-attention layers. Unlike FP8, Q8_0, or Q4_0—which are broadly compatible with vLLM and llama.cpp—TurboQuant is an experimental feature tied to the specific CUDA kernels in the turboquant research project.

### How does llmfit calculate total KV cache memory requirements?

llmfit multiplies three values: the **number of layers**, the **context length**, and the `bytes_per_element()` value of the selected `KvQuant` variant. This calculation occurs in [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs) within the memory estimation logic, allowing the system to predict whether a given model and context configuration will fit within available GPU memory before execution begins.