KV Cache Quantization Options in llmfit: Byte Sizes and Memory Optimization Guide

llmfit supports five distinct KV cache quantization schemes ranging from 2.0 bytes per element (FP16) down to approximately 0.34 bytes per element (TurboQuant), enabling VRAM reduction of up to 83% during inference.

The KV cache quantization options available in the llmfit repository allow developers to trade numerical precision for memory efficiency when running large language models. These compression schemes are defined in the core model logic and directly impact how much GPU memory is required for long-context inference.

Understanding KV Cache Quantization in llmfit

In transformer architectures, the key-value (KV) cache stores intermediate attention states to avoid recomputing them during autoregressive generation. As context length increases, this cache dominates memory consumption. The KvQuant enum in llmfit-core/src/models.rs (lines 601-618) implements the quantization schemes that compress these tensors, with the bytes_per_element method (lines 634-647) providing the precise memory cost for each variant.

Available KV Cache Quantization Options

The KvQuant enum defines five quantization variants, each optimized for different hardware capabilities and memory constraints.

FP16 (Default) — 2.0 Bytes per Element

KvQuant::Fp16 uses standard 16-bit floating-point (or BF16) representation. This is the baseline option that preserves full precision but consumes the most memory at 2.0 bytes per KV element. It is the default when no quantization is specified.

FP8 — 1.0 Byte per Element

KvQuant::Fp8 implements 8-bit floating-point quantization, reducing memory usage by 50% compared to FP16. This format is supported by modern inference engines like vLLM and recent llama.cpp builds (via the --cache-type-k fp8 flag) and costs 1.0 byte per element.

Q8_0 — 1.0 Byte per Element

KvQuant::Q8_0 provides 8-bit integer quantization, offering the same 1.0 byte per element footprint as FP8 but using integer arithmetic. This corresponds to llama.cpp's q8_0 format and vLLM's int8 implementation, suitable for hardware without native FP8 support.

Q4_0 — 0.5 Bytes per Element

KvQuant::Q4_0 compresses the cache to 4-bit integers, achieving 0.5 bytes per element (a 75% reduction from FP16). This aggressive quantization is compatible with llama.cpp's q4_0 and vLLM's int4 backends, though it may impact generation quality on sensitive tasks.

TurboQuant — ~0.34 Bytes per Element

KvQuant::TurboQuant is an experimental scheme achieving approximately 0.34 bytes per element (roughly 2.7 bits). It uses a hybrid 3-bit key plus 2-bit value approach via the turboquant research project. Note that this is CUDA-only and works only on full-attention layers.

Calculating Memory Usage with bytes_per_element

The bytes_per_element method in llmfit-core/src/models.rs returns the exact memory cost for each variant:

pub fn bytes_per_element(&self) -> f64 {
    match self {
        KvQuant::Fp16 => 2.0,
        KvQuant::Fp8 => 1.0,
        KvQuant::Q8_0 => 1.0,
        KvQuant::Q4_0 => 0.5,
        // TurboQuant uses ~2.7 bits per element on the compressible slice.
        KvQuant::TurboQuant => 0.34,
    }
}

These values feed into the memory-estimation logic (model.kv_cache_gb) to calculate total VRAM requirements based on context length and quantization choice.

Parsing and Selecting Quantization Schemes

You can instantiate quantization options from CLI arguments or configuration strings using the parse method:

use llmfit_core::models::KvQuant;

// Parse user input (e.g., from --kv-quant flag)
let kv = KvQuant::parse("q4_0").expect("unsupported quantization scheme");

// Retrieve precise byte cost
println!(
    "KV cache uses {:.2} bytes per element for {}",
    kv.bytes_per_element(),
    kv.label()
);

// Enumerate all supported options
for q in KvQuant::all() {
    println!("{} → {:.2} bytes/elem", q.label(), q.bytes_per_element());
}

This outputs:

KV cache uses 0.50 bytes per element for q4_0
fp16 → 2.00 bytes/elem
fp8 → 1.00 bytes/elem
q8_0 → 1.00 bytes/elem
q4_0 → 0.50 bytes/elem
tq → 0.34 bytes/elem

Integration with Memory Estimation

The quantization selection directly impacts memory planning in llmfit-core/src/plan.rs. When generating execution plans, the system uses the selected KvQuant variant to compute kv_cache_gb, determining whether a model fitting job will run within available VRAM constraints. The TUI layer in llmfit-tui/src/main.rs exposes this via the --kv-quant command-line flag, converting user input into the appropriate enum variant before invoking the planner.

Summary

  • Five quantization levels: FP16 (2.0B), FP8 (1.0B), Q8_0 (1.0B), Q4_0 (0.5B), and TurboQuant (~0.34B)
  • Memory savings: Range from 0% compression (FP16) to ~83% compression (TurboQuant) compared to baseline
  • Implementation location: KvQuant enum and bytes_per_element method defined in llmfit-core/src/models.rs
  • Usage: Parse from strings via KvQuant::parse(), retrieve bytes with bytes_per_element(), integrate into memory estimation via plan.rs
  • Hardware constraints: TurboQuant requires CUDA and full-attention layers; other formats work broadly with vLLM and llama.cpp backends

Frequently Asked Questions

What is the default KV cache quantization in llmfit?

FP16 is the default when no quantization is specified. This corresponds to KvQuant::Fp16 at 2.0 bytes per element, providing maximum accuracy at the cost of higher VRAM usage. You can override this by passing a specific variant to the --kv-quant flag.

How much memory does TurboQuant save compared to FP16?

TurboQuant reduces KV cache memory by approximately 83% compared to FP16. While FP16 consumes 2.0 bytes per element, TurboQuant uses roughly 0.34 bytes per element (about 2.7 bits). For a model with a 32k context window, this difference can save several gigabytes of VRAM.

Can I use TurboQuant on any hardware?

No, TurboQuant requires CUDA-compatible NVIDIA GPUs and only works on full-attention layers. Unlike FP8, Q8_0, or Q4_0—which are broadly compatible with vLLM and llama.cpp—TurboQuant is an experimental feature tied to the specific CUDA kernels in the turboquant research project.

How does llmfit calculate total KV cache memory requirements?

llmfit multiplies three values: the number of layers, the context length, and the bytes_per_element() value of the selected KvQuant variant. This calculation occurs in llmfit-core/src/plan.rs within the memory estimation logic, allowing the system to predict whether a given model and context configuration will fit within available GPU memory before execution begins.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →