How llmfit Uses TurboQuant for KV Cache Optimization: Implementation Guide
TurboQuant reduces the KV cache memory footprint to 0.34 bytes per element using 3-bit quantization for keys and 2-bit quantization for values, and llmfit implements this compression through backend-gated memory estimation that selectively applies mixed-precision quantization only to full-attention layers on CUDA-enabled vLLM systems.
The AlexsJones/llmfit repository provides memory planning tools for large language model deployment, and its TurboQuant KV cache optimization feature allows models to run in constrained GPU memory by compressing attention caches beyond standard fp16 precision. This implementation selectively targets architectural components while maintaining computational precision for linear and state-space layers.
What is TurboQuant KV Cache Compression?
TurboQuant is a KV cache compression scheme that achieves aggressive memory reduction by storing keys at 3-bit precision and values at 2-bit precision. This results in approximately 0.34 bytes per element, compared to 2 bytes per element in standard fp16 storage. The compression applies primarily to transformer attention mechanisms while preserving higher precision for other layer types.
Architectural Integration in llmfit
llmfit implements TurboQuant across three distinct architectural layers, combining quantization definitions, memory calculations, and backend validation.
Quantization Schema Definition
In llmfit-core/src/models.rs (lines 15-20), the KvQuant enum defines supported quantization formats, including the TurboQuant variant labeled as "tq". The implementation assigns 0.34 bytes per element to this compression scheme, distinguishing it from standard fp16 or int8 alternatives.
Mixed-Precision Memory Estimation
The Model::kv_cache_gb() method in models.rs (lines 94-101) computes KV cache sizes using a mixed-precision approach. For TurboQuant, the function compresses only the full-attention slice—layers participating in full self-attention—while maintaining fp16 precision for linear and state-space layers. This selective compression prevents accuracy degradation in non-attention components.
CUDA Backend Validation
TurboQuant requires specific hardware acceleration available only through vLLM on CUDA. In llmfit-core/src/plan.rs (lines 658-660), the Plan::new() constructor validates the system configuration, checking system.backend == GpuBackend::Cuda and returning an error for unsupported configurations such as CPU or ROCm execution.
Fit Detection and Flagging
The Fit struct in llmfit-core/src/fit.rs (lines 252-260) contains the fits_with_turboquant: bool flag. When a model exceeds memory capacity under fp16 KV caching but fits within constraints after TurboQuant compression, this flag enables the "TurboQuant+" fit status surfaced in llmfit-tui/src/tui_ui.rs (line 2294).
Practical Usage Examples
Enabling TurboQuant via CLI
To request TurboQuant compression from the command line, use the --kv-quant flag with the tq identifier:
llmfit fit --kv-quant tq llama-2-7b
When the system supports CUDA and the model fits with compression enabled, the output displays the TurboQuant status:
TurboQuant+: Would fit with 9.8x KV compression
Model KV cache (GB) Memory (GB) Fit
----------------------------------------------------
llama-2-7b 9.4 (tq) 12.1 TurboQuant+
If the backend lacks CUDA support, llmfit-tui/src/main.rs (lines 2030-2032) triggers a warning: "TurboQuant is experimental, not in upstream vLLM yet."
Programmatic Memory Calculation
Calculate KV cache sizes programmatically using the core library:
use llmfit_core::models::{Model, KvQuant};
fn main() -> anyhow::Result<()> {
// Load model metadata without initializing weights
let model = Model::load("meta-llama/Llama-2-7b-chat-hf")?;
// Calculate TurboQuant size at 8k context window
let kv_gb = model.kv_cache_gb(8192, KvQuant::TurboQuant);
println!("TurboQuant KV cache size: {:.2} GB", kv_gb);
Ok(())
}
This returns the compressed memory footprint, accounting for the mixed-precision layer handling.
Checking Model Compatibility
Verify whether a model fits within hardware constraints using the planning interface:
use llmfit_core::{plan::Plan, hardware::SystemSpecs, models::KvQuant};
fn main() -> anyhow::Result<()> {
// Detect hardware specifications
let specs = SystemSpecs::detect()?;
// Load model definition
let model = Model::load("meta-llama/Llama-2-7b-chat-hf")?;
// Create execution plan with TurboQuant enabled
let plan = Plan::new(&model, &specs, KvQuant::TurboQuant)?;
if plan.fits_with_turboquant {
println!("Model fits with TurboQuant compression");
} else {
println!("Model requires additional optimization");
}
Ok(())
}
Summary
- TurboQuant achieves 0.34 bytes per element through 3-bit key and 2-bit value quantization.
- llmfit applies compression selectively to full-attention layers only, preserving fp16 precision for linear and state-space components as implemented in
models.rs. - The implementation validates CUDA backend availability in
plan.rsbefore permitting TurboQuant usage. - The
fits_with_turboquantflag infit.rsenables UI differentiation between standard fits and compressed fits. - The TUI surface in
tui_ui.rsdisplays compression ratios when TurboQuant enables model deployment.
Frequently Asked Questions
What compression ratio does TurboQuant achieve?
TurboQuant reduces KV cache memory requirements by approximately 9.8x compared to fp16 storage, compressing keys to 3 bits and values to 2 bits for an effective rate of 0.34 bytes per element.
Why does TurboQuant only work with CUDA backends?
The quantization kernels and vLLM integration for TurboQuant are currently implemented exclusively for NVIDIA CUDA architectures. The Plan::new() constructor in llmfit-core/src/plan.rs explicitly checks for GpuBackend::Cuda and rejects other backends because the experimental compression operators are not yet available in upstream vLLM for CPU or ROCm execution.
Which model layers are compressed by TurboQuant?
Only layers participating in full self-attention receive TurboQuant compression. The kv_cache_gb() implementation in models.rs maintains fp16 precision for linear projections and state-space model layers to prevent accuracy degradation in non-attention computations.
How do I know if a model will fit with TurboQuant enabled?
The Fit struct's fits_with_turboquant boolean indicates whether a model that exceeds memory limits under fp16 KV caching will fit within available GPU memory when TurboQuant compression is applied. The CLI and TUI display "TurboQuant+" when this condition is met, as implemented in llmfit-core/src/fit.rs and surfaced through llmfit-tui/src/tui_ui.rs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →