TurboQuant KV Cache Quantization in MLX-VLM: Low-Bit Compression for Efficient Generation
TurboQuant KV Cache Quantization is a low-bit integer compression scheme that quantizes transformer key-value tensors on-the-fly to configurable bit widths, reducing memory usage during autoregressive generation while maintaining attention score fidelity.
MLX-VLM, the Apple MLX-based vision-language model framework, implements TurboQuant KV Cache Quantization to address the memory bottleneck of growing KV caches during long-context inference. Unlike standard full-precision caches that store float16 tensors, TurboQuant compresses keys and values to fractional bit widths (e.g., 3.5 bits) using statistical quantization and GPU-accelerated codecs.
Architecture of TurboQuant KV Cache Quantization
The implementation centers on the TurboQuantKVCache class in mlx_vlm/turboquant.py, which extends the base cache interface to provide quantized storage and retrieval.
The TurboQuantKVCache Class
At lines 4790–4812 of mlx_vlm/turboquant.py, the TurboQuantKVCache class inherits from _BaseCache and manages quantized KV states. It stores tensors in packed uint32 buffers alongside per-token norms, retaining the original tensor shape for efficient slicing while compressing the underlying data.
from mlx_vlm.turboquant import TurboQuantKVCache
# Initialize cache with 3.5-bit quantization
cache = TurboQuantKVCache(bits=3.5)
Quantization Codecs and Bit Splitting
The _ensure_codecs method (around line 19 in turboquant.py) constructs MSE-optimized integer codecs for keys and values. For fractional bit widths, the implementation splits bits between key and value tensors—assigning floor(bits) to keys and ceil(bits) to values—to achieve the target compression rate.
Fused vs. Separate Quantization Paths
TurboQuant optimizes throughput through _try_fused_kv_quantize (line 40), a Metal kernel that packs key and value tensors together when both have singleton head dimensions, halving dispatch overhead. When shapes prevent fusion, the system falls back to separate quantization via self.key_codec.quantize and self.value_codec.quantize within the update_and_fetch method.
How TurboQuant KV Cache Quantization Works
TurboQuant achieves high fidelity at low bit widths through three statistical techniques:
Random Hadamard Transform (Rotation) – Before quantization, vectors are multiplied by a deterministic rotation matrix (_rotation_matrix) to distribute information uniformly across dimensions, making low-bit scalar quantization effective.
Beta-PDF Lloyd-Max Quantization – The _TurboQuantMSECodec builds a codebook matching the KV vector distribution using a Beta-probability-density-function-based Lloyd-max quantizer. This preserves the most informative directions while discarding redundant precision.
MSE-Optimized Decoding – The codec minimizes mean-square error for the given bit budget. During attention scoring, Metal kernels such as _metal_mse_score (line 58) compute scores directly from the compressed representation without fully expanding tensors to float32.
Using TurboQuant KV Cache Quantization in MLX-VLM
Integration occurs primarily through the generation pipeline in mlx_vlm/generate.py, with automatic codec selection based on CLI arguments.
Enabling TurboQuant via CLI
Users trigger quantization using the --kv-bits and --kv-quant-scheme flags. Non-integer bit widths automatically select the TurboQuant scheme:
python -m mlx_vlm.main \
--model mlx-community/nanoLLaVA-1.5-8bit \
--kv-bits 3.5 \
--kv-quant-scheme turboquant \
--max-tokens 128
Argument parsing occurs at lines 29–38 of generate.py, while cache instantiation happens at lines 255–268 via maybe_quantize_kv_cache.
Programmatic API Integration
For custom generation loops, instantiate TurboQuantKVCache and pass it through generate_step:
from mlx_vlm.generate import generate_step
from mlx_vlm.turboquant import TurboQuantKVCache
def generate_with_quantization(model, input_ids, kv_bits=4.0):
cache = TurboQuantKVCache(bits=kv_bits)
generator = generate_step(
input_ids=input_ids,
model=model,
kv_bits=kv_bits,
kv_quant_scheme="turboquant",
kv_group_size=64,
quantized_kv_start=5000,
max_tokens=20,
)
for token, _ in generator:
yield token
Manual Cache Operations
For debugging or custom attention implementations, manually update and inspect the cache:
# After forward pass, quantize new KV tensors
kv_state_keys, kv_state_values = cache.update_and_fetch(raw_keys, raw_values)
# Inspect packed representation (uint32 buffers)
print("Packed shape:", cache.keys.shape)
# Dequantize for inspection
keys_fp, values_fp = cache.dequantize(kv_state_keys, kv_state_values)
The packed tensors use uint32 words where each word encodes multiple low-bit values according to the configured bit width.
Summary
- TurboQuantKVCache in
mlx_vlm/turboquant.pyprovides the core quantized cache implementation, storing KV tensors in packeduint32buffers with per-token norms. - Bit splitting allows fractional quantization (e.g., 3.5 bits) by distributing bits unevenly between keys and values using floor/ceil allocation in
_ensure_codecs. - Fused Metal kernels optimize performance when head dimensions permit combined key-value quantization via
_try_fused_kv_quantize, falling back to separate paths for incompatible shapes. - Statistical techniques including Random Hadamard Transform and Beta-PDF Lloyd-Max quantization preserve attention fidelity at low bit widths through the
_TurboQuantMSECodec. - CLI flags
--kv-bitsand--kv-quant-scheme turboquantenable the feature inmlx_vlm/generate.py, with automatic detection of fractional bit widths triggering TurboQuant selection.
Frequently Asked Questions
What bit widths does TurboQuant KV Cache Quantization support?
TurboQuant supports arbitrary fractional bit widths (e.g., 3.5, 4.2) in addition to integer values. The _ensure_codecs method automatically splits fractional bits between keys and values—typically assigning the floor to keys and ceiling to values—to achieve the target compression rate while maintaining balanced quality.
How does TurboQuant affect generation quality compared to full-precision caches?
According to the MSE-optimized codec implementation in mlx_vlm/turboquant.py, TurboQuant minimizes mean-square error through statistical quantization and random rotation. The Metal kernels compute attention scores directly from compressed representations, preserving fidelity while reducing memory usage by 4–5x at 3.5-bit precision compared to float16.
Can I use TurboQuant with any vision-language model in MLX-VLM?
TurboQuant works with any transformer model using the standard MLX-VLM cache interface. The maybe_quantize_kv_cache wrapper in generate.py (imported at line 14) automatically swaps in TurboQuantKVCache when --kv-bits is specified and the scheme is set to turboquant, requiring no model architecture changes.
Where is the dequantization logic implemented?
The dequantize method of TurboQuantKVCache (in mlx_vlm/turboquant.py) handles reconstruction of full-precision tensors. It uses the stored rotation matrix and codebooks to decode packed uint32 buffers back to float32 on-the-fly during attention scoring or when explicit retrieval is requested.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →