Quantization in AI Inference: A Complete Guide to Low-Precision Model Deployment
The quantization chapter in the HenryNdubuaku/maths-cs-ai-compendium repository provides a comprehensive walkthrough of reducing model weights and activations from high-precision floating-point formats to integer representations, covering post-training quantization, quantization-aware training, and state-of-the-art weight-only methods like GPTQ and AWQ.
The repository HenryNdubuaku/maths-cs-ai-compendium serves as a technical reference for machine learning engineers optimizing large language models. Located at chapter 17: AI inference/01. quantisation.md, the chapter delivers an end-to-end analysis of quantization techniques that make LLM inference faster, smaller, and cheaper by lowering precision from FP32 to INT8, INT4, or even binary formats.
Why Quantization Matters for LLM Inference
Modern large language models are bandwidth-bound during inference, meaning the speed of token generation depends on how quickly weights can be moved from memory to compute units. Reducing the number of bytes per weight directly alleviates this bottleneck, enabling real-time inference on edge devices and reducing cloud compute costs. The chapter emphasizes that quantization is not merely about compression—it is a fundamental optimization for the memory-bandwidth wall facing transformer architectures.
Number Formats and Granularity Levels
The chapter catalogs the numeric representations available for quantized deployment:
- INT8: The standard 8-bit integer format offering a balance between accuracy and efficiency
- INT4: Aggressive compression with 16x reduction in model size
- INT2 and Ternary (BitNet): Extreme quantization for ultra-low-power scenarios
- Mixed-precision schemes: Combining different bit widths for different layers
Granularity determines how scales are applied. Per-tensor scaling uses one scale factor for an entire weight matrix, while per-channel and per-group scaling offer finer granularity to preserve accuracy at the cost of slightly higher metadata overhead.
Post-Training Quantization vs. Quantization-Aware Training
Post-Training Quantization (PTQ)
PTQ converts pre-trained models to lower precision without retraining. The chapter details calibration-set collection and scaling strategies including min-max, MSE-optimal, and KL-based scaling. It provides the symmetric quantization formula where scale = tensor.abs().max() / qmax and quantized values are clamped to the integer range.
Quantization-Aware Training (QAT)
QAT integrates quantization into the training loop using fake-quant nodes that simulate low-precision arithmetic during forward passes. The straight-through estimator allows gradients to flow through these non-differentiable operations. The chapter recommends QAT specifically for low-bit INT4/INT2 configurations or edge deployments where accuracy loss from PTQ is unacceptable.
State-of-the-Art Weight-Only Quantization
The chapter dedicates significant coverage to modern weight-only methods that preserve model quality while maximizing compression:
- GPTQ: Uses Hessian-based error compensation to quantize weights layer-by-layer in a single pass
- AWQ: Applies activation-aware channel scaling to protect the most important weight parameters
- QuIP/QuIP#: Employs lattice codebooks for optimal vector quantization
- SpQR: Handles outlier values separately to prevent quantization error accumulation
- HQQ, AQLM, and BitNet: Alternative approaches for extreme compression and ternary weight representations
Activation and KV-Cache Quantization Strategies
Mixed-Precision Inference
Mixed-precision schemes store weights in INT4 while keeping activations in FP16, striking a balance between memory savings and computational precision. This approach requires careful handling of the dequantization points during the forward pass.
KV-Cache Quantization
The chapter addresses memory bottlenecks in long-context inference by quantizing the key-value cache. INT8 and INT4 KV-cache quantization reduce the memory footprint of attention mechanisms, allowing longer sequence lengths on consumer hardware. The implementation involves per-token or per-channel scaling factors stored alongside the cached tensors.
SmoothQuant for Activation Quantization
SmoothQuant shifts the quantization difficulty from activations (which have outliers) to weights (which are more uniform) by mathematically smoothing the activation distributions. This enables static activation quantization without the accuracy degradation typically associated with dynamic per-token scaling.
Practical Implementation Guide
The chapter provides minimal Python implementations demonstrating core quantization operations.
Symmetric INT8 Quantization
import torch
def symmetric_quantise(tensor: torch.Tensor, bits: int = 8):
qmax = 2 ** (bits - 1) - 1 # 127 for 8‑bit
qmin = -qmax - 1 # -128 for 8‑bit
scale = tensor.abs().max() / qmax # symmetric scale
quantised = torch.clamp(torch.round(tensor / scale), qmin, qmax).to(torch.int8)
return quantised, scale
def dequantise(quantised: torch.Tensor, scale: float):
return quantised.float() * scale
# Example
weight = torch.randn(4, 4) * 0.1
q_weight, s = symmetric_quantise(weight, bits=8)
recon = dequantise(q_weight, s)
print('Reconstruction error:', (weight - recon).abs().mean())
MSE-Optimal Calibration
import numpy as np
def ptq_mse(tensor: np.ndarray, bits: int = 8):
qmax = 2 ** (bits - 1) - 1
# Search over possible clipping values
best_clip, best_err = None, np.inf
for clip in np.linspace(tensor.min(), tensor.max(), 100):
scale = clip / qmax
quant = np.clip(np.round(tensor / scale), -qmax, qmax).astype(np.int8)
recon = quant.astype(np.float32) * scale
err = ((tensor - recon) ** 2).mean()
if err < best_err:
best_err, best_clip = err, clip
return best_clip / qmax
KV-Cache Quantization
import torch
def quantise_kv(cache: torch.Tensor, bits: int = 8):
# cache shape: (seq_len, num_heads, head_dim)
qmax = 2 ** (bits - 1) - 1
scale = cache.abs().max() / qmax
q_cache = torch.clamp(torch.round(cache / scale), -qmax, qmax).to(torch.int8)
return q_cache, scale
# De‑quantise on‑the‑fly during attention
def dequantise_kv(q_cache, scale):
return q_cache.float() * scale
Sensitivity Analysis and Future Directions
The chapter concludes with practical guidance on layer-wise sensitivity analysis, enabling per-layer bit-allocation where sensitive layers retain higher precision. It also discusses emerging formats like GGUF (used by llama.cpp) and the trade-offs of extreme low-bit schemes such as BitNet's ternary weights.
Summary
- Quantization in AI inference reduces FP32 weights to INT8/INT4/INT2 to overcome memory bandwidth bottlenecks in LLMs
- Post-training quantization (PTQ) offers fast deployment without retraining, while quantization-aware training (QAT) provides superior accuracy for low-bit regimes
- Modern weight-only methods (GPTQ, AWQ, QuIP, SpQR) use Hessian information and activation-aware scaling to minimize precision loss
- KV-cache quantization and SmoothQuant address memory constraints in long-context inference and activation outliers
- The
chapter 17: AI inference/01. quantisation.mdfile provides runnable PyTorch implementations for symmetric quantization, MSE-optimal calibration, and cache compression
Frequently Asked Questions
What is the difference between post-training quantization and quantization-aware training?
Post-training quantization converts already-trained models to lower precision using calibration data to determine scaling factors, making it fast to deploy. Quantization-aware training simulates low-precision operations during the training process using fake-quant nodes and the straight-through estimator, which typically yields higher accuracy for aggressive bit widths like INT4 or INT2 but requires full model retraining.
How does GPTQ differ from AWQ for weight-only quantization?
GPTQ quantizes weights layer-by-layer using approximate second-order information (Hessian-based compensation) to minimize reconstruction error in a single pass. AWQ instead protects salient weight channels by observing activation magnitudes, scaling weights to reduce the impact of quantization on the most important parameters without requiring computational heavy optimization.
What is KV-cache quantization and why is it important for inference?
KV-cache quantization compresses the stored key and value tensors in transformer attention mechanisms from FP16/FP32 to INT8 or INT4. This is crucial for long-context inference because the KV-cache grows linearly with sequence length and can exceed model weight size, often becoming the limiting factor for batch size and context length on consumer GPUs.
When should I use INT4 versus INT8 quantization for model deployment?
Use INT8 quantization when accuracy is critical and hardware supports efficient INT8 arithmetic, as it typically preserves 99%+ of model quality. Reserve INT4 for extreme memory-constrained environments or when using state-of-the-art weight-only methods like GPTQ or AWQ that compensate for the increased quantization error, as naive INT4 quantization can significantly degrade model performance.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →