Quantization in AI Inference: A Complete Guide to Low-Precision Model Deployment

The quantization chapter in the HenryNdubuaku/maths-cs-ai-compendium repository provides a comprehensive walkthrough of reducing model weights and activations from high-precision floating-point formats to integer representations, covering post-training quantization, quantization-aware training, and state-of-the-art weight-only methods like GPTQ and AWQ.

The repository HenryNdubuaku/maths-cs-ai-compendium serves as a technical reference for machine learning engineers optimizing large language models. Located at chapter 17: AI inference/01. quantisation.md, the chapter delivers an end-to-end analysis of quantization techniques that make LLM inference faster, smaller, and cheaper by lowering precision from FP32 to INT8, INT4, or even binary formats.

Why Quantization Matters for LLM Inference

Modern large language models are bandwidth-bound during inference, meaning the speed of token generation depends on how quickly weights can be moved from memory to compute units. Reducing the number of bytes per weight directly alleviates this bottleneck, enabling real-time inference on edge devices and reducing cloud compute costs. The chapter emphasizes that quantization is not merely about compression—it is a fundamental optimization for the memory-bandwidth wall facing transformer architectures.

Number Formats and Granularity Levels

The chapter catalogs the numeric representations available for quantized deployment:

  • INT8: The standard 8-bit integer format offering a balance between accuracy and efficiency
  • INT4: Aggressive compression with 16x reduction in model size
  • INT2 and Ternary (BitNet): Extreme quantization for ultra-low-power scenarios
  • Mixed-precision schemes: Combining different bit widths for different layers

Granularity determines how scales are applied. Per-tensor scaling uses one scale factor for an entire weight matrix, while per-channel and per-group scaling offer finer granularity to preserve accuracy at the cost of slightly higher metadata overhead.

Post-Training Quantization vs. Quantization-Aware Training

Post-Training Quantization (PTQ)

PTQ converts pre-trained models to lower precision without retraining. The chapter details calibration-set collection and scaling strategies including min-max, MSE-optimal, and KL-based scaling. It provides the symmetric quantization formula where scale = tensor.abs().max() / qmax and quantized values are clamped to the integer range.

Quantization-Aware Training (QAT)

QAT integrates quantization into the training loop using fake-quant nodes that simulate low-precision arithmetic during forward passes. The straight-through estimator allows gradients to flow through these non-differentiable operations. The chapter recommends QAT specifically for low-bit INT4/INT2 configurations or edge deployments where accuracy loss from PTQ is unacceptable.

State-of-the-Art Weight-Only Quantization

The chapter dedicates significant coverage to modern weight-only methods that preserve model quality while maximizing compression:

  • GPTQ: Uses Hessian-based error compensation to quantize weights layer-by-layer in a single pass
  • AWQ: Applies activation-aware channel scaling to protect the most important weight parameters
  • QuIP/QuIP#: Employs lattice codebooks for optimal vector quantization
  • SpQR: Handles outlier values separately to prevent quantization error accumulation
  • HQQ, AQLM, and BitNet: Alternative approaches for extreme compression and ternary weight representations

Activation and KV-Cache Quantization Strategies

Mixed-Precision Inference

Mixed-precision schemes store weights in INT4 while keeping activations in FP16, striking a balance between memory savings and computational precision. This approach requires careful handling of the dequantization points during the forward pass.

KV-Cache Quantization

The chapter addresses memory bottlenecks in long-context inference by quantizing the key-value cache. INT8 and INT4 KV-cache quantization reduce the memory footprint of attention mechanisms, allowing longer sequence lengths on consumer hardware. The implementation involves per-token or per-channel scaling factors stored alongside the cached tensors.

SmoothQuant for Activation Quantization

SmoothQuant shifts the quantization difficulty from activations (which have outliers) to weights (which are more uniform) by mathematically smoothing the activation distributions. This enables static activation quantization without the accuracy degradation typically associated with dynamic per-token scaling.

Practical Implementation Guide

The chapter provides minimal Python implementations demonstrating core quantization operations.

Symmetric INT8 Quantization

import torch

def symmetric_quantise(tensor: torch.Tensor, bits: int = 8):
    qmax = 2 ** (bits - 1) - 1          # 127 for 8‑bit

    qmin = -qmax - 1                    # -128 for 8‑bit

    scale = tensor.abs().max() / qmax   # symmetric scale

    quantised = torch.clamp(torch.round(tensor / scale), qmin, qmax).to(torch.int8)
    return quantised, scale

def dequantise(quantised: torch.Tensor, scale: float):
    return quantised.float() * scale

# Example

weight = torch.randn(4, 4) * 0.1
q_weight, s = symmetric_quantise(weight, bits=8)
recon = dequantise(q_weight, s)
print('Reconstruction error:', (weight - recon).abs().mean())

MSE-Optimal Calibration

import numpy as np

def ptq_mse(tensor: np.ndarray, bits: int = 8):
    qmax = 2 ** (bits - 1) - 1
    # Search over possible clipping values

    best_clip, best_err = None, np.inf
    for clip in np.linspace(tensor.min(), tensor.max(), 100):
        scale = clip / qmax
        quant = np.clip(np.round(tensor / scale), -qmax, qmax).astype(np.int8)
        recon = quant.astype(np.float32) * scale
        err = ((tensor - recon) ** 2).mean()
        if err < best_err:
            best_err, best_clip = err, clip
    return best_clip / qmax

KV-Cache Quantization

import torch

def quantise_kv(cache: torch.Tensor, bits: int = 8):
    # cache shape: (seq_len, num_heads, head_dim)

    qmax = 2 ** (bits - 1) - 1
    scale = cache.abs().max() / qmax
    q_cache = torch.clamp(torch.round(cache / scale), -qmax, qmax).to(torch.int8)
    return q_cache, scale

# De‑quantise on‑the‑fly during attention

def dequantise_kv(q_cache, scale):
    return q_cache.float() * scale

Sensitivity Analysis and Future Directions

The chapter concludes with practical guidance on layer-wise sensitivity analysis, enabling per-layer bit-allocation where sensitive layers retain higher precision. It also discusses emerging formats like GGUF (used by llama.cpp) and the trade-offs of extreme low-bit schemes such as BitNet's ternary weights.

Summary

  • Quantization in AI inference reduces FP32 weights to INT8/INT4/INT2 to overcome memory bandwidth bottlenecks in LLMs
  • Post-training quantization (PTQ) offers fast deployment without retraining, while quantization-aware training (QAT) provides superior accuracy for low-bit regimes
  • Modern weight-only methods (GPTQ, AWQ, QuIP, SpQR) use Hessian information and activation-aware scaling to minimize precision loss
  • KV-cache quantization and SmoothQuant address memory constraints in long-context inference and activation outliers
  • The chapter 17: AI inference/01. quantisation.md file provides runnable PyTorch implementations for symmetric quantization, MSE-optimal calibration, and cache compression

Frequently Asked Questions

What is the difference between post-training quantization and quantization-aware training?

Post-training quantization converts already-trained models to lower precision using calibration data to determine scaling factors, making it fast to deploy. Quantization-aware training simulates low-precision operations during the training process using fake-quant nodes and the straight-through estimator, which typically yields higher accuracy for aggressive bit widths like INT4 or INT2 but requires full model retraining.

How does GPTQ differ from AWQ for weight-only quantization?

GPTQ quantizes weights layer-by-layer using approximate second-order information (Hessian-based compensation) to minimize reconstruction error in a single pass. AWQ instead protects salient weight channels by observing activation magnitudes, scaling weights to reduce the impact of quantization on the most important parameters without requiring computational heavy optimization.

What is KV-cache quantization and why is it important for inference?

KV-cache quantization compresses the stored key and value tensors in transformer attention mechanisms from FP16/FP32 to INT8 or INT4. This is crucial for long-context inference because the KV-cache grows linearly with sequence length and can exceed model weight size, often becoming the limiting factor for batch size and context length on consumer GPUs.

When should I use INT4 versus INT8 quantization for model deployment?

Use INT8 quantization when accuracy is critical and hardware supports efficient INT8 arithmetic, as it typically preserves 99%+ of model quality. Reserve INT4 for extreme memory-constrained environments or when using state-of-the-art weight-only methods like GPTQ or AWQ that compensate for the increased quantization error, as naive INT4 quantization can significantly degrade model performance.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →