# How Quantisation Reduces AI Model Size Without Losing Performance: A Technical Deep-Dive

> Discover how quantisation slashes AI model size using low-bit integers, achieving up to 75% reduction with minimal performance loss through advanced calibration and error compensation techniques.

- Repository: [Henry Ndubuaku/maths-cs-ai-compendium](https://github.com/HenryNdubuaku/maths-cs-ai-compendium)
- Tags: deep-dive
- Published: 2026-07-16

---

**Quantisation reduces AI model size by converting 32-bit or 16-bit floating-point weights into low-bit integers like INT8 or INT4, cutting memory usage by up to 75% while preserving accuracy through calibrated scaling, per-channel granularity, and error-compensation algorithms.**

The `HenryNdubuaku/maths-cs-ai-compendium` repository provides a comprehensive technical reference for optimizing large model inference. According to the source analysis in `chapter 17 - AI inference/01. quantisation.md`, quantisation exploits the statistical distribution of neural network weights to shrink numerical precision without degrading model quality.

## Why Quantise?

Quantisation delivers four measurable benefits that make massive models deployable on consumer hardware.

**Memory Reduction**
Converting from FP16 to INT8 halves storage requirements; INT4 reduces them by four times. A 70-billion-parameter model requires approximately 140 GB in FP16, but fits into roughly 35 GB when quantised to INT4, enabling single-GPU deployment on an A100.

**Throughput Gains**
Modern tensor cores provide approximately 2× faster compute for each precision step. For example, the NVIDIA H100 delivers 989 TFLOPS in FP8 compared to only 67 TFLOPS in FP32, as detailed in `chapter 16 - SIMD and GPU programming/04. GPU architecture and CUDA.md`.

**Bandwidth Savings**
Inference workloads are typically memory-bandwidth bound. Smaller weights require fewer bytes to move from VRAM to compute units, yielding near-linear speed-ups in tokens-per-second generation.

**Energy Efficiency**
Fewer bits per operation lower power draw dramatically. This scaling translates to significant cost reductions in data-center deployments and extends battery life for edge devices.

## How Accuracy Is Preserved

The preservation of model quality relies on minimising quantisation error through sophisticated mathematical techniques.

### Scale and Zero-Point Calibration

Each tensor is represented as `x_q = round(x/scale) + zero_point`. Proper scaling ensures the integer range tightly covers the original value distribution, keeping reconstruction error minimal.

### Granularity Control

The granularity of scaling dramatically impacts fidelity. According to lines 57-64 in the quantisation file:

- **Per-tensor** quantisation uses one scale for the entire matrix. This is fastest but yields lower accuracy.
- **Per-channel** quantisation assigns a unique scale per output channel or row, improving fidelity with minimal computational overhead.
- **Per-group** and **per-token** methods further refine scaling for extreme compression scenarios or large language model activations.

### Calibration Techniques

Post-training quantisation (PTQ) gathers activation statistics on a small calibration set. Techniques include percentile clipping, MSE-optimal scaling, and entropy-based calibration, which select scales that minimise quantisation error (lines 71-77).

### Quantisation-Aware Training

QAT inserts fake-quantisation nodes during the training process, allowing the model to learn weight configurations that tolerate quantisation noise (lines 107-115). This approach recovers accuracy lost by PTQ, particularly when targeting aggressive formats like INT4 or INT2.

### Weight-Only Compression Algorithms

Advanced algorithms apply second-order error compensation to maintain accuracy under extreme compression. As implemented in lines 124-182, these include:

- **GPTQ**: Column-wise quantisation with Hessian-based error compensation for high-quality INT4 weight-only compression.
- **AWQ**: Activation-aware scaling of important channels before quantisation.
- **SpQR**: Stores outlier weights in full precision while quantising the remainder aggressively.
- **HQQ**: Zero-shot quantisation requiring no calibration data.
- **AQLM**: Additive multi-codebook vector quantisation achieving 2-bit effective precision.

## Key Quantisation Techniques

Modern deployment pipelines utilise specialised algorithms tailored to specific hardware constraints.

**GPTQ** remains the standard for high-quality INT4 weight-only quantisation in large language models, using column-wise processing with Hessian-based error compensation to minimise layer-wise distortion.

**AWQ** delivers faster processing than GPTQ by identifying and scaling salient channels before quantisation, then rescaling activations during inference to maintain model fidelity.

**SpQR** achieves extreme compression by storing approximately 0.1% of outlier weights in full precision while aggressively quantising the remaining parameters, introducing minimal overhead for significant memory savings.

**KV-Cache Quantisation** targets the attention mechanism's key-value cache, converting these buffers to INT8 or INT4. This enables 100,000-token contexts on a single GPU by reducing the cache memory footprint from hundreds of gigabytes to tens of gigabytes.

**Emerging Formats** like MX (Microscaling) block-floating-point use shared exponents per block with per-element mantissas, representing an industry standard for next-generation GPU architectures.

## Practical Implementation

The following examples demonstrate core quantisation patterns from the compendium.

Simple symmetric INT8 quantisation in PyTorch:

```python
import torch

def quantise_int8(tensor, bits=8):
    qmax = 2 ** (bits - 1) - 1          # 127 for INT8

    scale = tensor.abs().max() / qmax
    q = torch.clamp(torch.round(tensor / scale), -qmax, qmax).to(torch.int8)
    return q, scale

def dequantise(q, scale):
    return q.float() * scale

# Example weight matrix

W = torch.randn(512, 512)
W_q, scl = quantise_int8(W)
W_hat = dequantise(W_q, scl)

print('Mean absolute error:', (W - W_hat).abs().mean())
print('Compression factor:', W.numel() * 4 / (W_q.numel() + 4))  # +4 bytes for scale

```

Per-channel versus per-tensor granularity in JAX:

```python
import jax.numpy as jnp
import jax

key = jax.random.PRNGKey(0)
acts = jax.random.normal(key, (32, 512)) * 0.1
acts = acts.at[:, 0].set(acts[:, 0] * 100)   # outlier channel

# Per-tensor

scale_t = jnp.max(jnp.abs(acts)) / 127.0
q_t = jnp.clip(jnp.round(acts / scale_t), -127, 127)
rec_t = q_t * scale_t

# Per-channel

scale_c = jnp.max(jnp.abs(acts), axis=0) / 127.0
q_c = jnp.clip(jnp.round(acts / scale_c), -127, 127)
rec_c = q_c * scale_c

print('Tensor error:', jnp.abs(acts - rec_t).mean())
print('Channel error:', jnp.abs(acts - rec_c).mean())

```

Memory footprint calculation for KV-cache quantisation:

```python
def kv_cache_gb(layers, heads, d_head, seq_len, bytes_per_elem):
    return 2 * layers * heads * d_head * seq_len * bytes_per_elem / 1e9

# 70 B model, 80 layers, 64 heads, 128-dim heads

fp16 = kv_cache_gb(80, 64, 128, 131072, 2)   # FP16 → 330 GB

int8 = kv_cache_gb(80, 64, 128, 131072, 1)   # INT8  → 165 GB

int4 = kv_cache_gb(80, 64, 128, 131072, 0.5) # INT4  →  82 GB

print(f'FP16: {fp16:.1f} GB, INT8: {int8:.1f} GB, INT4: {int4:.1f} GB')

```

## Summary

- **Quantisation** converts FP32/FP16 weights to low-bit integers (INT8/INT4), reducing a 70B model from 140 GB to 35 GB.
- **Accuracy preservation** relies on calibrated scaling (scale and zero-point), per-channel granularity, and advanced algorithms like GPTQ and AWQ.
- **Performance gains** come from tensor core acceleration (2× per precision step) and bandwidth savings, with H100s reaching 989 TFLOPS in low-precision modes.
- **Implementation** requires choosing between PTQ for speed of deployment or QAT for maximum accuracy recovery at ultra-low bit widths.
- **KV-cache quantisation** specifically enables long-context inference (100K+ tokens) on single-GPU setups by compressing attention buffers.

## Frequently Asked Questions

### What is the difference between Post-Training Quantisation and Quantisation-Aware Training?

Post-Training Quantisation (PTQ) applies conversion to a fully trained model using calibration data to determine optimal scales, making it fast to implement but potentially lossy at extreme compression ratios. Quantisation-Aware Training (QAT) simulates low-precision arithmetic during training by inserting fake-quantisation nodes, allowing the network to adapt its weights to quantisation noise and recover accuracy that PTQ loses, particularly for INT4 and INT2 formats.

### How does per-channel quantisation improve accuracy over per-tensor methods?

Per-channel quantisation calculates a unique scale factor for each output channel or row of a weight matrix, accommodating outliers that might otherwise dominate a single global scale. According to the compendium analysis, this approach dramatically improves fidelity with minimal overhead compared to per-tensor quantisation, which uses one scale for the entire matrix and suffers when weight distributions vary across dimensions.

### Why does lowering precision from FP16 to INT8 not always halve the model accuracy?

Neural network weights typically follow smooth statistical distributions that concentrate around zero, meaning the dynamic range required is often much smaller than the full floating-point range. By using calibration techniques like MSE-optimal or entropy-based scaling (lines 71-77), quantisation maps the dense regions of the distribution precisely to the available integer range, keeping reconstruction error below perceptible thresholds while eliminating the redundant precision used for representational headroom.

### Can quantisation work for the KV-cache in transformer models?

Yes, KV-cache quantisation specifically targets the key-value buffers that grow linearly with sequence length during autoregressive generation. Converting these caches from FP16 to INT8 or INT4 reduces memory consumption from hundreds of gigabytes to manageable sizes—enabling 100,000-token contexts on a single GPU. This technique is crucial for deploying long-context large language models without requiring multi-node inference infrastructure.