Methods for Model Quantization in Deep Learning: 5 Techniques Explained
Model quantization in deep learning reduces the numerical precision of neural network weights, activations, and KV-cache entries—typically from FP32 to INT8, INT4, or FP8—to shrink memory footprints and accelerate inference on specialized hardware.
Model quantization is essential for deploying large language models on resource-constrained devices. According to the rohitg00/ai-engineering-from-scratch curriculum—specifically Phase 10’s Lesson 11 on Quantization—mastering these techniques allows engineers to trade minimal accuracy for significant gains in speed and efficiency.
Post-Training Quantization (PTQ)
Post-Training Quantization (PTQ) converts FP16 weights to lower-precision formats using per-tensor or per-channel scaling factors without requiring retraining. As documented in phases/10-llms-from-scratch/11-quantization/docs/en.md, this method is ideal for quick deployment when speed and memory budgets are tight.
PTQ works by calculating an abs_max value across the tensor, deriving a scale factor, and clipping values to the quantized range:
def quantize_symmetric(tensor, num_bits=8):
qmin = -(2 ** (num_bits - 1))
qmax = 2 ** (num_bits - 1) - 1
abs_max = np.max(np.abs(tensor))
if abs_max == 0:
return np.zeros_like(tensor, dtype=np.int32), 1.0
scale = abs_max / qmax
quantized = np.clip(np.round(tensor / scale), qmin, qmax).astype(np.int32)
return quantized, float(scale)
This approach is computationally cheap but may introduce higher error for weight matrices with irregular value distributions.
Quantization-Aware Training (QAT)
Quantization-Aware Training (QAT) inserts fake-quantize and de-quantize nodes into the forward pass during training. The model learns to adapt to rounding errors via the straight-through estimator, maximizing quality at aggressive bit-widths like INT4.
As implemented in the curriculum’s reference materials, QAT is the preferred method when you can afford a full training run and need the highest accuracy at low precision. Unlike PTQ, QAT allows the network to adjust its weights to compensate for quantization noise before deployment.
GPTQ: Hessian-Guided One-Shot Quantization
GPTQ is a one-shot PTQ method that quantizes each layer sequentially using a small calibration set to compute a second-order (Hessian) estimate of weight importance. This makes INT4 quantization practical for large LLMs with minimal quality loss.
The simulated_gptq function in phases/10-llms-from-scratch/11-quantization/code/main.py demonstrates the core logic:
def simulated_gptq(weight_matrix, calibration_inputs, num_bits=4):
n_in, n_out = weight_matrix.shape
qmin = -(2 ** (num_bits - 1))
qmax = 2 ** (num_bits - 1) - 1
# Approximate Hessian
H = np.zeros((n_in, n_in))
for x in calibration_inputs:
x = x.reshape(-1, 1) if x.ndim == 1 else x
for row in range(x.shape[0]):
xi = x[row].reshape(-1, 1)
H += xi @ xi.T
H /= len(calibration_inputs)
H += np.eye(n_in) * 1e-4
weight_importance = np.diag(H)
# Quantize each column with column-wise scale
quantized = np.zeros_like(weight_matrix, dtype=np.int32)
scales = np.zeros(n_out)
for col in range(n_out):
w_col = weight_matrix[:, col]
abs_max = np.max(np.abs(w_col))
if abs_max == 0:
scales[col] = 1.0
continue
scale = abs_max / qmax
scales[col] = scale
quantized[:, col] = np.clip(np.round(w_col / scale), qmin, qmax).astype(np.int32)
return quantized, scales, {}
GPTQ requires more computation than basic PTQ but delivers superior accuracy for large models where activations are sensitive to weight perturbations.
AWQ: Activation-Aware Weight Quantization
AWQ detects “salient” weights that interact with large activations, scales them before quantization, and leaves the rest at lower precision. This method is faster than GPTQ while preserving quality, making it suitable for fast INT4 pipelines where calibration time is limited.
According to the lesson documentation, AWQ protects the most impactful weights by considering activation magnitudes during the quantization process, whereas GPTQ relies solely on Hessian-based importance.
Mixed-Precision and GGUF Formats
Mixed-precision quantization stores different layers in different precisions—for example, keeping first and last layers at INT8 while compressing middle layers to INT4. The GGUF container format, used by llama.cpp, implements this strategy to optimize CPU and Apple Silicon inference where memory is the primary constraint.
The curriculum’s outputs/skill-quantization.md provides a decision framework that selects optimal layer-wise precision based on hardware constraints and quality tolerance.
Implementation: Bit-Level Logic to Per-Channel Scaling
The reference implementation in phases/10-llms-from-scratch/11-quantization/code/main.py provides utilities for understanding number formats and implementing quantization strategies.
To inspect FP32 bit-level representation:
def float_to_fp32_bits(value):
bits = np.float32(value).view(np.uint32)
sign = (bits >> 31) & 1
exponent = (bits >> 23) & 0xFF
mantissa = bits & 0x7FFFFF
return {"sign": int(sign), "exponent": int(exponent), "mantissa": int(mantissa)}
For weight matrices where per-tensor scaling introduces too much error, use per-channel quantization:
def quantize_per_channel(tensor, num_bits=8, axis=0):
qmin = -(2 ** (num_bits - 1))
qmax = 2 ** (num_bits - 1) - 1
if axis == 0:
abs_max = np.max(np.abs(tensor), axis=1, keepdims=True)
else:
abs_max = np.max(np.abs(tensor), axis=0, keepdims=True)
abs_max = np.where(abs_max == 0, 1.0, abs_max)
scales = abs_max / qmax
quantized = np.clip(np.round(tensor / scales), qmin, qmax).astype(np.int32)
return quantized, scales.squeeze()
Per-channel quantization reduces error for weight matrices by calculating separate scale factors for each output channel rather than using a global tensor scale.
The Sensitivity Hierarchy: What to Quantize
Not all components tolerate quantization equally. The curriculum defines a sensitivity hierarchy from most robust to most sensitive:
- Weights — Most robust to low-precision formats like INT4.
- Activations — Moderately robust; often quantized to INT8.
- KV-cache — Sensitive; typically requires INT8 or higher.
- Attention logits — Most sensitive; usually kept at FP16 or BF16 to avoid catastrophic quality loss.
Quantizing attention logits to INT4 usually degrades model behavior significantly, whereas weights can often tolerate aggressive compression without perceptible output degradation.
Summary
- PTQ offers the fastest deployment path by converting FP16 weights to INT8 or INT4 without retraining.
- QAT provides maximum quality at low bit-widths by simulating quantization during the training forward pass.
- GPTQ uses Hessian estimates from calibration data to enable practical INT4 quantization for large LLMs.
- AWQ accelerates INT4 inference by protecting salient weights identified through activation analysis.
- The GGUF format supports mixed-precision strategies that optimize specific layers for CPU inference.
- Reference implementations in
phases/10-llms-from-scratch/11-quantization/code/main.pydemonstrate bit-level manipulation, symmetric quantization, and per-channel scaling.
Frequently Asked Questions
What is the difference between PTQ and QAT?
Post-Training Quantization (PTQ) converts pre-trained weights to lower precision using calibration data but requires no gradient updates, making it fast but potentially less accurate. Quantization-Aware Training (QAT) integrates fake-quantization operations into the training graph, allowing the model to learn compensation for quantization error via backpropagation.
When should I use GPTQ versus AWQ?
Use GPTQ when you have time for calibration and need the highest accuracy at INT4 for large models, as it uses second-order Hessian information to minimize layer-wise error. Choose AWQ when calibration speed is critical and you need fast INT4 inference, as it skips the expensive Hessian computation in favor of activation-aware weight scaling.
How does per-channel quantization improve accuracy over per-tensor?
Per-channel quantization calculates separate scale factors for each output channel of a weight matrix rather than using a single global scale, which better accommodates the varying magnitude distributions across different neurons. As shown in phases/10-llms-from-scratch/11-quantization/code/main.py, this method significantly reduces mean-squared error compared to naive per-tensor symmetric quantization.
Why are attention logits usually kept at higher precision?
Attention logits exhibit the highest sensitivity to quantization in the transformer architecture; compressing them to INT4 or INT8 typically causes catastrophic degradation in model coherence and output quality. According to the sensitivity hierarchy in the curriculum documentation, weights are the most robust component, while logits require FP16 or BF16 precision to maintain generation quality.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →