# Methods for Model Quantization in Deep Learning: 5 Techniques Explained

> Explore 5 essential model quantization techniques in deep learning. Shrink memory and boost inference speed by converting FP32 to INT8, INT4, or FP8.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: deep-dive
- Published: 2026-07-19

---

**Model quantization in deep learning reduces the numerical precision of neural network weights, activations, and KV-cache entries—typically from FP32 to INT8, INT4, or FP8—to shrink memory footprints and accelerate inference on specialized hardware.**

Model quantization is essential for deploying large language models on resource-constrained devices. According to the rohitg00/ai-engineering-from-scratch curriculum—specifically Phase 10’s Lesson 11 on Quantization—mastering these techniques allows engineers to trade minimal accuracy for significant gains in speed and efficiency.

## Post-Training Quantization (PTQ)

**Post-Training Quantization (PTQ)** converts FP16 weights to lower-precision formats using per-tensor or per-channel scaling factors without requiring retraining. As documented in [`phases/10-llms-from-scratch/11-quantization/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/10-llms-from-scratch/11-quantization/docs/en.md), this method is ideal for quick deployment when speed and memory budgets are tight.

PTQ works by calculating an `abs_max` value across the tensor, deriving a scale factor, and clipping values to the quantized range:

```python
def quantize_symmetric(tensor, num_bits=8):
    qmin = -(2 ** (num_bits - 1))
    qmax = 2 ** (num_bits - 1) - 1
    abs_max = np.max(np.abs(tensor))
    if abs_max == 0:
        return np.zeros_like(tensor, dtype=np.int32), 1.0
    scale = abs_max / qmax
    quantized = np.clip(np.round(tensor / scale), qmin, qmax).astype(np.int32)
    return quantized, float(scale)

```

This approach is computationally cheap but may introduce higher error for weight matrices with irregular value distributions.

## Quantization-Aware Training (QAT)

**Quantization-Aware Training (QAT)** inserts fake-quantize and de-quantize nodes into the forward pass during training. The model learns to adapt to rounding errors via the straight-through estimator, maximizing quality at aggressive bit-widths like INT4.

As implemented in the curriculum’s reference materials, QAT is the preferred method when you can afford a full training run and need the highest accuracy at low precision. Unlike PTQ, QAT allows the network to adjust its weights to compensate for quantization noise before deployment.

## GPTQ: Hessian-Guided One-Shot Quantization

**GPTQ** is a one-shot PTQ method that quantizes each layer sequentially using a small calibration set to compute a second-order (Hessian) estimate of weight importance. This makes INT4 quantization practical for large LLMs with minimal quality loss.

The `simulated_gptq` function in [`phases/10-llms-from-scratch/11-quantization/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/10-llms-from-scratch/11-quantization/code/main.py) demonstrates the core logic:

```python
def simulated_gptq(weight_matrix, calibration_inputs, num_bits=4):
    n_in, n_out = weight_matrix.shape
    qmin = -(2 ** (num_bits - 1))
    qmax = 2 ** (num_bits - 1) - 1
    # Approximate Hessian

    H = np.zeros((n_in, n_in))
    for x in calibration_inputs:
        x = x.reshape(-1, 1) if x.ndim == 1 else x
        for row in range(x.shape[0]):
            xi = x[row].reshape(-1, 1)
            H += xi @ xi.T
    H /= len(calibration_inputs)
    H += np.eye(n_in) * 1e-4
    weight_importance = np.diag(H)
    # Quantize each column with column-wise scale

    quantized = np.zeros_like(weight_matrix, dtype=np.int32)
    scales = np.zeros(n_out)
    for col in range(n_out):
        w_col = weight_matrix[:, col]
        abs_max = np.max(np.abs(w_col))
        if abs_max == 0:
            scales[col] = 1.0
            continue
        scale = abs_max / qmax
        scales[col] = scale
        quantized[:, col] = np.clip(np.round(w_col / scale), qmin, qmax).astype(np.int32)
    return quantized, scales, {}

```

GPTQ requires more computation than basic PTQ but delivers superior accuracy for large models where activations are sensitive to weight perturbations.

## AWQ: Activation-Aware Weight Quantization

**AWQ** detects “salient” weights that interact with large activations, scales them before quantization, and leaves the rest at lower precision. This method is faster than GPTQ while preserving quality, making it suitable for fast INT4 pipelines where calibration time is limited.

According to the lesson documentation, AWQ protects the most impactful weights by considering activation magnitudes during the quantization process, whereas GPTQ relies solely on Hessian-based importance.

## Mixed-Precision and GGUF Formats

**Mixed-precision quantization** stores different layers in different precisions—for example, keeping first and last layers at INT8 while compressing middle layers to INT4. The **GGUF** container format, used by [`llama.cpp`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/llama.cpp), implements this strategy to optimize CPU and Apple Silicon inference where memory is the primary constraint.

The curriculum’s [`outputs/skill-quantization.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/outputs/skill-quantization.md) provides a decision framework that selects optimal layer-wise precision based on hardware constraints and quality tolerance.

## Implementation: Bit-Level Logic to Per-Channel Scaling

The reference implementation in [`phases/10-llms-from-scratch/11-quantization/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/10-llms-from-scratch/11-quantization/code/main.py) provides utilities for understanding number formats and implementing quantization strategies.

To inspect FP32 bit-level representation:

```python
def float_to_fp32_bits(value):
    bits = np.float32(value).view(np.uint32)
    sign = (bits >> 31) & 1
    exponent = (bits >> 23) & 0xFF
    mantissa = bits & 0x7FFFFF
    return {"sign": int(sign), "exponent": int(exponent), "mantissa": int(mantissa)}

```

For weight matrices where per-tensor scaling introduces too much error, use **per-channel quantization**:

```python
def quantize_per_channel(tensor, num_bits=8, axis=0):
    qmin = -(2 ** (num_bits - 1))
    qmax = 2 ** (num_bits - 1) - 1
    if axis == 0:
        abs_max = np.max(np.abs(tensor), axis=1, keepdims=True)
    else:
        abs_max = np.max(np.abs(tensor), axis=0, keepdims=True)
    abs_max = np.where(abs_max == 0, 1.0, abs_max)
    scales = abs_max / qmax
    quantized = np.clip(np.round(tensor / scales), qmin, qmax).astype(np.int32)
    return quantized, scales.squeeze()

```

Per-channel quantization reduces error for weight matrices by calculating separate scale factors for each output channel rather than using a global tensor scale.

## The Sensitivity Hierarchy: What to Quantize

Not all components tolerate quantization equally. The curriculum defines a sensitivity hierarchy from most robust to most sensitive:

1. **Weights** — Most robust to low-precision formats like INT4.
2. **Activations** — Moderately robust; often quantized to INT8.
3. **KV-cache** — Sensitive; typically requires INT8 or higher.
4. **Attention logits** — Most sensitive; usually kept at FP16 or BF16 to avoid catastrophic quality loss.

Quantizing attention logits to INT4 usually degrades model behavior significantly, whereas weights can often tolerate aggressive compression without perceptible output degradation.

## Summary

- **PTQ** offers the fastest deployment path by converting FP16 weights to INT8 or INT4 without retraining.
- **QAT** provides maximum quality at low bit-widths by simulating quantization during the training forward pass.
- **GPTQ** uses Hessian estimates from calibration data to enable practical INT4 quantization for large LLMs.
- **AWQ** accelerates INT4 inference by protecting salient weights identified through activation analysis.
- The **GGUF** format supports mixed-precision strategies that optimize specific layers for CPU inference.
- Reference implementations in [`phases/10-llms-from-scratch/11-quantization/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/10-llms-from-scratch/11-quantization/code/main.py) demonstrate bit-level manipulation, symmetric quantization, and per-channel scaling.

## Frequently Asked Questions

### What is the difference between PTQ and QAT?

**Post-Training Quantization (PTQ)** converts pre-trained weights to lower precision using calibration data but requires no gradient updates, making it fast but potentially less accurate. **Quantization-Aware Training (QAT)** integrates fake-quantization operations into the training graph, allowing the model to learn compensation for quantization error via backpropagation.

### When should I use GPTQ versus AWQ?

Use **GPTQ** when you have time for calibration and need the highest accuracy at INT4 for large models, as it uses second-order Hessian information to minimize layer-wise error. Choose **AWQ** when calibration speed is critical and you need fast INT4 inference, as it skips the expensive Hessian computation in favor of activation-aware weight scaling.

### How does per-channel quantization improve accuracy over per-tensor?

**Per-channel quantization** calculates separate scale factors for each output channel of a weight matrix rather than using a single global scale, which better accommodates the varying magnitude distributions across different neurons. As shown in [`phases/10-llms-from-scratch/11-quantization/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/10-llms-from-scratch/11-quantization/code/main.py), this method significantly reduces mean-squared error compared to naive per-tensor symmetric quantization.

### Why are attention logits usually kept at higher precision?

Attention logits exhibit the highest sensitivity to quantization in the transformer architecture; compressing them to INT4 or INT8 typically causes catastrophic degradation in model coherence and output quality. According to the sensitivity hierarchy in the curriculum documentation, weights are the most robust component, while logits require FP16 or BF16 precision to maintain generation quality.