# Implementing LLM Quantization for Inference: A Complete Guide to INT8, GPTQ, and GGUF

> Master LLM quantization for inference. Learn INT8, GPTQ, and GGUF to cut model size by 75% without losing accuracy. Accelerate your AI models today.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: tutorial
- Published: 2026-07-26

---

**Implementing LLM quantization for inference reduces model size by 50-75% while maintaining accuracy through techniques like INT8 symmetric quantization, GPTQ for INT4 compression, and the GGUF container format used by llama.cpp.**

The **AI-Engineering-From-Scratch** repository provides a comprehensive curriculum for implementing LLM quantization for inference, walking learners from mathematical foundations to production deployment. In Phase 10, Lesson 11, the curriculum teaches how to shrink large language model weights from FP16 to INT8/INT4 and package them in the GGUF format for edge and server deployment.

## Number Formats for Quantized Inference

Understanding binary number representations is essential before implementing compression strategies. The curriculum defines the hierarchy of precision formats used in modern inference:

- **FP32** (32-bit): `[1 sign][8 exponent][23 mantissa]` — Used for training and high-precision reference
- **BF16** (16-bit): `[1 sign][8 exponent][7 mantissa]` — Google's training-phase mixed-precision format
- **FP16** (16-bit): `[1 sign][5 exponent][10 mantissa]` — Baseline inference standard
- **FP8 (E4M3)** (8-bit): `[1 sign][4 exponent][3 mantissa]` — H100 GPU inference speed-up
- **INT8** (8-bit): `[1 sign][7 value]` — Integer arithmetic on GPUs and CPUs
- **INT4** (4-bit): `[1 sign][3 value]` — Extreme compression for laptops and mobile devices

These definitions are documented in [`phases/10-llms-from-scratch/11-quantization/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/10-llms-from-scratch/11-quantization/docs/en.md) (lines 35-43), providing the theoretical backbone for quantization decisions.

## Core Quantization Operations

The reference implementation in [`phases/10-llms-from-scratch/11-quantization/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/10-llms-from-scratch/11-quantization/code/main.py) demonstrates three fundamental quantization schemes.

### Symmetric Quantization

Symmetric quantization maps floating-point values to integers around zero using a single scale factor:

```python
scale = max(abs(tensor)) / 127
quantized = np.clip(np.round(tensor / scale), -128, 127).astype(np.int32)

```

Dequantization reverses the process:

```python
reconstructed = quantized.astype(np.float64) * scale

```

### Asymmetric Quantization

For non-centered distributions, asymmetric quantization adds a **zero-point** offset:

```python
scale = (tensor.max() - tensor.min()) / (qmax - qmin)
zero_point = int(np.round(qmin - tensor.min() / scale))
quantized = np.clip(np.round(tensor / scale) + zero_point, qmin, qmax)

```

### Per-Channel Quantization

Production implementations typically apply different scales per output channel. The `quantize_per_channel` function in [`main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/main.py) repeats scale computation for each channel along the specified axis, reducing error compared to global quantization.

## Sensitivity Hierarchy in Transformers

Not all tensors tolerate quantization equally. The curriculum visualizes a sensitivity hierarchy (shown in the Mermaid diagram at lines 22-40 of the docs) that drives deployment strategy:

1. **Weights** — Most tolerant; INT4 viable with advanced methods
2. **Activations** — Moderate sensitivity; require INT8 or per-channel scaling
3. **KV Cache** — High sensitivity; errors compound across long contexts
4. **Attention Logits** — Most sensitive; typically retained in FP16/BF16

This hierarchy determines where to apply aggressive compression versus preserving precision.

## PTQ vs QAT: Choosing Your Strategy

The curriculum compares two primary quantization paradigms:

| Aspect | PTQ (Post-Training) | QAT (Quant-Aware Training) |
|--------|---------------------|----------------------------|
| **Time Cost** | Minutes to hours | Full training run |
| **INT8 Quality** | Excellent (< 0.1% loss) | Excellent |
| **INT4 Quality** | Good with GPTQ/AWQ (1-3% loss) | Superior (< 1% loss) |
| **Calibration Data** | 128-1024 examples | Full dataset |

PTQ is the pragmatic choice for most production pipelines, while QAT suits scenarios where maximum quality at INT4 is required.

## GPTQ, AWQ, and the GGUF Format

Three advanced methods bridge the gap between theory and production deployment:

**GPTQ** (Gradient-based Post-training Quantization) uses Hessian-guided, one-shot PTQ to make INT4 practical for LLMs. It analyzes layer-wise error sensitivity to minimize accuracy degradation.

**AWQ** (Activation-aware Weight Quantization) identifies and rescales the ~1% most salient weights, delivering 1.5-2× speed-up over GPTQ with similar compression ratios.

**GGUF** (GPT-Generated Unified Format) is the container specification used by llama.cpp, storing mixed-precision weight blocks (e.g., Q4_K_M) alongside tokenizer and metadata. These files load directly into inference engines like llama.cpp, vLLM, or SGLang.

## Practical Implementation: Python Reference

The repository provides a self-contained reference script in [`main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/main.py) that demonstrates quantization workflows. Below is a practical example for implementing INT8 per-channel quantization:

```python
import numpy as np

def quantize_per_channel(tensor, num_bits=8, axis=0):
    """Per-channel symmetric quantization as implemented in the curriculum."""
    scales = []
    quantized_slices = []
    
    for i in range(tensor.shape[axis]):
        channel = tensor.take(i, axis=axis)
        max_val = np.max(np.abs(channel))
        scale = max_val / (2 ** (num_bits - 1) - 1)
        
        if scale == 0:
            scale = 1e-8
            
        q = np.clip(np.round(channel / scale), 
                   -(2 ** (num_bits - 1)), 
                   2 ** (num_bits - 1) - 1).astype(np.int32)
        
        scales.append(scale)
        quantized_slices.append(q)
    
    return np.stack(quantized_slices, axis=axis), np.array(scales)

def quantization_error(original, reconstructed):
    """Calculate MSE, SNR, and cosine similarity metrics."""
    mse = np.mean((original - reconstructed) ** 2)
    signal_power = np.mean(original ** 2)
    snr_db = 10 * np.log10(signal_power / mse) if mse > 0 else float('inf')
    cosine_sim = np.dot(original.flatten(), reconstructed.flatten()) / (
        np.linalg.norm(original) * np.linalg.norm(reconstructed)
    )
    return {
        "mse": mse,
        "snr_db": snr_db,
        "cosine_similarity": cosine_sim,
        "max_error": np.max(np.abs(original - reconstructed))
    }

# Usage: Quantize a 768×3072 weight matrix

weights = np.random.randn(768, 3072).astype(np.float32)
q_int8, scales = quantize_per_channel(weights, num_bits=8, axis=0)
reconstructed = q_int8.astype(np.float32) * scales.reshape(-1, 1)

error_metrics = quantization_error(weights, reconstructed)
print(f"INT8 MSE: {error_metrics['mse']:.6f}")
print(f"SNR: {error_metrics['snr_db']:.2f} dB")

```

The `bit_width_sweep` function in the same file enables rapid experimentation across 4-bit, 8-bit, and 16-bit configurations to evaluate compression ratios versus error metrics.

## Production Deployment Pipeline

The curriculum connects quantization theory to deployment in the "Speculative Decoding Server" capstone (Phase 19, Lesson 14). The end-to-end workflow for implementing LLM quantization for inference follows four stages:

1. **Train** the base model in FP16 or BF16 precision
2. **Apply GPTQ** or AWQ to generate a quantized GGUF file (e.g., `model-Q4_K_M.gguf`)
3. **Deploy** using llama.cpp for edge devices or vLLM for GPU clusters
4. **Measure** latency, tokens-per-second, and quality metrics (perplexity/benchmarks)

The skill artifact at [`phases/10-llms-from-scratch/11-quantization/outputs/skill-quantization.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/10-llms-from-scratch/11-quantization/outputs/skill-quantization.md) provides configuration flags for each inference engine, ensuring quantization decisions propagate correctly to production.

## Summary

- **INT8 symmetric quantization** is the production sweet spot, offering 2× compression with negligible accuracy loss on modern GPUs.
- **INT4 with GPTQ/AWQ** enables deployment on resource-constrained devices when paired with the GGUF container format.
- The reference implementation in [`main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/main.py) provides production-ready functions for `quantize_per_channel`, `quantize_asymmetric`, and error analysis.
- Quantization sensitivity varies across transformer components—keep attention logits in higher precision while aggressively compressing weights.
- The curriculum's skill artifacts bridge experimentation and deployment, documenting engine-specific configurations for llama.cpp and vLLM.

## Frequently Asked Questions

### What is the difference between GPTQ and AWQ for LLM quantization?

**GPTQ** uses second-order Hessian information to minimize quantization error in a one-shot post-training step, making INT4 practical for large models. **AWQ** (Activation-aware Weight Quantization) improves upon this by identifying and rescaling the most salient 1% of weights based on activation magnitudes, typically delivering 1.5-2× throughput improvement over GPTQ with similar bit-widths. Both methods produce weights suitable for GGUF packaging.

### When should I use INT8 versus INT4 quantization?

Choose **INT8** for server-class GPUs where memory bandwidth is the bottleneck but compute resources are abundant—this provides 2× compression with less than 0.1% accuracy degradation. Opt for **INT4** only when deploying to edge devices, laptops, or mobile hardware with severe memory constraints, as INT4 requires advanced methods like GPTQ or AWQ to maintain acceptable quality (typically 1-3% loss).

### How does the GGUF format differ from other model serialization formats?

**GGUF** is specifically optimized for quantized inference on CPU and heterogeneous hardware, storing weights in mixed-precision blocks (e.g., Q4_K_M) alongside vocabulary and metadata in a single file. Unlike PyTorch's pickle-based format or SafeTensors, GGUF is designed for memory-mapped loading and efficient CPU inference via llama.cpp, making it the standard container for quantized LLM deployment outside of pure GPU environments.

### Can I implement quantization without retraining my model?

Yes, through **Post-Training Quantization (PTQ)**. The reference implementation demonstrates symmetric and asymmetric PTQ that converts existing FP16 weights to INT8 or INT4 using calibration data (128-1024 examples) without gradient updates. This differs from Quantization-Aware Training (QAT), which simulates quantization during forward passes throughout training or fine-tuning to achieve superior INT4 results.

### What metrics should I monitor when evaluating quantized models?

Monitor **MSE** (mean squared error) and **cosine similarity** between original and reconstructed weights during development, then validate with **perplexity** and downstream task benchmarks on a holdout set. The `quantization_error` function in the curriculum's [`main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/main.py) calculates these automatically, including SNR in decibels to quantify signal preservation across quantization noise.