Implementing LLM Quantization for Inference: A Complete Guide to INT8, GPTQ, and GGUF

Implementing LLM quantization for inference reduces model size by 50-75% while maintaining accuracy through techniques like INT8 symmetric quantization, GPTQ for INT4 compression, and the GGUF container format used by llama.cpp.

The AI-Engineering-From-Scratch repository provides a comprehensive curriculum for implementing LLM quantization for inference, walking learners from mathematical foundations to production deployment. In Phase 10, Lesson 11, the curriculum teaches how to shrink large language model weights from FP16 to INT8/INT4 and package them in the GGUF format for edge and server deployment.

Number Formats for Quantized Inference

Understanding binary number representations is essential before implementing compression strategies. The curriculum defines the hierarchy of precision formats used in modern inference:

  • FP32 (32-bit): [1 sign][8 exponent][23 mantissa] — Used for training and high-precision reference
  • BF16 (16-bit): [1 sign][8 exponent][7 mantissa] — Google's training-phase mixed-precision format
  • FP16 (16-bit): [1 sign][5 exponent][10 mantissa] — Baseline inference standard
  • FP8 (E4M3) (8-bit): [1 sign][4 exponent][3 mantissa] — H100 GPU inference speed-up
  • INT8 (8-bit): [1 sign][7 value] — Integer arithmetic on GPUs and CPUs
  • INT4 (4-bit): [1 sign][3 value] — Extreme compression for laptops and mobile devices

These definitions are documented in phases/10-llms-from-scratch/11-quantization/docs/en.md (lines 35-43), providing the theoretical backbone for quantization decisions.

Core Quantization Operations

The reference implementation in phases/10-llms-from-scratch/11-quantization/code/main.py demonstrates three fundamental quantization schemes.

Symmetric Quantization

Symmetric quantization maps floating-point values to integers around zero using a single scale factor:

scale = max(abs(tensor)) / 127
quantized = np.clip(np.round(tensor / scale), -128, 127).astype(np.int32)

Dequantization reverses the process:

reconstructed = quantized.astype(np.float64) * scale

Asymmetric Quantization

For non-centered distributions, asymmetric quantization adds a zero-point offset:

scale = (tensor.max() - tensor.min()) / (qmax - qmin)
zero_point = int(np.round(qmin - tensor.min() / scale))
quantized = np.clip(np.round(tensor / scale) + zero_point, qmin, qmax)

Per-Channel Quantization

Production implementations typically apply different scales per output channel. The quantize_per_channel function in main.py repeats scale computation for each channel along the specified axis, reducing error compared to global quantization.

Sensitivity Hierarchy in Transformers

Not all tensors tolerate quantization equally. The curriculum visualizes a sensitivity hierarchy (shown in the Mermaid diagram at lines 22-40 of the docs) that drives deployment strategy:

  1. Weights — Most tolerant; INT4 viable with advanced methods
  2. Activations — Moderate sensitivity; require INT8 or per-channel scaling
  3. KV Cache — High sensitivity; errors compound across long contexts
  4. Attention Logits — Most sensitive; typically retained in FP16/BF16

This hierarchy determines where to apply aggressive compression versus preserving precision.

PTQ vs QAT: Choosing Your Strategy

The curriculum compares two primary quantization paradigms:

Aspect PTQ (Post-Training) QAT (Quant-Aware Training)
Time Cost Minutes to hours Full training run
INT8 Quality Excellent (< 0.1% loss) Excellent
INT4 Quality Good with GPTQ/AWQ (1-3% loss) Superior (< 1% loss)
Calibration Data 128-1024 examples Full dataset

PTQ is the pragmatic choice for most production pipelines, while QAT suits scenarios where maximum quality at INT4 is required.

GPTQ, AWQ, and the GGUF Format

Three advanced methods bridge the gap between theory and production deployment:

GPTQ (Gradient-based Post-training Quantization) uses Hessian-guided, one-shot PTQ to make INT4 practical for LLMs. It analyzes layer-wise error sensitivity to minimize accuracy degradation.

AWQ (Activation-aware Weight Quantization) identifies and rescales the ~1% most salient weights, delivering 1.5-2× speed-up over GPTQ with similar compression ratios.

GGUF (GPT-Generated Unified Format) is the container specification used by llama.cpp, storing mixed-precision weight blocks (e.g., Q4_K_M) alongside tokenizer and metadata. These files load directly into inference engines like llama.cpp, vLLM, or SGLang.

Practical Implementation: Python Reference

The repository provides a self-contained reference script in main.py that demonstrates quantization workflows. Below is a practical example for implementing INT8 per-channel quantization:

import numpy as np

def quantize_per_channel(tensor, num_bits=8, axis=0):
    """Per-channel symmetric quantization as implemented in the curriculum."""
    scales = []
    quantized_slices = []
    
    for i in range(tensor.shape[axis]):
        channel = tensor.take(i, axis=axis)
        max_val = np.max(np.abs(channel))
        scale = max_val / (2 ** (num_bits - 1) - 1)
        
        if scale == 0:
            scale = 1e-8
            
        q = np.clip(np.round(channel / scale), 
                   -(2 ** (num_bits - 1)), 
                   2 ** (num_bits - 1) - 1).astype(np.int32)
        
        scales.append(scale)
        quantized_slices.append(q)
    
    return np.stack(quantized_slices, axis=axis), np.array(scales)

def quantization_error(original, reconstructed):
    """Calculate MSE, SNR, and cosine similarity metrics."""
    mse = np.mean((original - reconstructed) ** 2)
    signal_power = np.mean(original ** 2)
    snr_db = 10 * np.log10(signal_power / mse) if mse > 0 else float('inf')
    cosine_sim = np.dot(original.flatten(), reconstructed.flatten()) / (
        np.linalg.norm(original) * np.linalg.norm(reconstructed)
    )
    return {
        "mse": mse,
        "snr_db": snr_db,
        "cosine_similarity": cosine_sim,
        "max_error": np.max(np.abs(original - reconstructed))
    }

# Usage: Quantize a 768×3072 weight matrix

weights = np.random.randn(768, 3072).astype(np.float32)
q_int8, scales = quantize_per_channel(weights, num_bits=8, axis=0)
reconstructed = q_int8.astype(np.float32) * scales.reshape(-1, 1)

error_metrics = quantization_error(weights, reconstructed)
print(f"INT8 MSE: {error_metrics['mse']:.6f}")
print(f"SNR: {error_metrics['snr_db']:.2f} dB")

The bit_width_sweep function in the same file enables rapid experimentation across 4-bit, 8-bit, and 16-bit configurations to evaluate compression ratios versus error metrics.

Production Deployment Pipeline

The curriculum connects quantization theory to deployment in the "Speculative Decoding Server" capstone (Phase 19, Lesson 14). The end-to-end workflow for implementing LLM quantization for inference follows four stages:

  1. Train the base model in FP16 or BF16 precision
  2. Apply GPTQ or AWQ to generate a quantized GGUF file (e.g., model-Q4_K_M.gguf)
  3. Deploy using llama.cpp for edge devices or vLLM for GPU clusters
  4. Measure latency, tokens-per-second, and quality metrics (perplexity/benchmarks)

The skill artifact at phases/10-llms-from-scratch/11-quantization/outputs/skill-quantization.md provides configuration flags for each inference engine, ensuring quantization decisions propagate correctly to production.

Summary

  • INT8 symmetric quantization is the production sweet spot, offering 2× compression with negligible accuracy loss on modern GPUs.
  • INT4 with GPTQ/AWQ enables deployment on resource-constrained devices when paired with the GGUF container format.
  • The reference implementation in main.py provides production-ready functions for quantize_per_channel, quantize_asymmetric, and error analysis.
  • Quantization sensitivity varies across transformer components—keep attention logits in higher precision while aggressively compressing weights.
  • The curriculum's skill artifacts bridge experimentation and deployment, documenting engine-specific configurations for llama.cpp and vLLM.

Frequently Asked Questions

What is the difference between GPTQ and AWQ for LLM quantization?

GPTQ uses second-order Hessian information to minimize quantization error in a one-shot post-training step, making INT4 practical for large models. AWQ (Activation-aware Weight Quantization) improves upon this by identifying and rescaling the most salient 1% of weights based on activation magnitudes, typically delivering 1.5-2× throughput improvement over GPTQ with similar bit-widths. Both methods produce weights suitable for GGUF packaging.

When should I use INT8 versus INT4 quantization?

Choose INT8 for server-class GPUs where memory bandwidth is the bottleneck but compute resources are abundant—this provides 2× compression with less than 0.1% accuracy degradation. Opt for INT4 only when deploying to edge devices, laptops, or mobile hardware with severe memory constraints, as INT4 requires advanced methods like GPTQ or AWQ to maintain acceptable quality (typically 1-3% loss).

How does the GGUF format differ from other model serialization formats?

GGUF is specifically optimized for quantized inference on CPU and heterogeneous hardware, storing weights in mixed-precision blocks (e.g., Q4_K_M) alongside vocabulary and metadata in a single file. Unlike PyTorch's pickle-based format or SafeTensors, GGUF is designed for memory-mapped loading and efficient CPU inference via llama.cpp, making it the standard container for quantized LLM deployment outside of pure GPU environments.

Can I implement quantization without retraining my model?

Yes, through Post-Training Quantization (PTQ). The reference implementation demonstrates symmetric and asymmetric PTQ that converts existing FP16 weights to INT8 or INT4 using calibration data (128-1024 examples) without gradient updates. This differs from Quantization-Aware Training (QAT), which simulates quantization during forward passes throughout training or fine-tuning to achieve superior INT4 results.

What metrics should I monitor when evaluating quantized models?

Monitor MSE (mean squared error) and cosine similarity between original and reconstructed weights during development, then validate with perplexity and downstream task benchmarks on a holdout set. The quantization_error function in the curriculum's main.py calculates these automatically, including SNR in decibels to quantify signal preservation across quantization noise.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →