Memory Reduction from Quantizing a 70B Parameter Model to INT4: A 4× Compression Guide
Quantizing a 70B parameter model from FP16 to INT4 reduces memory usage by approximately 4×, shrinking requirements from ~140 GB to ~35 GB.
The HenryNdubuaku/maths-cs-ai-compendium repository documents the precise arithmetic behind this compression, demonstrating how 4-bit quantization enables large language models to fit on single high-end GPUs. This guide breaks down the source calculations and provides runnable code to verify the memory savings.
The Mathematics of 70B Parameter Quantization
FP16 Baseline Memory Requirements
In half-precision floating-point format (FP16), each parameter consumes 16 bits (2 bytes) of storage. For a 70-billion-parameter model, the raw weight storage calculates as:
70,000,000,000 parameters × 16 bits = 1,120,000,000,000 bits
1,120,000,000,000 bits ÷ 8 = 140,000,000,000 bytes
140,000,000,000 bytes ÷ (1024³) ≈ 130.4 GB (using binary) or 140 GB (using decimal/commonly cited)
According to chapter 17: AI inference/01. quantisation.md, this places the FP16 model at approximately 140 GB of memory required for the weight tensors alone.
INT4 Compression Calculation
INT4 quantization compresses each weight to 4 bits (0.5 bytes), yielding a direct 4:1 compression ratio:
70,000,000,000 parameters × 4 bits = 280,000,000,000 bits
280,000,000,000 bits ÷ 8 = 35,000,000,000 bytes ≈ 35 GB
As implemented in the repository's analysis, this reduces the model footprint from 140 GB to 35 GB, achieving the 4× memory reduction characteristic of FP16-to-INT4 quantization.
Source Evidence from maths-cs-ai-compendium
The memory reduction figures are explicitly documented in the repository's inference optimization documentation:
-
chapter 17: AI inference/01. quantisation.md: Contains the foundational calculation showing that "A 70B parameter model in float16 requires 140 GB of memory... Quantise to INT4 and it fits in 35 GB." -
chapter 17: AI inference/05. scaling and deployment.md: Details the practical deployment implications, noting that "INT4 vs FP16 is 4x less memory" and enabling single-GPU inference for models previously requiring multi-GPU setups. -
chapter 16: SIMD and GPU programming/06. RISC‑V and embedded systems.md: Provides context on INT4 usage in edge computing, validating the low-bit quantization approach across hardware architectures.
Practical Implementation: Verifying Memory Reduction
Use the following Python snippet to calculate the exact memory requirements for any parameter count and bit width:
def calculate_memory_gb(num_params: int, bits_per_param: int) -> float:
"""Calculate memory footprint in gigabytes."""
bytes_total = (num_params * bits_per_param) / 8
return bytes_total / (1024 ** 3)
# 70B parameter model calculations
params = 70_000_000_000
fp16_memory = calculate_memory_gb(params, 16) # ~130.4 GiB
int4_memory = calculate_memory_gb(params, 4) # ~32.6 GiB
print(f"FP16 memory: {fp16_memory:.1f} GB")
print(f"INT4 memory: {int4_memory:.1f} GB")
print(f"Reduction ratio: {fp16_memory / int4_memory:.1f}x")
Loading 70B Models with INT4 Quantization
The following implementation uses the transformers library with bitsandbytes integration to load a 70B parameter model with 4-bit quantization, achieving the documented memory reduction:
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
import torch
model_id = "meta-llama/Llama-2-70b-chat-hf"
# Configure INT4 quantization
quantization_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4", # Normalized float4 for better accuracy
bnb_4bit_compute_dtype=torch.float16, # Compute in FP16 while storing in INT4
bnb_4bit_use_double_quant=True, # Nested quantization for additional memory savings
)
# Load model with ~35GB memory footprint instead of ~140GB
model = AutoModelForCausalLM.from_pretrained(
model_id,
quantization_config=quantization_config,
device_map="auto", # Automatically distributes across available GPUs
torch_dtype=torch.float16,
)
tokenizer = AutoTokenizer.from_pretrained(model_id)
print("Model loaded with INT4 quantization - approximately 35GB VRAM usage")
Deployment Implications of 4× Memory Reduction
Single GPU Feasibility: The reduction from 140 GB to 35 GB enables deployment on single high-end GPUs (such as the NVIDIA A100 40GB or 80GB), eliminating the need for multi-GPU inference pipelines.
Bandwidth and Throughput: As noted in 05. scaling and deployment.md, the 4× compression reduces memory bandwidth requirements proportionally, allowing faster token generation during inference.
Cost Optimization: The memory reduction translates directly to infrastructure cost savings—models that previously required 4-8 GPUs can now run on a single device, reducing both hardware acquisition and operational expenses.
Summary
- 4× memory reduction: Quantizing from FP16 (16-bit) to INT4 (4-bit) compresses a 70B parameter model from approximately 140 GB to 35 GB.
- Binary arithmetic: The reduction factor is derived from the bit-width ratio (16 ÷ 4 = 4).
- Single GPU deployment: The 35 GB footprint fits within high-end consumer and enterprise GPUs, democratizing access to large models.
- Source documentation: Calculations are verified in
chapter 17: AI inference/01. quantisation.mdand deployment strategies in05. scaling and deployment.md.
Frequently Asked Questions
What is the exact memory reduction factor when quantizing from FP16 to INT4?
The reduction factor is exactly 4× (or 75% memory savings), calculated by dividing the FP16 bit-width (16 bits) by the INT4 bit-width (4 bits). For a 70B parameter model, this reduces storage from approximately 140 GB to 35 GB.
How does INT4 quantization affect model accuracy?
Modern INT4 quantization techniques (such as 4-bit Normalized Float or NF4) typically retain 95-99% of the original FP16 model accuracy. The bitsandbytes library implements quantization-aware scaling factors that minimize precision loss during the FP16-to-INT4 conversion.
Can a 70B INT4 model fit on a single consumer GPU?
Yes, with approximately 35 GB of VRAM required for the weights, the quantized model fits on single high-end GPUs such as the NVIDIA RTX 4090 (24GB) with additional offloading techniques, or more comfortably on datacenter GPUs like the A100 (40GB/80GB) and H100 (80GB). This is infeasible with the 140 GB FP16 version.
Where is the memory calculation documented in the HenryNdubuaku/maths-cs-ai-compendium repository?
The specific calculation "A 70B parameter model in float16 requires 140 GB of memory... Quantise to INT4 and it fits in 35 GB" is documented in chapter 17: AI inference/01. quantisation.md, with practical deployment implications covered in 05. scaling and deployment.md.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →