Memory Requirements for Different Model Quantizations in DS4
DS4 supports GGUF-style quantization formats ranging from 2-bit to 16-bit precision, with memory usage scaling linearly from approximately 0.25 KB to 2 KB per 1,024 parameters, and provides the ds4_context_memory_estimate() function to calculate total consumption including KV-cache overhead.
The antirez/ds4 repository implements efficient inference for large language models using GGUF-style quantization formats that trade numerical precision for reduced memory footprint. Understanding the memory requirements for different model quantizations is essential for deploying models on resource-constrained hardware or multi-GPU configurations. DS4 provides specific APIs and lookup tables in gguf-tools/quants.c that map quantization types to their exact bit-per-weight ratios, enabling precise memory budgeting before model loading.
Quantization Formats and Memory Footprint
DS4 implements several quantization strategies defined in gguf-tools/quants.c, each offering a different trade-off between model fidelity and memory consumption. The memory occupied by model weights scales directly with the bits-per-weight ratio.
| Quantization | Bits per weight | Approx. Bytes per 1K parameters |
|---|---|---|
| Q8_0 | 8 | ~1 KB |
| Q4_K | 4 | ~0.5 KB |
| Q2_K | 2 | ~0.25 KB |
| IQ2_XXS | ~2.5 | ~0.31 KB |
| FP16 | 16 | ~2 KB |
| BF16 | 16 | ~2 KB |
| FP8 E4M3 | 8 | ~1 KB |
Values represent storage for 1,024 parameters. Actual inference memory includes additional overhead for the KV-cache, which stores intermediate activations for each token in the context window.
Memory Calculation Formula
The total memory required for a DS4 model can be estimated using the formula implemented in ds4_context.c:
memory ≈ (model_parameters × bits_per_weight / 8) + KV-cache overhead
The KV-cache overhead depends on the context length, number of transformer layers, attention heads, and the selected quantization type. Notably, the KV-cache uses the same quantization format as the model weights, meaning a Q4_K model will store cached keys and values using 4-bit quantization. The ds4_context_memory_estimate() function aggregates these factors to return a total byte estimate, accounting for device-specific alignment and padding requirements.
Programmatic Memory Estimation
DS4 exposes C APIs to calculate memory requirements before allocating GPU or CPU buffers. These functions are utilized throughout the test suite, particularly in tests/test_engine_mgpu_placement.c, to validate that requested contexts fit within available device memory.
Estimating Context Memory with ds4_context_memory_estimate
The ds4_context_memory_estimate() function, defined in ds4_context.c, computes total memory consumption given a device type, desired context length, and quantization bit depth.
#include "ds4.h"
#include "ds4_context.h"
int main(void) {
/* 7-B model ≈ 7,000,000,000 parameters */
const int64_t model_params = 7LL * 1000 * 1000 * 1000;
/* Q4_K quantization = 4 bits per weight */
const int quant_bits = 4;
/* Desired context length in tokens */
const int64_t ctx_len = 4096;
/* Estimate memory for CUDA device */
size_t mem_bytes = ds4_context_memory_estimate(DS4_DEVICE_CUDA,
ctx_len,
quant_bits);
printf("Estimated memory: %.2f GiB\n",
mem_bytes / (1024.0 * 1024 * 1024));
return 0;
}
Calculating Per-Token Memory Overhead
For custom memory planners, DS4 provides bit-depth constants in gguf-tools/quants.c that allow manual calculation of storage requirements per token block.
#include <stdio.h>
static void print_mem_per_token(int bits) {
double bytes_per_1k = (bits / 8.0) * 1024.0 / 1024.0; /* Convert to KB */
printf(" %2d-bit → %.3f KB per 1K tokens\n", bits, bytes_per_1k);
}
int main(void) {
printf("Per-token memory usage by quantization:\n");
print_mem_per_token(8); /* Q8_0 */
print_mem_per_token(4); /* Q4_K */
print_mem_per_token(2); /* Q2_K */
print_mem_per_token(3); /* IQ2_XXS approximated */
return 0;
}
Key Source Files
The quantization system and memory estimation logic are implemented across the following source files in antirez/ds4:
gguf-tools/quants.c: Implements low-level quantization routines (ds4q_quantize_*) and defines bit-per-weight constants for each format.gguf-tools/deepseek4-quantize.c: Provides the high-level quantization pipeline and validation checks viads4q_can_quantize().ds4_context.c: Containsds4_context_memory_estimate(), which aggregates model parameters, quantization bits, and KV-cache requirements.tests/test_engine_mgpu_placement.c: Uses memory estimation functions to validate context placement across multiple GPUs.
Summary
- DS4 supports Q2_K through FP16 quantization formats, with memory usage ranging from ~0.25 KB to ~2 KB per 1,024 parameters.
- Memory scales linearly with bit depth: Q4_K uses 50% less memory than Q8_0, and 75% less than FP16.
- The
ds4_context_memory_estimate()function inds4_context.cprovides device-aware total memory calculation, including KV-cache overhead. - KV-cache quantization matches the weight quantization type, affecting total memory proportionally with context length.
- Reference
gguf-tools/quants.cfor the authoritative definitions of bit-per-weight ratios used in all calculations.
Frequently Asked Questions
What is the most memory-efficient quantization format available in DS4?
Q2_K provides the smallest footprint at approximately 0.25 KB per 1,024 parameters. For scenarios requiring slightly higher precision, IQ2_XXS offers a compromise at ~0.31 KB per 1,024 parameters by utilizing extra scaling factors to reduce quantization error while maintaining near-2-bit storage efficiency.
How does DS4 calculate total GPU memory requirements for inference?
The repository uses the ds4_context_memory_estimate() function implemented in ds4_context.c. This function accepts the target device (DS4_DEVICE_CUDA or DS4_DEVICE_CPU), desired context length, and quantization bit depth, then returns the total bytes required for both model weights and the KV-cache.
Does the KV-cache use the same quantization precision as the model weights?
Yes. According to the implementation in gguf-tools/quants.c and the estimation logic in ds4_context.c, the KV-cache is stored using the identical quantization type selected for the model weights. Therefore, choosing Q4_K quantization reduces both model storage and context cache memory by 50% compared to Q8_0.
What is the practical memory difference between Q8_0 and FP16 quantization?
Q8_0 consumes exactly 50% less memory than FP16. Specifically, Q8_0 requires approximately 1 KB of storage per 1,024 parameters (8 bits per weight), while FP16 requires 2 KB per 1,024 parameters (16 bits per weight). This translates to a ~7 GB savings for a 7-billion parameter model when using Q8_0 instead of FP16.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →