Estimating Memory Requirements for Different Context Sizes in ds4: A Complete Guide
Use the ds4_context_memory_estimate() API to calculate GPU/CPU memory consumption for any token context length before allocating model state.
The ds4 library by antirez provides explicit memory budgeting tools that let you predict RAM and VRAM usage across different context window sizes before loading a model. Understanding how to estimate these requirements prevents out-of-memory errors during inference and helps you right-size your hardware for specific token lengths. This guide walks through the architecture, API usage, and source code implementation for memory estimation in ds4.
How ds4 Calculates Memory for Context Windows
The estimation engine follows a three-layer architecture that separates the public API from backend-specific calculations.
The Entry Point: ds4_context_memory_estimate
The primary interface is ds4_context_memory_estimate(), defined in ds4.c. This function takes a backend identifier and a context size in tokens, then returns a populated ds4_context_memory struct.
Internally, it delegates immediately to the prefill-aware variant with a zero-length prefill chunk:
ds4_context_memory ds4_context_memory_estimate(ds4_backend backend,
int ctx_size) {
return ds4_context_memory_estimate_with_prefill_mode(backend, ctx_size, 0);
}
This design ensures that all memory calculations account for both the base context and any potential prefilling overhead.
Prefill-Aware Estimation Logic
The core logic resides in ds4_context_memory_estimate_with_prefill_mode(). This function branches based on the selected backend:
- GLM-DSA graph backend: Routes to GLM-specific graph memory calculators.
- Metal backend: Uses Metal-specific capacity functions and compression ratio lookups.
The function aggregates raw KV cache bytes, compressed cache bytes, and scratch workspace bytes into the final total.
Backend-Specific Memory Models
Different hardware backends compute memory layouts using distinct algorithms optimized for their execution graphs.
GLM-DSA Graph Backend Calculations
For GLM models, the estimator calculates several capacity tiers:
- Full-attention capacity:
glm_graph_full_attention_cap()determines the maximum tokens that can use full attention. - Compact-cache initialization:
glm_graph_compact_cache_initial_cap()derives the initial compressed cache size from the context length and full-attention capacity. - Workspace sizing:
glm_graph_workspace_bytes_for_cap()sums tensor allocations for embeddings, Q-K-V matrices, and indexer buffers. - Final aggregation:
glm_graph_context_memory_estimate_for_compact_cap_slice()combines raw KV cache, compressed cache, and workspace buffers into the total byte count.
Metal Backend Memory Layout
The Metal (non-graph) backend uses a different partitioning strategy:
- Prefill capacity:
metal_graph_prefill_cap_for_prompt()calculates tokens available for prefix caching based on prompt length. - Raw capacity:
metal_graph_raw_cap_for_context()determines the uncompressed KV cache size for the full context window. - Compression ratios: The estimator queries
ds4_layer_compress_ratiofor each layer, using the smallest non-zero ratio to setcomp_cap(compressed capacity). - Buffer summation: Raw bytes, compressed bytes (including indexer heads for 4:1 ratios), and scratch buffers (prefill-stage plus attention-stage) sum to form
total_bytes.
Interpreting the ds4_context_memory Structure
The ds4_context_memory struct breaks down memory usage into discrete components:
| Field | Description |
|---|---|
raw_cap |
Token capacity for the uncompressed KV cache |
prefill_cap |
Tokens that can be prefilled without reallocation |
comp_cap |
Token capacity for the compressed cache (if applicable) |
raw_bytes |
Byte size of the raw KV cache |
compressed_bytes |
Byte size of the compressed cache |
scratch_bytes |
Workspace and scratch buffer allocation |
total_bytes |
Aggregate memory footprint for the requested context size |
Use total_bytes to validate that your GPU or CPU memory budget can accommodate the desired context length before initializing the model.
Practical Code Examples
The following C program demonstrates how to call the estimation API for various context sizes:
#include "ds4.h"
#include <stdio.h>
int main(void) {
/* Estimate memory for a 4096-token context on the CUDA backend */
ds4_context_memory mem = ds4_context_memory_estimate(DS4_BACKEND_CUDA, 4096);
printf("Context size: %d tokens\n", 4096);
printf("Raw KV cache : %llu bytes\n", (unsigned long long)mem.raw_bytes);
printf("Compressed cache : %llu bytes\n", (unsigned long long)mem.compressed_bytes);
printf("Scratch/workspace : %llu bytes\n", (unsigned long long)mem.scratch_bytes);
printf("TOTAL : %llu bytes (≈ %.2f MiB)\n",
(unsigned long long)mem.total_bytes,
mem.total_bytes / (1024.0 * 1024.0));
return 0;
}
Typical memory scaling for a standard model configuration looks like this:
| Context (tokens) | raw_bytes | compressed_bytes | scratch_bytes | total_bytes |
|---|---|---|---|---|
| 512 | ~4 MiB | 0 MiB | ~2 MiB | ~6 MiB |
| 2048 | ~16 MiB | ~1 MiB | ~6 MiB | ~23 MiB |
| 4096 | ~32 MiB | ~2 MiB | ~12 MiB | ~46 MiB |
| 8192 | ~64 MiB | ~4 MiB | ~24 MiB | ~92 MiB |
Exact values vary by model family and per-layer compression ratios configured in the backend.
Key Source Files and Functions
Understanding the codebase structure helps when modifying estimation logic or debugging memory issues:
- [
ds4.c](https://github.com/antirez/ds4/blob/main/ds4.c) – Contains the main implementation ofds4_context_memory_estimate(),ds4_context_memory_estimate_with_prefill_mode(), and all GLM-DSA graph helpers. - [
ds4.h](https://github.com/antirez/ds4/blob/main/ds4.h) – Defines theds4_context_memorystruct, backend enums, and public API signatures. - [
tests/test_engine_mgpu_placement.c](https://github.com/antirez/ds4/blob/main/tests/test_engine_mgpu_placement.c) – Demonstrates real-world usage of the estimation API in the test suite. - [
ds4_cli.c](https://github.com/antirez/ds4/blob/main/ds4_cli.c) – Command-line interface that prints memory estimates for user-specified context sizes. - [
ds4_server.c](https://github.com/antirez/ds4/blob/main/ds4_server.c) – Server implementation that validates memory budgets against estimated requirements before model allocation.
Summary
- Use
ds4_context_memory_estimate()to predict memory consumption for any token context length before loading a model. - The API delegates to backend-specific calculators for GLM-DSA graphs and Metal execution paths.
- Key components include raw KV caches, compressed caches (when compression ratios are configured), and scratch workspaces.
- Always check the
total_bytesfield against your hardware limits to prevent allocation failures.
Frequently Asked Questions
How do I estimate memory for a 4096-token context in ds4?
Call ds4_context_memory_estimate(DS4_BACKEND_CUDA, 4096) or substitute your specific backend enum. The function returns a struct where total_bytes indicates the required GPU/CPU memory. Divide by (1024.0 * 1024.0) to convert to megabytes.
What is the difference between raw_cap and comp_cap in ds4?
raw_cap specifies the token capacity for the uncompressed key-value cache, while comp_cap defines the capacity for the compressed cache when layer compression is enabled. If no compression ratios are set, comp_cap will be zero and compressed_bytes will be null.
Does ds4 account for prefill tokens in memory estimates?
Yes. The internal function ds4_context_memory_estimate_with_prefill_mode() accepts a prefill size parameter. The public API ds4_context_memory_estimate() passes zero by default, but the architecture supports calculating additional scratch memory required for prefilling operations.
Which backend requires more memory: GLM-DSA or Metal?
Memory requirements depend on the specific model configuration and compression settings rather than the backend alone. GLM-DSA graphs typically allocate larger workspace buffers for tensor operations, while Metal backends may reserve significant scratch space for prefill stages. Always use ds4_context_memory_estimate() to compare specific context sizes for your target backend.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →