How to Run Performance Benchmarks with ds4-bench: A Complete Guide

The ds4-bench tool measures inference throughput for DeepSeek V4 and GLM 5.2 models by isolating prefill and generation phases across configurable context sizes.

Running performance benchmarks with ds4-bench allows developers to quantify the latency and throughput characteristics of the ds4 inference engine. This C-based benchmarking utility, maintained in the antirez/ds4 repository, provides granular control over backend selection, context scaling, and KV cache management to produce reproducible performance metrics.

Core Architecture Components

Engine and Backend Selection

The benchmark relies on the ds4_engine structure defined in [ds4.c](https://github.com/antirez/ds4/blob/main/ds4.c#L5-L12) to load GGUF models and select computational shapes. The engine initializes with one of three shapes: DS4_SHAPE_FLASH, DS4_SHAPE_PRO, or DS4_SHAPE_GLM52, stored in the global g_ds4_shape variable.

Backend selection occurs through the default_backend() helper in [ds4_bench.c](https://github.com/antirez/ds4/blob/main/ds4_bench.c#L51-L58). The system defaults to Metal on macOS, CUDA on Linux systems with NVIDIA GPUs, and CPU when GPU support is disabled or unavailable. The backend enum in [ds4.h](https://github.com/antirez/ds4/blob/main/ds4.h#L19-L23) defines these execution targets.

Session Management

Live inference state resides in the ds4_session structure declared in [ds4.h](https://github.com/antirez/ds4/blob/main/ds4.h#L30-L38). Sessions encapsulate the KV cache, logits buffers, and the mutable inference timeline. The benchmark creates sessions via ds4_session_create() and manipulates them through two primary operations:

  • ds4_session_sync(): Processes prompt tokens during the prefill phase.
  • ds4_session_eval(): Performs autoregressive token generation.

Tokenization support in [ds4.c](https://github.com/antirez/ds4/blob/main/ds4.c#L61-L66) provides ds4_tokenize_text() for plain text and ds4_encode_chat_prompt() for chat-style prompts, converting input into ds4_tokens arrays suitable for session consumption.

The Benchmark Loop

The central measurement logic resides in main() within [ds4_bench.c](https://github.com/antirez/ds4/blob/main/ds4_bench.c#L54-L84). The tool iterates through context frontiers using next_frontier() (lines 71-84), which computes subsequent token counts based on linear (--step-incr) or exponential (--step-mul) growth patterns.

For each frontier, the loop executes:

  1. Prefill measurement: Times ds4_session_sync() using bench_now_sec().
  2. Optional snapshot: Saves the KV cache via ds4_session_save_snapshot() if the payload size permits.
  3. Generation measurement: Runs ds4_session_eval() for --gen-tokens iterations, recording first-token and steady-state latency.
  4. State restoration: Reloads via ds4_session_load_snapshot() or replays the prefix to maintain measurement isolation.

Snapshot Handling and Output

Snapshot functionality, implemented around lines 730-770 of ds4_bench.c, creates compact binary representations of the session state. This mechanism is automatically disabled for very large payloads unless the DS4_BENCH_FORCE_SNAPSHOT environment variable is set.

The tool emits structured output through two channels:

  • CSV reporting: Writes headers at line 71 containing ctx_tokens, prefill_tps, and gen_tps metrics.
  • JSON logits: The write_frontier_logits_json() function (lines 97-110) dumps full vocabulary logits when --dump-frontier-logits-dir is specified.

Step-by-Step Benchmark Execution

When running performance benchmarks with ds4-bench, the execution flow follows this pipeline:

  1. Configuration parsing: The bench_config struct (lines 28-62) gathers CLI options including model path, backend type, context boundaries, and SSD-streaming flags via parse_options().

  2. Engine initialization: Calls ds4_engine_open() or ds4_engine_create_with_gpu_config() (lines 103-118) with the parsed configuration.

  3. Prompt ingestion: Loads either --prompt-file (plain text) or --chat-prompt-file (JSON), tokenizing through the appropriate ds4 API.

  4. Frontier iteration: Starting at --ctx-start, incrementally increases context size until reaching --ctx-max.

  5. Isolated timing: Measures only the newest prefill interval and generation forward pass, explicitly excluding snapshot save/restore overhead from throughput calculations.

Practical Benchmark Examples

Basic CPU Benchmark

Run a baseline measurement using the CPU backend with linear context scaling:

./ds4-bench \
  -m ds4flash.gguf \
  --prompt-file prompt.txt \
  --ctx-start 2048 \
  --ctx-max 32768 \
  --step-incr 2048 \
  --gen-tokens 128 \
  --backend cpu \
  --csv bench_results.csv

This configuration tests the Flash model variant from 2,048 to 32,768 tokens in 2,048-token increments, generating 128 tokens at each step to measure decode throughput.

Metal-Accelerated Benchmark on macOS

Leverage Apple Silicon GPU acceleration with exponential context growth:

./ds4-bench \
  -m ds4pro.gguf \
  --prompt-file prompt.txt \
  --ctx-start 4096 \
  --ctx-max 65536 \
  --step-mul 1.5 \
  --gen-tokens 64 \
  --backend metal \
  --prefill-chunk 4096 \
  --csv metal_bench.csv

The --step-mul 1.5 parameter enables exponential scaling, while --prefill-chunk optimizes the batch size for Metal shader dispatch defined in the metal/ directory.

CUDA with SSD Streaming

Evaluate large-context performance using SSD-backed KV caching:

./ds4-bench \
  -m ds4pro.gguf \
  --prompt-file prompt.txt \
  --ctx-start 8192 \
  --ctx-max 131072 \
  --step-incr 8192 \
  --gen-tokens 32 \
  --backend cuda \
  --ssd-streaming \
  --ssd-streaming-cache-experts 64 \
  --expert-profile ./expert_profile.json \
  --csv cuda_streaming.csv

This example activates SSD streaming for contexts up to 131,072 tokens, limiting in-memory expert caching to 64 entries and recording per-layer expert usage patterns for cache-policy optimization.

Dumping Frontier Logits for Analysis

Capture raw model outputs for quantization debugging or probability analysis:

mkdir logits_dir
./ds4-bench \
  -m ds4flash.gguf \
  --prompt-file prompt.txt \
  --ctx-start 2048 \
  --ctx-max 16384 \
  --step-incr 2048 \
  --gen-tokens 0 \
  --dump-frontier-logits-dir logits_dir \
  --csv logits_summary.csv

Setting --gen-tokens 0 isolates prefill performance exclusively, while generating frontier_000001.logits.json files containing complete vocabulary logits for each context size.

Key Implementation Files

Understanding these source files is essential for advanced benchmark customization:

Summary

Running performance benchmarks with ds4-bench provides hardware-specific throughput metrics for DeepSeek and GLM model inference. Key takeaways include:

  • The tool isolates prefill (prompt processing) and generation (token decoding) phases into separate measurable windows.
  • Context frontiers enable systematic scaling studies from minimal to maximum supported context lengths using linear or exponential stepping.
  • Snapshot mechanisms eliminate repeated prefill costs between measurement iterations without polluting timing data.
  • Backend selection (Metal, CUDA, CPU) is automatic but overridable via --backend, with each path utilizing device-specific kernels.
  • Structured CSV and JSON outputs facilitate automated analysis of throughput trends and model behavior across context sizes.

Frequently Asked Questions

What is ds4-bench and what does it measure?

ds4-bench is the official benchmarking utility for the antirez/ds4 inference engine. It measures tokens-per-second throughput separately for the prefill phase (processing input prompts) and the generation phase (decoding output tokens), enabling precise performance characterization across different hardware backends and context lengths.

How does ds4-bench handle large context windows efficiently?

The tool implements session snapshots that serialize the KV cache to a compact binary format via ds4_session_save_snapshot(). When iterating through context frontiers, the benchmark restores this state using ds4_session_load_snapshot() rather than re-processing the entire prompt prefix, significantly reducing benchmark duration for large context studies while maintaining measurement accuracy.

What is the difference between --step-incr and --step-mul?

--step-incr specifies a fixed token count added to the context size at each frontier (linear growth), suitable for uniform sampling across context ranges. --step-mul applies a multiplicative factor (e.g., 1.5 or 2.0) to generate exponentially spaced frontiers, which is more efficient for identifying performance inflection points or testing very large maximum contexts with fewer intermediate steps.

Can I benchmark custom prompts or chat conversations?

Yes. The tool accepts --prompt-file for plain text inputs or --chat-prompt-file for JSON-formatted chat conversations. Both paths utilize the ds4 tokenization pipeline—ds4_tokenize_text() for raw text and ds4_encode_chat_prompt() for structured chat templates—ensuring the benchmark reflects real-world prompt complexity and token distributions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →