# How to Run Performance Benchmarks with ds4-bench: A Complete Guide

> Master performance benchmarking with ds4-bench. This guide details measuring DeepSeek V4 and GLM 5.2 inference throughput across context sizes. Optimize your models now.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: performance
- Published: 2026-08-08

---

**The `ds4-bench` tool measures inference throughput for DeepSeek V4 and GLM 5.2 models by isolating prefill and generation phases across configurable context sizes.**

Running performance benchmarks with **ds4-bench** allows developers to quantify the latency and throughput characteristics of the **ds4** inference engine. This C-based benchmarking utility, maintained in the `antirez/ds4` repository, provides granular control over backend selection, context scaling, and KV cache management to produce reproducible performance metrics.

## Core Architecture Components

### Engine and Backend Selection

The benchmark relies on the **`ds4_engine`** structure defined in [[`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c)](https://github.com/antirez/ds4/blob/main/ds4.c#L5-L12) to load GGUF models and select computational shapes. The engine initializes with one of three shapes: `DS4_SHAPE_FLASH`, `DS4_SHAPE_PRO`, or `DS4_SHAPE_GLM52`, stored in the global `g_ds4_shape` variable.

Backend selection occurs through the `default_backend()` helper in [[`ds4_bench.c`](https://github.com/antirez/ds4/blob/main/ds4_bench.c)](https://github.com/antirez/ds4/blob/main/ds4_bench.c#L51-L58). The system defaults to **Metal** on macOS, **CUDA** on Linux systems with NVIDIA GPUs, and **CPU** when GPU support is disabled or unavailable. The backend enum in [[`ds4.h`](https://github.com/antirez/ds4/blob/main/ds4.h)](https://github.com/antirez/ds4/blob/main/ds4.h#L19-L23) defines these execution targets.

### Session Management

Live inference state resides in the **`ds4_session`** structure declared in [[`ds4.h`](https://github.com/antirez/ds4/blob/main/ds4.h)](https://github.com/antirez/ds4/blob/main/ds4.h#L30-L38). Sessions encapsulate the KV cache, logits buffers, and the mutable inference timeline. The benchmark creates sessions via `ds4_session_create()` and manipulates them through two primary operations:

- **`ds4_session_sync()`**: Processes prompt tokens during the prefill phase.
- **`ds4_session_eval()`**: Performs autoregressive token generation.

Tokenization support in [[`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c)](https://github.com/antirez/ds4/blob/main/ds4.c#L61-L66) provides `ds4_tokenize_text()` for plain text and `ds4_encode_chat_prompt()` for chat-style prompts, converting input into `ds4_tokens` arrays suitable for session consumption.

### The Benchmark Loop

The central measurement logic resides in `main()` within [[`ds4_bench.c`](https://github.com/antirez/ds4/blob/main/ds4_bench.c)](https://github.com/antirez/ds4/blob/main/ds4_bench.c#L54-L84). The tool iterates through **context frontiers** using `next_frontier()` ([lines 71-84](https://github.com/antirez/ds4/blob/main/ds4_bench.c#L71-L84)), which computes subsequent token counts based on linear (`--step-incr`) or exponential (`--step-mul`) growth patterns.

For each frontier, the loop executes:

1. **Prefill measurement**: Times `ds4_session_sync()` using `bench_now_sec()`.
2. **Optional snapshot**: Saves the KV cache via `ds4_session_save_snapshot()` if the payload size permits.
3. **Generation measurement**: Runs `ds4_session_eval()` for `--gen-tokens` iterations, recording first-token and steady-state latency.
4. **State restoration**: Reloads via `ds4_session_load_snapshot()` or replays the prefix to maintain measurement isolation.

### Snapshot Handling and Output

Snapshot functionality, implemented around [lines 730-770 of ds4_bench.c](https://github.com/antirez/ds4/blob/main/ds4_bench.c#L730-L770), creates compact binary representations of the session state. This mechanism is automatically disabled for very large payloads unless the `DS4_BENCH_FORCE_SNAPSHOT` environment variable is set.

The tool emits structured output through two channels:

- **CSV reporting**: Writes headers at [line 71](https://github.com/antirez/ds4/blob/main/ds4_bench.c#L71-L73) containing `ctx_tokens`, `prefill_tps`, and `gen_tps` metrics.
- **JSON logits**: The `write_frontier_logits_json()` function ([lines 97-110](https://github.com/antirez/ds4/blob/main/ds4_bench.c#L97-L110)) dumps full vocabulary logits when `--dump-frontier-logits-dir` is specified.

## Step-by-Step Benchmark Execution

When running performance benchmarks with ds4-bench, the execution flow follows this pipeline:

1. **Configuration parsing**: The `bench_config` struct ([lines 28-62](https://github.com/antirez/ds4/blob/main/ds4_bench.c#L28-L62)) gathers CLI options including model path, backend type, context boundaries, and SSD-streaming flags via `parse_options()`.

2. **Engine initialization**: Calls `ds4_engine_open()` or `ds4_engine_create_with_gpu_config()` ([lines 103-118](https://github.com/antirez/ds4/blob/main/ds4_bench.c#L103-L118)) with the parsed configuration.

3. **Prompt ingestion**: Loads either `--prompt-file` (plain text) or `--chat-prompt-file` (JSON), tokenizing through the appropriate ds4 API.

4. **Frontier iteration**: Starting at `--ctx-start`, incrementally increases context size until reaching `--ctx-max`.

5. **Isolated timing**: Measures only the newest prefill interval and generation forward pass, explicitly excluding snapshot save/restore overhead from throughput calculations.

## Practical Benchmark Examples

### Basic CPU Benchmark

Run a baseline measurement using the CPU backend with linear context scaling:

```bash
./ds4-bench \
  -m ds4flash.gguf \
  --prompt-file prompt.txt \
  --ctx-start 2048 \
  --ctx-max 32768 \
  --step-incr 2048 \
  --gen-tokens 128 \
  --backend cpu \
  --csv bench_results.csv

```

This configuration tests the Flash model variant from 2,048 to 32,768 tokens in 2,048-token increments, generating 128 tokens at each step to measure decode throughput.

### Metal-Accelerated Benchmark on macOS

Leverage Apple Silicon GPU acceleration with exponential context growth:

```bash
./ds4-bench \
  -m ds4pro.gguf \
  --prompt-file prompt.txt \
  --ctx-start 4096 \
  --ctx-max 65536 \
  --step-mul 1.5 \
  --gen-tokens 64 \
  --backend metal \
  --prefill-chunk 4096 \
  --csv metal_bench.csv

```

The `--step-mul 1.5` parameter enables exponential scaling, while `--prefill-chunk` optimizes the batch size for Metal shader dispatch defined in the [`metal/`](https://github.com/antirez/ds4/tree/main/metal) directory.

### CUDA with SSD Streaming

Evaluate large-context performance using SSD-backed KV caching:

```bash
./ds4-bench \
  -m ds4pro.gguf \
  --prompt-file prompt.txt \
  --ctx-start 8192 \
  --ctx-max 131072 \
  --step-incr 8192 \
  --gen-tokens 32 \
  --backend cuda \
  --ssd-streaming \
  --ssd-streaming-cache-experts 64 \
  --expert-profile ./expert_profile.json \
  --csv cuda_streaming.csv

```

This example activates SSD streaming for contexts up to 131,072 tokens, limiting in-memory expert caching to 64 entries and recording per-layer expert usage patterns for cache-policy optimization.

### Dumping Frontier Logits for Analysis

Capture raw model outputs for quantization debugging or probability analysis:

```bash
mkdir logits_dir
./ds4-bench \
  -m ds4flash.gguf \
  --prompt-file prompt.txt \
  --ctx-start 2048 \
  --ctx-max 16384 \
  --step-incr 2048 \
  --gen-tokens 0 \
  --dump-frontier-logits-dir logits_dir \
  --csv logits_summary.csv

```

Setting `--gen-tokens 0` isolates prefill performance exclusively, while generating [`frontier_000001.logits.json`](https://github.com/antirez/ds4/blob/main/frontier_000001.logits.json) files containing complete vocabulary logits for each context size.

## Key Implementation Files

Understanding these source files is essential for advanced benchmark customization:

- **[[`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c)](https://github.com/antirez/ds4/blob/main/ds4.c)**: Implements `ds4_engine` initialization, shape selection (`DS4_SHAPE_FLASH`, `DS4_SHAPE_PRO`, `DS4_SHAPE_GLM52`), and tokenization APIs.
- **[[`ds4_bench.c`](https://github.com/antirez/ds4/blob/main/ds4_bench.c)](https://github.com/antirez/ds4/blob/main/ds4_bench.c)**: Contains the complete benchmark implementation including CLI parsing (`bench_config`), the measurement loop (`next_frontier()`), and output formatters.
- **[[`ds4.h`](https://github.com/antirez/ds4/blob/main/ds4.h)](https://github.com/antirez/ds4/blob/main/ds4.h)**: Defines the public API surface including `ds4_engine_options` ([lines 27-73](https://github.com/antirez/ds4/blob/main/ds4.h#L27-L73)), backend enums, and `ds4_session` lifecycle functions.
- **[[`ds4_gpu_mgpu.h`](https://github.com/antirez/ds4/blob/main/ds4_gpu_mgpu.h)](https://github.com/antirez/ds4/blob/main/ds4_gpu_mgpu.h)**: Provides multi-GPU configuration structures when using `--gpu-devices` for tensor parallelism.

## Summary

Running performance benchmarks with ds4-bench provides hardware-specific throughput metrics for DeepSeek and GLM model inference. Key takeaways include:

- The tool isolates **prefill** (prompt processing) and **generation** (token decoding) phases into separate measurable windows.
- **Context frontiers** enable systematic scaling studies from minimal to maximum supported context lengths using linear or exponential stepping.
- **Snapshot mechanisms** eliminate repeated prefill costs between measurement iterations without polluting timing data.
- Backend selection (Metal, CUDA, CPU) is automatic but overridable via `--backend`, with each path utilizing device-specific kernels.
- Structured CSV and JSON outputs facilitate automated analysis of throughput trends and model behavior across context sizes.

## Frequently Asked Questions

### What is ds4-bench and what does it measure?

**ds4-bench** is the official benchmarking utility for the `antirez/ds4` inference engine. It measures **tokens-per-second throughput** separately for the prefill phase (processing input prompts) and the generation phase (decoding output tokens), enabling precise performance characterization across different hardware backends and context lengths.

### How does ds4-bench handle large context windows efficiently?

The tool implements **session snapshots** that serialize the KV cache to a compact binary format via `ds4_session_save_snapshot()`. When iterating through context frontiers, the benchmark restores this state using `ds4_session_load_snapshot()` rather than re-processing the entire prompt prefix, significantly reducing benchmark duration for large context studies while maintaining measurement accuracy.

### What is the difference between `--step-incr` and `--step-mul`?

**`--step-incr`** specifies a fixed token count added to the context size at each frontier (linear growth), suitable for uniform sampling across context ranges. **`--step-mul`** applies a multiplicative factor (e.g., 1.5 or 2.0) to generate exponentially spaced frontiers, which is more efficient for identifying performance inflection points or testing very large maximum contexts with fewer intermediate steps.

### Can I benchmark custom prompts or chat conversations?

Yes. The tool accepts **`--prompt-file`** for plain text inputs or **`--chat-prompt-file`** for JSON-formatted chat conversations. Both paths utilize the ds4 tokenization pipeline—`ds4_tokenize_text()` for raw text and `ds4_encode_chat_prompt()` for structured chat templates—ensuring the benchmark reflects real-world prompt complexity and token distributions.