# How to Benchmark Inference Performance with the ds4-bench Tool

> Benchmark inference performance with ds4-bench. Measure pre-fill latency and greedy decode speed for your models. Optimize your AI applications today.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: how-to-guide
- Published: 2026-08-08

---

**The `ds4-bench` tool is a purpose-built throughput benchmark that walks a fixed prompt through increasing context-size frontiers, measuring pre-fill latency and greedy decode performance using high-resolution monotonic timers.**

The antirez/ds4 repository provides a lightweight inference engine for large language models, shipping with `ds4-bench` for reproducible performance profiling. This command-line utility iterates through context length frontiers to capture both **pre-fill tokens-per-second (TPS)** and **generation TPS** across CPU, Metal, CUDA, and distributed backends. Understanding how to benchmark inference performance with ds4-bench enables accurate capacity planning and hardware optimization for production deployments.

## Architecture and Implementation Details

The benchmark is implemented entirely in [`ds4_bench.c`](https://github.com/antirez/ds4/blob/main/ds4_bench.c), utilizing the public API declared in [`ds4.h`](https://github.com/antirez/ds4/blob/main/ds4.h) and the core engine logic from [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c). The tool orchestrates engine creation, prompt tokenization, and timed evaluation loops to produce standardized CSV reports.

### Command-Line Configuration

Benchmark behavior is controlled through the `parse_options` function (lines 200-272 in [`ds4_bench.c`](https://github.com/antirez/ds4/blob/main/ds4_bench.c)), which populates a `bench_config` struct from flags such as `--model`, `--ctx-start`, `--ctx-max`, `--step-incr`, and `--backend`. Validation helpers like `parse_int` and `parse_double_arg` ensure numeric arguments are sane before the benchmark begins.

Key configuration parameters include:

- `--prompt-file` or `--chat-prompt-file`: Path to the text file containing the fixed prompt used for all frontiers
- `--ctx-start` and `--ctx-max`: Define the range of context lengths to test
- `--step-incr` or `--step-mul`: Arithmetic or multiplicative step size between frontiers
- `--gen-tokens`: Number of tokens to generate during the decode phase (default is typically short, e.g., 32-128)
- `--backend`: Execution backend (`metal`, `cuda`, `cpu`, etc.)

### Engine and Session Initialization

After parsing, the tool constructs a `ds4_engine_options` struct and creates the inference engine via `ds4_engine_open` or `ds4_engine_create_with_gpu_config` (lines 594-618). For GPU benchmarks, arguments parsed by [`ds4_gpu_args.c`](https://github.com/antirez/ds4/blob/main/ds4_gpu_args.c) configure VRAM limits and device selection.

The prompt text is loaded via `read_file` and tokenized using `ds4_tokenize_text` or encoded as a chat prompt via `ds4_encode_chat_prompt`. A `ds4_session` is then allocated with the target context size (`cfg.ctx_alloc`). In distributed mode, the tool waits for route readiness through `wait_distributed_route` (defined in [`ds4_distributed.c`](https://github.com/antirez/ds4/blob/main/ds4_distributed.c)).

### The Benchmark Loop

The core measurement logic resides in the frontier iteration loop (lines 543-615). For each context size frontier, the tool executes:

1. **Pre-fill Phase**: Calls `ds4_session_sync` with a prefix slice of the prompt to measure prompt processing throughput
2. **Log Frontier Logits** (optional): Writes per-frontier logits via `write_frontier_logits_json`
3. **Snapshot Creation**: Optionally saves a session snapshot to avoid replaying the prefix after generation
4. **Generation Phase**: Performs a short greedy decode of `cfg.gen_tokens` tokens using `ds4_session_eval`
5. **Session Restoration**: Reloads the snapshot or re-syncs the prefix to maintain consistent state
6. **CSV Reporting**: Appends metrics for the current frontier

### Timing and Memory Tracking

High-resolution timing is captured via `bench_now_sec`, which uses `clock_gettime(CLOCK_MONOTONIC)` (lines 63-68) to ensure nanosecond-precision measurements for both pre-fill and generation phases. Memory footprint estimation is provided by `log_context_memory` (lines 86-103), which calculates the KV-cache size based on the selected backend and pre-fill chunk size.

## Running ds4-bench: Command Examples

The following invocations demonstrate how to benchmark inference performance across different hardware configurations and modes.

### Basic Metal Benchmark on Apple Silicon

```bash
./ds4-bench \
  --model ds4flash.gguf \
  --prompt-file my_prompt.txt \
  --ctx-start 2048 \
  --ctx-max 32768 \
  --step-incr 2048 \
  --backend metal \
  --csv bench_results.csv

```

This command iterates from 2048 to 32768 tokens in 2048-token steps, outputting pre-fill and generation TPS to `bench_results.csv`.

### Multi-GPU CUDA Testing

```bash
./ds4-bench \
  --model ds4flash.gguf \
  --prompt-file my_prompt.txt \
  --gpu-devices 0,1 \
  --gpu-vram auto \
  --backend cuda \
  --step-mul 1.5 \
  --gen-tokens 64 \
  --show-output

```

Here, the benchmark utilizes multiple GPU devices with automatic VRAM detection and multiplicative step sizing (1.5x per frontier), generating 64 tokens per iteration while displaying output.

### SSD-Streaming for Large Models

```bash
./ds4-bench \
  --model large.gguf \
  --prompt-file story_prompt.txt \
  --ssd-streaming \
  --ssd-streaming-cache-experts 16 \
  --ssd-streaming-full-layers 12 \
  --ctx-max 131072 \
  --gen-tokens 128

```

This configuration enables SSD-streaming mode for models exceeding VRAM capacity, caching 16 experts and 12 full layers while testing up to 131072 context length.

### Distributed Coordinator Mode

```bash

# Terminal 1: Start worker processes

./ds4 --role worker --listen 0.0.0.0:12345

# Terminal 2: Run benchmark as coordinator

./ds4-bench \
  --model ds4flash.gguf \
  --prompt-file my_prompt.txt \
  --dist-coord 127.0.0.1:12345 \
  --ctx-max 65536

```

The distributed mode requires separate worker processes; the benchmark acts as a coordinator, routing inference requests across the network.

## Understanding CSV Output

The benchmark produces a CSV file (or stdout) with the following columns:

- `ctx_tokens`: Total context size for the frontier
- `prefill_tokens`: Number of tokens processed in the pre-fill phase
- `prefill_tps`: Pre-fill throughput (tokens per second)
- `gen_tokens`: Number of tokens generated
- `gen_tps`: Generation throughput (tokens per second)
- `gen_first_ms`: Time to first token in milliseconds
- `gen_steady_tokens` and `gen_steady_tps`: Steady-state generation metrics
- `kvcache_bytes`: Size of the KV-cache snapshot in bytes

These metrics enable plotting throughput degradation curves as context length increases, identifying memory bandwidth bottlenecks and optimal batch configurations.

## Summary

- **ds4-bench** is implemented in [`ds4_bench.c`](https://github.com/antirez/ds4/blob/main/ds4_bench.c) and provides reproducible inference benchmarking across multiple backends.
- The tool measures **pre-fill TPS** via `ds4_session_sync` and **generation TPS** via `ds4_session_eval`, using `clock_gettime(CLOCK_MONOTONIC)` for nanosecond precision.
- Configuration is handled through `parse_options`, supporting context size frontiers, GPU device selection, and distributed coordination.
- Output is standardized as CSV containing per-frontier metrics including KV-cache size, enabling quantitative comparison of hardware and model configurations.
- SSD-streaming and distributed modes allow benchmarking of models that exceed single-machine memory limits.

## Frequently Asked Questions

### What metrics does ds4-bench measure?

The tool captures **pre-fill throughput** (tokens processed per second during prompt ingestion) and **generation throughput** (tokens generated per second during decoding). It also records time-to-first-token latency, steady-state generation speeds, and KV-cache memory footprint for each context size frontier.

### How does ds4-bench handle distributed benchmarking?

In distributed mode, `ds4-bench` acts as a coordinator connecting to worker processes started with `./ds4 --role worker`. The tool waits for route readiness via `wait_distributed_route` in [`ds4_distributed.c`](https://github.com/antirez/ds4/blob/main/ds4_distributed.c), then distributes the inference workload across the cluster while aggregating timing results locally.

### Can I use custom prompts for benchmarking?

Yes. Use the `--prompt-file` flag to specify a text file containing your fixed prompt, or `--chat-prompt-file` to automatically format the content as a chat prompt using `ds4_encode_chat_prompt`. The same prompt is reused across all context size frontiers to ensure consistent measurements.

### What is the difference between pre-fill and generation throughput?

**Pre-fill throughput** measures how quickly the model processes the input prompt (the "pre-fill" or "prompt encoding" phase), while **generation throughput** measures the speed of producing new tokens autoregressively. Pre-fill is typically compute-bound and highly parallelizable, whereas generation is often memory-bandwidth bound and slower due to sequential dependency.