# How ds4-bench Measures Context Frontier Throughput in LLM Inference

> Discover how ds4-bench measures context frontier throughput for LLM inference. Learn about its evaluation of token processing speed during prefill and decode phases at varying context lengths.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: performance
- Published: 2026-08-09

---

**ds4-bench evaluates context frontier throughput by timing how many tokens per second a model can process during the prefill phase (KV-cache population) and generate during the decode phase at increasing context lengths.**

The `ds4-bench` tool in the antirez/ds4 repository provides a rigorous methodology for benchmarking large language model performance across varying context window sizes. By systematically measuring **context frontier throughput**—the tokens-per-second rate at each incremental context length—developers can identify performance bottlenecks in both prompt processing and token generation. This analysis examines the precise timing mechanisms and calculation methods implemented in the source code.

## Understanding Context Frontier Throughput

In LLM inference, the context frontier represents the current sequence length being processed. As the benchmark progresses from smaller to larger context windows, it measures how efficiently the model handles increasingly longer prompts. The throughput calculation splits into two distinct operational phases: the **prefill** phase, where new tokens are processed into the KV-cache, and the **generation** phase, where the model produces new tokens autoregressively.

## The Two-Phase Measurement Process

The benchmark implements separate timers for each phase in [`ds4_bench.c`](https://github.com/antirez/ds4/blob/main/ds4_bench.c), capturing nanosecond-precision timestamps to compute accurate throughput metrics.

### Prefill Phase: Measuring KV-Cache Population

During the prefill phase, `ds4-bench` measures how quickly the model processes the incremental portion of the prompt. The timing begins immediately before calling `ds4_session_sync()` at line 2890:

```c
prefill_t0 = bench_now_sec();
// ... sync call executes ...
prefill_t1 = bench_now_sec();

```

According to the source code at lines 2898-2900 of [`ds4_bench.c`](https://github.com/antirez/ds4/blob/main/ds4_bench.c), the benchmark calculates the elapsed time and token count:

- `prefill_sec = prefill_t1 - prefill_t0`
- `prefill_tokens = frontier - previous`

The **prefill throughput** equals `prefill_tokens / prefill_sec`, representing the rate at which new tokens enter the KV-cache (line 3002).

### Generation Phase: Measuring Decode Performance

After prefill completion, the benchmark measures generation throughput by timing the decode loop. The timer initializes before the first generation step (line 3332) and stops after producing the configured number of tokens (line 3364):

```c
gen_t0 = bench_now_sec();
// ... generation loop executes ...
gen_t1 = bench_now_sec();

```

The calculation at lines 3376-3380 derives:

- `gen_sec = gen_t1 - gen_t0`
- `gen_done` = actual tokens generated

**Generation throughput** computes as `gen_done / gen_sec` (line 3004). The benchmark also captures **first-token latency** (`gen_first_sec`) and steady-state throughput excluding the initial token.

## How the Benchmark Iterates Through Context Sizes

The `next_frontier()` function (lines 2871-2884) controls the progression through context lengths, applying either fixed increments or multiplicative steps based on configuration. For each frontier, the benchmark executes a four-step workflow:

1. **Prefill** the session to the current context length using `ds4_session_sync`
2. **Optionally snapshot** the session state to avoid redundant re-prefilling
3. **Generate** a fixed number of tokens (`cfg.gen_tokens`) while timing decode steps
4. **Write** results to CSV containing frontier size, prefill metrics, generation metrics, and KV-cache size (lines 3000-3009)

## Running ds4-bench: Practical Example

Execute the benchmark from the command line to measure performance from 2K to 32K tokens:

```bash
./ds4-bench \
  --model mymodel.gguf \
  --prompt-file myprompt.txt \
  --ctx-start 2048 \
  --ctx-max 32768 \
  --step-incr 2048 \
  --gen-tokens 128 \
  --csv results.csv

```

The tool generates a `results.csv` file with columns mapping directly to the internal measurements:

```

ctx_tokens,prefill_tokens,prefill_tps,gen_tokens,gen_tps,gen_first_ms,gen_steady_tokens,gen_steady_tps,kvcache_bytes
2048,2048,1234.56,128,567.89,12.3,127,578.12,12345678
4096,2048,1150.00,128,550.00,13.0,127,560.00,12456789

```

Each row represents a context frontier, with `prefill_tps` showing the throughput for new tokens added and `gen_tps` showing generation speed at that context length.

## Key Source Files and Implementation Details

The measurement logic resides primarily in [`ds4_bench.c`](https://github.com/antirez/ds4/blob/main/ds4_bench.c), which implements the timing harness and CSV output formatting. Supporting files include:

- **[`ds4_help.c`](https://github.com/antirez/ds4/blob/main/ds4_help.c)**: Contains `DS4_HELP_BENCH` documentation describing command-line options and benchmark purpose
- **[`ds4.h`](https://github.com/antirez/ds4/blob/main/ds4.h)**: Defines core API functions including `ds4_session_sync` and `ds4_session_eval` used during measurement
- **[`ds4_gpu_args.c`](https://github.com/antirez/ds4/blob/main/ds4_gpu_args.c)**: Parses GPU-specific arguments affecting the backend performance characteristics during benchmarking

## Summary

- **ds4-bench** measures **context frontier throughput** by walking a fixed token sequence through incrementally larger context windows.
- The **prefill phase** timer (lines 2890-2900) calculates tokens-per-second for KV-cache population using `bench_now_sec()` before and after `ds4_session_sync()`.
- The **generation phase** timer (lines 3332-3364) measures decode throughput by tracking elapsed time across the generation loop.
- Results export to CSV format (lines 3000-3009) including prefill TPS, generation TPS, first-token latency, and KV-cache utilization.
- The `next_frontier()` function controls stepping logic through context sizes using either additive or multiplicative increments.

## Frequently Asked Questions

### What is the difference between prefill and generation throughput?

**Prefill throughput** measures how many tokens per second the model can process when adding new prompt tokens to the KV-cache, calculated as `prefill_tokens / prefill_sec`. **Generation throughput** measures the autoregressive decode speed when producing new tokens, calculated as `gen_done / gen_sec`. Prefill typically achieves higher throughput due to parallel processing, while generation is sequential and slower.

### How does ds4-bench handle increasing context sizes?

The benchmark uses the `next_frontier()` function in [`ds4_bench.c`](https://github.com/antirez/ds4/blob/main/ds4_bench.c) (lines 2871-2884) to step through context lengths, either adding a fixed increment (`--step-incr`) or multiplying by a factor. For each frontier, it optionally snapshots the session state to avoid re-processing previous tokens, then measures performance at that specific context length.

### What does the gen_first_ms column in the CSV output represent?

The `gen_first_ms` column records the **first-token latency**—the time elapsed from starting generation until the first token is produced. This metric helps identify cold-start penalties or initialization overhead in the inference backend, distinct from the steady-state generation rate shown in `gen_steady_tps`.

### Why does the benchmark measure both prefill and generation phases separately?

LLM inference exhibits fundamentally different performance characteristics during these phases. Prefill benefits from parallelism across the attention heads when processing the prompt, while generation proceeds token-by-token in an autoregressive manner. Measuring both phases separately at each context frontier reveals how KV-cache size and memory bandwidth affect each operation type as sequences grow longer.