How ds4-bench Measures Context Frontier Throughput in LLM Inference

ds4-bench evaluates context frontier throughput by timing how many tokens per second a model can process during the prefill phase (KV-cache population) and generate during the decode phase at increasing context lengths.

The ds4-bench tool in the antirez/ds4 repository provides a rigorous methodology for benchmarking large language model performance across varying context window sizes. By systematically measuring context frontier throughput—the tokens-per-second rate at each incremental context length—developers can identify performance bottlenecks in both prompt processing and token generation. This analysis examines the precise timing mechanisms and calculation methods implemented in the source code.

Understanding Context Frontier Throughput

In LLM inference, the context frontier represents the current sequence length being processed. As the benchmark progresses from smaller to larger context windows, it measures how efficiently the model handles increasingly longer prompts. The throughput calculation splits into two distinct operational phases: the prefill phase, where new tokens are processed into the KV-cache, and the generation phase, where the model produces new tokens autoregressively.

The Two-Phase Measurement Process

The benchmark implements separate timers for each phase in ds4_bench.c, capturing nanosecond-precision timestamps to compute accurate throughput metrics.

Prefill Phase: Measuring KV-Cache Population

During the prefill phase, ds4-bench measures how quickly the model processes the incremental portion of the prompt. The timing begins immediately before calling ds4_session_sync() at line 2890:

prefill_t0 = bench_now_sec();
// ... sync call executes ...
prefill_t1 = bench_now_sec();

According to the source code at lines 2898-2900 of ds4_bench.c, the benchmark calculates the elapsed time and token count:

  • prefill_sec = prefill_t1 - prefill_t0
  • prefill_tokens = frontier - previous

The prefill throughput equals prefill_tokens / prefill_sec, representing the rate at which new tokens enter the KV-cache (line 3002).

Generation Phase: Measuring Decode Performance

After prefill completion, the benchmark measures generation throughput by timing the decode loop. The timer initializes before the first generation step (line 3332) and stops after producing the configured number of tokens (line 3364):

gen_t0 = bench_now_sec();
// ... generation loop executes ...
gen_t1 = bench_now_sec();

The calculation at lines 3376-3380 derives:

  • gen_sec = gen_t1 - gen_t0
  • gen_done = actual tokens generated

Generation throughput computes as gen_done / gen_sec (line 3004). The benchmark also captures first-token latency (gen_first_sec) and steady-state throughput excluding the initial token.

How the Benchmark Iterates Through Context Sizes

The next_frontier() function (lines 2871-2884) controls the progression through context lengths, applying either fixed increments or multiplicative steps based on configuration. For each frontier, the benchmark executes a four-step workflow:

  1. Prefill the session to the current context length using ds4_session_sync
  2. Optionally snapshot the session state to avoid redundant re-prefilling
  3. Generate a fixed number of tokens (cfg.gen_tokens) while timing decode steps
  4. Write results to CSV containing frontier size, prefill metrics, generation metrics, and KV-cache size (lines 3000-3009)

Running ds4-bench: Practical Example

Execute the benchmark from the command line to measure performance from 2K to 32K tokens:

./ds4-bench \
  --model mymodel.gguf \
  --prompt-file myprompt.txt \
  --ctx-start 2048 \
  --ctx-max 32768 \
  --step-incr 2048 \
  --gen-tokens 128 \
  --csv results.csv

The tool generates a results.csv file with columns mapping directly to the internal measurements:


ctx_tokens,prefill_tokens,prefill_tps,gen_tokens,gen_tps,gen_first_ms,gen_steady_tokens,gen_steady_tps,kvcache_bytes
2048,2048,1234.56,128,567.89,12.3,127,578.12,12345678
4096,2048,1150.00,128,550.00,13.0,127,560.00,12456789

Each row represents a context frontier, with prefill_tps showing the throughput for new tokens added and gen_tps showing generation speed at that context length.

Key Source Files and Implementation Details

The measurement logic resides primarily in ds4_bench.c, which implements the timing harness and CSV output formatting. Supporting files include:

  • ds4_help.c: Contains DS4_HELP_BENCH documentation describing command-line options and benchmark purpose
  • ds4.h: Defines core API functions including ds4_session_sync and ds4_session_eval used during measurement
  • ds4_gpu_args.c: Parses GPU-specific arguments affecting the backend performance characteristics during benchmarking

Summary

  • ds4-bench measures context frontier throughput by walking a fixed token sequence through incrementally larger context windows.
  • The prefill phase timer (lines 2890-2900) calculates tokens-per-second for KV-cache population using bench_now_sec() before and after ds4_session_sync().
  • The generation phase timer (lines 3332-3364) measures decode throughput by tracking elapsed time across the generation loop.
  • Results export to CSV format (lines 3000-3009) including prefill TPS, generation TPS, first-token latency, and KV-cache utilization.
  • The next_frontier() function controls stepping logic through context sizes using either additive or multiplicative increments.

Frequently Asked Questions

What is the difference between prefill and generation throughput?

Prefill throughput measures how many tokens per second the model can process when adding new prompt tokens to the KV-cache, calculated as prefill_tokens / prefill_sec. Generation throughput measures the autoregressive decode speed when producing new tokens, calculated as gen_done / gen_sec. Prefill typically achieves higher throughput due to parallel processing, while generation is sequential and slower.

How does ds4-bench handle increasing context sizes?

The benchmark uses the next_frontier() function in ds4_bench.c (lines 2871-2884) to step through context lengths, either adding a fixed increment (--step-incr) or multiplying by a factor. For each frontier, it optionally snapshots the session state to avoid re-processing previous tokens, then measures performance at that specific context length.

What does the gen_first_ms column in the CSV output represent?

The gen_first_ms column records the first-token latency—the time elapsed from starting generation until the first token is produced. This metric helps identify cold-start penalties or initialization overhead in the inference backend, distinct from the steady-state generation rate shown in gen_steady_tps.

Why does the benchmark measure both prefill and generation phases separately?

LLM inference exhibits fundamentally different performance characteristics during these phases. Prefill benefits from parallelism across the attention heads when processing the prompt, while generation proceeds token-by-token in an autoregressive manner. Measuring both phases separately at each context frontier reveals how KV-cache size and memory bandwidth affect each operation type as sequences grow longer.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →