How to Benchmark Inference Performance with the ds4-bench Tool
The ds4-bench tool is a purpose-built throughput benchmark that walks a fixed prompt through increasing context-size frontiers, measuring pre-fill latency and greedy decode performance using high-resolution monotonic timers.
The antirez/ds4 repository provides a lightweight inference engine for large language models, shipping with ds4-bench for reproducible performance profiling. This command-line utility iterates through context length frontiers to capture both pre-fill tokens-per-second (TPS) and generation TPS across CPU, Metal, CUDA, and distributed backends. Understanding how to benchmark inference performance with ds4-bench enables accurate capacity planning and hardware optimization for production deployments.
Architecture and Implementation Details
The benchmark is implemented entirely in ds4_bench.c, utilizing the public API declared in ds4.h and the core engine logic from ds4.c. The tool orchestrates engine creation, prompt tokenization, and timed evaluation loops to produce standardized CSV reports.
Command-Line Configuration
Benchmark behavior is controlled through the parse_options function (lines 200-272 in ds4_bench.c), which populates a bench_config struct from flags such as --model, --ctx-start, --ctx-max, --step-incr, and --backend. Validation helpers like parse_int and parse_double_arg ensure numeric arguments are sane before the benchmark begins.
Key configuration parameters include:
--prompt-fileor--chat-prompt-file: Path to the text file containing the fixed prompt used for all frontiers--ctx-startand--ctx-max: Define the range of context lengths to test--step-incror--step-mul: Arithmetic or multiplicative step size between frontiers--gen-tokens: Number of tokens to generate during the decode phase (default is typically short, e.g., 32-128)--backend: Execution backend (metal,cuda,cpu, etc.)
Engine and Session Initialization
After parsing, the tool constructs a ds4_engine_options struct and creates the inference engine via ds4_engine_open or ds4_engine_create_with_gpu_config (lines 594-618). For GPU benchmarks, arguments parsed by ds4_gpu_args.c configure VRAM limits and device selection.
The prompt text is loaded via read_file and tokenized using ds4_tokenize_text or encoded as a chat prompt via ds4_encode_chat_prompt. A ds4_session is then allocated with the target context size (cfg.ctx_alloc). In distributed mode, the tool waits for route readiness through wait_distributed_route (defined in ds4_distributed.c).
The Benchmark Loop
The core measurement logic resides in the frontier iteration loop (lines 543-615). For each context size frontier, the tool executes:
- Pre-fill Phase: Calls
ds4_session_syncwith a prefix slice of the prompt to measure prompt processing throughput - Log Frontier Logits (optional): Writes per-frontier logits via
write_frontier_logits_json - Snapshot Creation: Optionally saves a session snapshot to avoid replaying the prefix after generation
- Generation Phase: Performs a short greedy decode of
cfg.gen_tokenstokens usingds4_session_eval - Session Restoration: Reloads the snapshot or re-syncs the prefix to maintain consistent state
- CSV Reporting: Appends metrics for the current frontier
Timing and Memory Tracking
High-resolution timing is captured via bench_now_sec, which uses clock_gettime(CLOCK_MONOTONIC) (lines 63-68) to ensure nanosecond-precision measurements for both pre-fill and generation phases. Memory footprint estimation is provided by log_context_memory (lines 86-103), which calculates the KV-cache size based on the selected backend and pre-fill chunk size.
Running ds4-bench: Command Examples
The following invocations demonstrate how to benchmark inference performance across different hardware configurations and modes.
Basic Metal Benchmark on Apple Silicon
./ds4-bench \
--model ds4flash.gguf \
--prompt-file my_prompt.txt \
--ctx-start 2048 \
--ctx-max 32768 \
--step-incr 2048 \
--backend metal \
--csv bench_results.csv
This command iterates from 2048 to 32768 tokens in 2048-token steps, outputting pre-fill and generation TPS to bench_results.csv.
Multi-GPU CUDA Testing
./ds4-bench \
--model ds4flash.gguf \
--prompt-file my_prompt.txt \
--gpu-devices 0,1 \
--gpu-vram auto \
--backend cuda \
--step-mul 1.5 \
--gen-tokens 64 \
--show-output
Here, the benchmark utilizes multiple GPU devices with automatic VRAM detection and multiplicative step sizing (1.5x per frontier), generating 64 tokens per iteration while displaying output.
SSD-Streaming for Large Models
./ds4-bench \
--model large.gguf \
--prompt-file story_prompt.txt \
--ssd-streaming \
--ssd-streaming-cache-experts 16 \
--ssd-streaming-full-layers 12 \
--ctx-max 131072 \
--gen-tokens 128
This configuration enables SSD-streaming mode for models exceeding VRAM capacity, caching 16 experts and 12 full layers while testing up to 131072 context length.
Distributed Coordinator Mode
# Terminal 1: Start worker processes
./ds4 --role worker --listen 0.0.0.0:12345
# Terminal 2: Run benchmark as coordinator
./ds4-bench \
--model ds4flash.gguf \
--prompt-file my_prompt.txt \
--dist-coord 127.0.0.1:12345 \
--ctx-max 65536
The distributed mode requires separate worker processes; the benchmark acts as a coordinator, routing inference requests across the network.
Understanding CSV Output
The benchmark produces a CSV file (or stdout) with the following columns:
ctx_tokens: Total context size for the frontierprefill_tokens: Number of tokens processed in the pre-fill phaseprefill_tps: Pre-fill throughput (tokens per second)gen_tokens: Number of tokens generatedgen_tps: Generation throughput (tokens per second)gen_first_ms: Time to first token in millisecondsgen_steady_tokensandgen_steady_tps: Steady-state generation metricskvcache_bytes: Size of the KV-cache snapshot in bytes
These metrics enable plotting throughput degradation curves as context length increases, identifying memory bandwidth bottlenecks and optimal batch configurations.
Summary
- ds4-bench is implemented in
ds4_bench.cand provides reproducible inference benchmarking across multiple backends. - The tool measures pre-fill TPS via
ds4_session_syncand generation TPS viads4_session_eval, usingclock_gettime(CLOCK_MONOTONIC)for nanosecond precision. - Configuration is handled through
parse_options, supporting context size frontiers, GPU device selection, and distributed coordination. - Output is standardized as CSV containing per-frontier metrics including KV-cache size, enabling quantitative comparison of hardware and model configurations.
- SSD-streaming and distributed modes allow benchmarking of models that exceed single-machine memory limits.
Frequently Asked Questions
What metrics does ds4-bench measure?
The tool captures pre-fill throughput (tokens processed per second during prompt ingestion) and generation throughput (tokens generated per second during decoding). It also records time-to-first-token latency, steady-state generation speeds, and KV-cache memory footprint for each context size frontier.
How does ds4-bench handle distributed benchmarking?
In distributed mode, ds4-bench acts as a coordinator connecting to worker processes started with ./ds4 --role worker. The tool waits for route readiness via wait_distributed_route in ds4_distributed.c, then distributes the inference workload across the cluster while aggregating timing results locally.
Can I use custom prompts for benchmarking?
Yes. Use the --prompt-file flag to specify a text file containing your fixed prompt, or --chat-prompt-file to automatically format the content as a chat prompt using ds4_encode_chat_prompt. The same prompt is reused across all context size frontiers to ensure consistent measurements.
What is the difference between pre-fill and generation throughput?
Pre-fill throughput measures how quickly the model processes the input prompt (the "pre-fill" or "prompt encoding" phase), while generation throughput measures the speed of producing new tokens autoregressively. Pre-fill is typically compute-bound and highly parallelizable, whereas generation is often memory-bandwidth bound and slower due to sequential dependency.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →