How to Use ds4-eval for Capability Regression Testing in the ds4 Inference Engine
ds4-eval is the built-in benchmark harness that loads a real model, runs fixed prompt-answer pairs from the eval_cases array, and grades output against expected answers to catch regressions in the full inference pipeline.
The ds4-eval tool ships with the ds4 inference engine (antirez/ds4) as a complete regression testing framework. It exercises every layer of the model execution path—from token generation and context window management to KV-store syncing and sampling logic—making it essential for validating changes to CUDA kernels, routing logic, or the KV-store implementation.
What ds4-eval Tests
Unlike unit tests that mock components, ds4-eval runs end-to-end inference. This design catches subtle bugs that only appear when subsystems interact: memory layout shifts affecting generation quality, sampling temperature changes altering answer distributions, or context window edge cases corrupting outputs.
The harness validates against a curated dataset drawn from GPQA Diamond, SuperGPQA, AIME 2025, and COMPSEC (lines 95–770 of ds4_eval.c). Each eval_case entry contains:
source– originating dataset nameid– stable identifier for filteringdomain,title,question– human-readable metadatachoice[]– multiple-choice options where applicableanswer– expected answer token (e.g.,"B","70","true")
Architecture of the Evaluation Pipeline
The tool operates through three logical layers defined in ds4_eval.c and supporting files:
Model Loading and Session Creation
ds4_load_model() in ds4.c reads the checkpoint and creates a ds4_t handle. ds4_session_new() allocates the session and its KV store, establishing the runtime environment used by production servers.
Prompt Preprocessing and Pre-fill
The harness constructs a chat-formatted prompt via ds4_prompt_chat(), then calls ds4_session_sync() to pre-fill the model's context with the prompt text before sampling begins. This matches the production prefill behavior in ds4_server.c.
Token-wise Sampling and Grading
A generation loop in ds4_eval.c repeatedly calls ds4_session_step() to produce one token at a time, applies temperature and ranking logic, appends to the running answer, and finally compares against the expected result. Status codes: EVAL_PASSED, EVAL_FAILED, or EVAL_SKIPPED.
Building and Running ds4-eval
Compile the binary from the repository root:
make
This produces ./ds4-eval alongside other ds4 binaries.
Basic Execution
Run with default model path detection:
./ds4-eval
Specify a checkpoint directory explicitly:
DS4_MODEL=/path/to/model ./ds4-eval
Target a specific GPU in multi-GPU systems:
DS4_GPU=1 DS4_MODEL=/path/to/model ./ds4-eval
Filtering and Debugging
Run only cases from a specific dataset:
DS4_EVAL_FILTER=GPQA ./ds4-eval
Filter to a single case by ID:
DS4_EVAL_FILTER=recNu3MXkvWUzHZr9 ./ds4-eval
Capture the full ANSI UI for later analysis:
./ds4-eval | tee ds4-eval.log
Interpreting Results
The harness renders a two-pane terminal UI (defined by color macros at the top of ds4_eval.c): prompts on the left, generated answers on the right, updating in-place as tokens stream.
Terminal output follows this pattern:
▶️ Question: [prompt text]
🟢 Answer so far: [accumulating generation]
✅ Case recNu3MXkvWUzHZr9 … PASSED
❌ Case 001b51d76b4d4229 … FAILED (got "A", expected "C")
Final summary format:
🟢 23 / 24 cases passed – regression test SUCCESS
Exit Codes
0– all cases passed1– one or more cases failed2– internal error (model loading failure, etc.)
Integration into Development Workflow
Run ds4-eval after any change touching the inference path. This includes:
- New CUDA kernels in
ds4_cuda*.cufiles - KV-store modifications in
ds4_kv.cords4_session.c - Sampling logic changes in
ds4_sampling.c - Context window or attention mechanism updates
The harness runs the identical code path as ds4_server.c, ensuring that optimizations or refactors don't degrade model capabilities. Because the eval_cases array is version-controlled, results are reproducible across commits and machines.
Key Source Files
| File | Purpose |
|---|---|
ds4_eval.c |
Harness implementation, eval_cases array, sampling loop, ANSI UI |
ds4.c / ds4.h |
Core inference API: ds4_load_model(), ds4_t handle definition |
ds4_session.c |
Session lifecycle: KV-store management, ds4_session_step(), ds4_session_sync() |
ds4_help.c |
Unified help text system for all ds4 binaries |
Makefile |
Build rules compiling ds4-eval with the engine |
Summary
ds4-evalprovides end-to-end capability regression testing for the ds4 inference engine- It exercises the full production code path including token generation, KV-store operations, and sampling
- The
eval_casesarray contains validated questions from GPQA, SuperGPQA, AIME 2025, and COMPSEC - Use
DS4_MODEL,DS4_GPU, andDS4_EVAL_FILTERenvironment variables to control execution - Exit code
0confirms all tests passed; non-zero codes indicate failures or errors - Integrate into CI/CD pipelines to catch regressions before production deployment
Frequently Asked Questions
What makes ds4-eval different from other LLM benchmarks?
ds4-eval executes the exact inference stack used in production, not a separate evaluation framework. According to the antirez/ds4 source code, it calls ds4_session_step() and ds4_session_sync() directly—the same functions driving ds4_server.c. This catches integration bugs that isolated benchmarks miss.
How do I add custom test cases to ds4-eval?
Extend the eval_cases array in ds4_eval.c (around line 95). Each entry requires source, id, domain, title, question, optional choice[] array, and answer string. Rebuild with make to incorporate new cases.
Can ds4-eval run without a GPU?
The tool requires CUDA-capable hardware. The DS4_GPU environment variable selects among multiple GPUs (0, 1, etc.) but doesn't enable CPU fallback. Check ds4.c for the ds4_load_model() implementation that initializes CUDA contexts.
Why does ds4-eval show different results than my manual prompting?
The harness uses fixed sampling parameters and chat-formatted prompts via ds4_prompt_chat(). Temperature, top-p, and other sampling settings are controlled within ds4_eval.c rather than user-provided. For debugging, run with DS4_EVAL_FILTER on a single case and compare token-by-token output against manual ds4_session calls.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →