How to Use the `ds4-eval` Tool for Capability Regression Testing in the ds4 Inference Engine
ds4-eval is a built‑in benchmark harness that runs a fixed set of prompt‑answer pairs through the ds4 inference engine to detect regressions in model execution, output quality, and performance.
This guide walks through using ds4‑eval for capability regression testing after any change that touches the model's execution path—whether that's new CUDA kernels, routing logic, KV‑store implementations, or sampling modifications. The tool exercises the exact end‑to‑end path used in production, including token generation, context window management, and KV‑store syncing.
What ds4‑eval Actually Tests
The harness validates three interconnected layers of the inference stack:
| Layer | Key Functions | Source File |
|---|---|---|
| Model loading & session creation | ds4_load_model(), ds4_session_new() |
ds4.c |
| Prompt preprocessing & pre‑fill | ds4_prompt_chat(), ds4_session_sync() |
ds4_eval.c |
| Token‑wise sampling & grading | ds4_session_step(), answer comparison |
ds4_eval.c |
Because ds4‑eval calls the same ds4_session_step() loop used by ds4_server.c, it catches regressions that unit tests would miss.
The eval_cases Benchmark Suite
The test cases live in lines 95–~770 of ds4_eval.c. Each entry in the eval_cases array contains:
source— dataset name (GPQA Diamond, SuperGPQA, AIME 2025, COMPSEC)id— stable identifier for filteringdomain,title,question— human‑readable metadatachoice[]— multiple‑choice options where applicableanswer— expected answer token (e.g.,"B","70","true")
Building and Running ds4‑eval
Build the Project
make
This produces ./ds4-eval alongside other binaries.
Basic Invocation
# Use default model path
./ds4-eval
# Or specify explicitly
DS4_MODEL=/path/to/checkpoint ./ds4-eval
GPU Selection
DS4_GPU=1 DS4_MODEL=/path/to/model ./ds4-eval
Filtering Test Cases
The DS4_EVAL_FILTER variable accepts dataset names or case IDs:
# Run only GPQA cases
DS4_EVAL_FILTER=GPQA ./ds4-eval
# Run a specific case by ID
DS4_EVAL_FILTER=recNu3MXkvWUzHZr9 ./ds4-eval
Capturing Output
./ds4-eval | tee ds4-eval.log
Understanding the Live UI
ds4‑eval renders a two‑pane ANSI display (defined by color macros at the top of ds4_eval.c):
- Left pane: Current prompt text
- Right pane: Generated answer, updating token‑by‑token
Sample final output:
✅ Case recNu3MXkvWUzHZr9 … PASSED
❌ Case 001b51d76b4d4229 … FAILED (got "A", expected "C")
Summary line:
🟢 23 / 24 cases passed – regression test SUCCESS
Exit Codes for CI/CD Integration
| Exit Code | Meaning |
|---|---|
0 |
All cases passed |
1 |
At least one case failed |
2 |
Internal error (model loading, etc.) |
Use these in pre‑commit hooks or CI pipelines:
#!/bin/bash
DS4_EVAL_FILTER=GPQA ./ds4-eval || exit 1
Key Source Files for Deep Dives
| File | Purpose |
|---|---|
ds4_eval.c |
Harness implementation, eval_cases array, sampling loop |
ds4.c / ds4.h |
Core inference API |
ds4_session.c |
KV‑store handling, token generation, sampling |
ds4_help.c |
--help text generation |
Makefile |
Build rules for ds4-eval |
When to Run Regression Tests
Trigger ds4‑eval after any modification to:
- CUDA kernels in the attention or MLP layers
- KV‑cache memory layout or synchronization
- Token sampling logic (temperature, top‑p, ranking)
- Context window management
- Routing or MoE dispatch code
The harness validates that output quality (measured by case pass/fail) and execution correctness (measured by crash‑freedom) remain intact.
Summary
ds4‑evalis the canonical regression harness for the ds4 inference engine, shipping in the main repository.- It runs 24+ curated cases from GPQA Diamond, SuperGPQA, AIME 2025, and COMPSEC.
- The tool exercises the full production inference path via
ds4_session_step()andds4_session_sync(). - Filter with
DS4_EVAL_FILTERto debug specific failures without running the full suite. - Exit codes 0/1/2 enable seamless CI/CD integration.
- Source references:
ds4_eval.c(harness),ds4.c/ds4_session.c(engine),Makefile(build).
Frequently Asked Questions
How do I add my own test cases to ds4‑eval?
Append new entries to the eval_cases array in ds4_eval.c (around line 95). Each entry requires source, id, question, and answer fields. Rebuild with make and run. The harness will automatically include your case in subsequent runs.
Can I run ds4‑eval without a GPU?
No. The harness requires CUDA initialization through ds4_load_model(), which fails gracefully with exit code 2 if no GPU is available. The tool is designed to validate GPU inference paths specifically.
What's the difference between ds4‑eval and ds4‑server benchmarking?
ds4‑server handles concurrent client requests and measures throughput under load. ds4‑eval runs sequential, deterministic cases with known answers to catch correctness regressions, not performance variations. Use both: ds4‑eval for correctness gates, ds4‑server stress tests for performance validation.
Why does ds4‑eval sometimes skip cases?
Cases marked internally (or filtered via DS4_EVAL_FILTER) report SKIPPED in the UI. The harness also skips cases when the model's vocabulary lacks required tokens, or when context window exceeds force the prompt to truncate below a minimum threshold.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →