How to Use ds4-eval for Capability Regression Testing in the ds4 Inference Engine

ds4-eval is the built-in benchmark harness that loads a real model, runs fixed prompt-answer pairs from the eval_cases array, and grades output against expected answers to catch regressions in the full inference pipeline.

The ds4-eval tool ships with the ds4 inference engine (antirez/ds4) as a complete regression testing framework. It exercises every layer of the model execution path—from token generation and context window management to KV-store syncing and sampling logic—making it essential for validating changes to CUDA kernels, routing logic, or the KV-store implementation.

What ds4-eval Tests

Unlike unit tests that mock components, ds4-eval runs end-to-end inference. This design catches subtle bugs that only appear when subsystems interact: memory layout shifts affecting generation quality, sampling temperature changes altering answer distributions, or context window edge cases corrupting outputs.

The harness validates against a curated dataset drawn from GPQA Diamond, SuperGPQA, AIME 2025, and COMPSEC (lines 95–770 of ds4_eval.c). Each eval_case entry contains:

  • source – originating dataset name
  • id – stable identifier for filtering
  • domain, title, question – human-readable metadata
  • choice[] – multiple-choice options where applicable
  • answer – expected answer token (e.g., "B", "70", "true")

Architecture of the Evaluation Pipeline

The tool operates through three logical layers defined in ds4_eval.c and supporting files:

Model Loading and Session Creation

ds4_load_model() in ds4.c reads the checkpoint and creates a ds4_t handle. ds4_session_new() allocates the session and its KV store, establishing the runtime environment used by production servers.

Prompt Preprocessing and Pre-fill

The harness constructs a chat-formatted prompt via ds4_prompt_chat(), then calls ds4_session_sync() to pre-fill the model's context with the prompt text before sampling begins. This matches the production prefill behavior in ds4_server.c.

Token-wise Sampling and Grading

A generation loop in ds4_eval.c repeatedly calls ds4_session_step() to produce one token at a time, applies temperature and ranking logic, appends to the running answer, and finally compares against the expected result. Status codes: EVAL_PASSED, EVAL_FAILED, or EVAL_SKIPPED.

Building and Running ds4-eval

Compile the binary from the repository root:

make

This produces ./ds4-eval alongside other ds4 binaries.

Basic Execution

Run with default model path detection:

./ds4-eval

Specify a checkpoint directory explicitly:

DS4_MODEL=/path/to/model ./ds4-eval

Target a specific GPU in multi-GPU systems:

DS4_GPU=1 DS4_MODEL=/path/to/model ./ds4-eval

Filtering and Debugging

Run only cases from a specific dataset:

DS4_EVAL_FILTER=GPQA ./ds4-eval

Filter to a single case by ID:

DS4_EVAL_FILTER=recNu3MXkvWUzHZr9 ./ds4-eval

Capture the full ANSI UI for later analysis:

./ds4-eval | tee ds4-eval.log

Interpreting Results

The harness renders a two-pane terminal UI (defined by color macros at the top of ds4_eval.c): prompts on the left, generated answers on the right, updating in-place as tokens stream.

Terminal output follows this pattern:


▶️  Question: [prompt text]
🟢  Answer so far: [accumulating generation]
✅  Case recNu3MXkvWUzHZr9 … PASSED
❌  Case 001b51d76b4d4229 … FAILED (got "A", expected "C")

Final summary format:


🟢  23 / 24 cases passed – regression test SUCCESS

Exit Codes

  • 0 – all cases passed
  • 1 – one or more cases failed
  • 2 – internal error (model loading failure, etc.)

Integration into Development Workflow

Run ds4-eval after any change touching the inference path. This includes:

  • New CUDA kernels in ds4_cuda*.cu files
  • KV-store modifications in ds4_kv.c or ds4_session.c
  • Sampling logic changes in ds4_sampling.c
  • Context window or attention mechanism updates

The harness runs the identical code path as ds4_server.c, ensuring that optimizations or refactors don't degrade model capabilities. Because the eval_cases array is version-controlled, results are reproducible across commits and machines.

Key Source Files

File Purpose
ds4_eval.c Harness implementation, eval_cases array, sampling loop, ANSI UI
ds4.c / ds4.h Core inference API: ds4_load_model(), ds4_t handle definition
ds4_session.c Session lifecycle: KV-store management, ds4_session_step(), ds4_session_sync()
ds4_help.c Unified help text system for all ds4 binaries
Makefile Build rules compiling ds4-eval with the engine

Summary

  • ds4-eval provides end-to-end capability regression testing for the ds4 inference engine
  • It exercises the full production code path including token generation, KV-store operations, and sampling
  • The eval_cases array contains validated questions from GPQA, SuperGPQA, AIME 2025, and COMPSEC
  • Use DS4_MODEL, DS4_GPU, and DS4_EVAL_FILTER environment variables to control execution
  • Exit code 0 confirms all tests passed; non-zero codes indicate failures or errors
  • Integrate into CI/CD pipelines to catch regressions before production deployment

Frequently Asked Questions

What makes ds4-eval different from other LLM benchmarks?

ds4-eval executes the exact inference stack used in production, not a separate evaluation framework. According to the antirez/ds4 source code, it calls ds4_session_step() and ds4_session_sync() directly—the same functions driving ds4_server.c. This catches integration bugs that isolated benchmarks miss.

How do I add custom test cases to ds4-eval?

Extend the eval_cases array in ds4_eval.c (around line 95). Each entry requires source, id, domain, title, question, optional choice[] array, and answer string. Rebuild with make to incorporate new cases.

Can ds4-eval run without a GPU?

The tool requires CUDA-capable hardware. The DS4_GPU environment variable selects among multiple GPUs (0, 1, etc.) but doesn't enable CPU fallback. Check ds4.c for the ds4_load_model() implementation that initializes CUDA contexts.

Why does ds4-eval show different results than my manual prompting?

The harness uses fixed sampling parameters and chat-formatted prompts via ds4_prompt_chat(). Temperature, top-p, and other sampling settings are controlled within ds4_eval.c rather than user-provided. For debugging, run with DS4_EVAL_FILTER on a single case and compare token-by-token output against manual ds4_session calls.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →