How to Use the `ds4-eval` Tool for Capability Regression Testing in the ds4 Inference Engine

ds4-eval is a built‑in benchmark harness that runs a fixed set of prompt‑answer pairs through the ds4 inference engine to detect regressions in model execution, output quality, and performance.

This guide walks through using ds4‑eval for capability regression testing after any change that touches the model's execution path—whether that's new CUDA kernels, routing logic, KV‑store implementations, or sampling modifications. The tool exercises the exact end‑to‑end path used in production, including token generation, context window management, and KV‑store syncing.

What ds4‑eval Actually Tests

The harness validates three interconnected layers of the inference stack:

Layer Key Functions Source File
Model loading & session creation ds4_load_model(), ds4_session_new() ds4.c
Prompt preprocessing & pre‑fill ds4_prompt_chat(), ds4_session_sync() ds4_eval.c
Token‑wise sampling & grading ds4_session_step(), answer comparison ds4_eval.c

Because ds4‑eval calls the same ds4_session_step() loop used by ds4_server.c, it catches regressions that unit tests would miss.

The eval_cases Benchmark Suite

The test cases live in lines 95–~770 of ds4_eval.c. Each entry in the eval_cases array contains:

  • source — dataset name (GPQA Diamond, SuperGPQA, AIME 2025, COMPSEC)
  • id — stable identifier for filtering
  • domain, title, question — human‑readable metadata
  • choice[] — multiple‑choice options where applicable
  • answer — expected answer token (e.g., "B", "70", "true")

Building and Running ds4‑eval

Build the Project

make

This produces ./ds4-eval alongside other binaries.

Basic Invocation


# Use default model path

./ds4-eval

# Or specify explicitly

DS4_MODEL=/path/to/checkpoint ./ds4-eval

GPU Selection

DS4_GPU=1 DS4_MODEL=/path/to/model ./ds4-eval

Filtering Test Cases

The DS4_EVAL_FILTER variable accepts dataset names or case IDs:


# Run only GPQA cases

DS4_EVAL_FILTER=GPQA ./ds4-eval

# Run a specific case by ID

DS4_EVAL_FILTER=recNu3MXkvWUzHZr9 ./ds4-eval

Capturing Output

./ds4-eval | tee ds4-eval.log

Understanding the Live UI

ds4‑eval renders a two‑pane ANSI display (defined by color macros at the top of ds4_eval.c):

  • Left pane: Current prompt text
  • Right pane: Generated answer, updating token‑by‑token

Sample final output:


✅  Case recNu3MXkvWUzHZr9 … PASSED
❌  Case 001b51d76b4d4229 … FAILED (got "A", expected "C")

Summary line:


🟢  23 / 24 cases passed – regression test SUCCESS

Exit Codes for CI/CD Integration

Exit Code Meaning
0 All cases passed
1 At least one case failed
2 Internal error (model loading, etc.)

Use these in pre‑commit hooks or CI pipelines:

#!/bin/bash
DS4_EVAL_FILTER=GPQA ./ds4-eval || exit 1

Key Source Files for Deep Dives

File Purpose
ds4_eval.c Harness implementation, eval_cases array, sampling loop
ds4.c / ds4.h Core inference API
ds4_session.c KV‑store handling, token generation, sampling
ds4_help.c --help text generation
Makefile Build rules for ds4-eval

When to Run Regression Tests

Trigger ds4‑eval after any modification to:

  • CUDA kernels in the attention or MLP layers
  • KV‑cache memory layout or synchronization
  • Token sampling logic (temperature, top‑p, ranking)
  • Context window management
  • Routing or MoE dispatch code

The harness validates that output quality (measured by case pass/fail) and execution correctness (measured by crash‑freedom) remain intact.

Summary

  • ds4‑eval is the canonical regression harness for the ds4 inference engine, shipping in the main repository.
  • It runs 24+ curated cases from GPQA Diamond, SuperGPQA, AIME 2025, and COMPSEC.
  • The tool exercises the full production inference path via ds4_session_step() and ds4_session_sync().
  • Filter with DS4_EVAL_FILTER to debug specific failures without running the full suite.
  • Exit codes 0/1/2 enable seamless CI/CD integration.
  • Source references: ds4_eval.c (harness), ds4.c/ds4_session.c (engine), Makefile (build).

Frequently Asked Questions

How do I add my own test cases to ds4‑eval?

Append new entries to the eval_cases array in ds4_eval.c (around line 95). Each entry requires source, id, question, and answer fields. Rebuild with make and run. The harness will automatically include your case in subsequent runs.

Can I run ds4‑eval without a GPU?

No. The harness requires CUDA initialization through ds4_load_model(), which fails gracefully with exit code 2 if no GPU is available. The tool is designed to validate GPU inference paths specifically.

What's the difference between ds4‑eval and ds4‑server benchmarking?

ds4‑server handles concurrent client requests and measures throughput under load. ds4‑eval runs sequential, deterministic cases with known answers to catch correctness regressions, not performance variations. Use both: ds4‑eval for correctness gates, ds4‑server stress tests for performance validation.

Why does ds4‑eval sometimes skip cases?

Cases marked internally (or filtered via DS4_EVAL_FILTER) report SKIPPED in the UI. The harness also skips cases when the model's vocabulary lacks required tokens, or when context window exceeds force the prompt to truncate below a minimum threshold.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →