# How to Use ds4-eval for Capability Regression Testing in the ds4 Inference Engine

> Learn how to use ds4-eval for capability regression testing. This benchmark tool helps catch regressions in the ds4 inference engine by running fixed prompt-answer pairs and grading output.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: how-to-guide
- Published: 2026-08-05

---

**`ds4-eval` is the built-in benchmark harness that loads a real model, runs fixed prompt-answer pairs from the `eval_cases` array, and grades output against expected answers to catch regressions in the full inference pipeline.**

The `ds4-eval` tool ships with the **ds4** inference engine (antirez/ds4) as a complete regression testing framework. It exercises every layer of the model execution path—from token generation and context window management to KV-store syncing and sampling logic—making it essential for validating changes to CUDA kernels, routing logic, or the KV-store implementation.

## What ds4-eval Tests

Unlike unit tests that mock components, `ds4-eval` runs **end-to-end inference**. This design catches subtle bugs that only appear when subsystems interact: memory layout shifts affecting generation quality, sampling temperature changes altering answer distributions, or context window edge cases corrupting outputs.

The harness validates against a curated dataset drawn from **GPQA Diamond**, **SuperGPQA**, **AIME 2025**, and **COMPSEC** (lines 95–770 of [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c)). Each `eval_case` entry contains:

- `source` – originating dataset name
- `id` – stable identifier for filtering
- `domain`, `title`, `question` – human-readable metadata
- `choice[]` – multiple-choice options where applicable
- `answer` – expected answer token (e.g., `"B"`, `"70"`, `"true"`)

## Architecture of the Evaluation Pipeline

The tool operates through three logical layers defined in [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c) and supporting files:

### Model Loading and Session Creation

`ds4_load_model()` in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) reads the checkpoint and creates a `ds4_t` handle. `ds4_session_new()` allocates the session and its KV store, establishing the runtime environment used by production servers.

### Prompt Preprocessing and Pre-fill

The harness constructs a chat-formatted prompt via `ds4_prompt_chat()`, then calls `ds4_session_sync()` to pre-fill the model's context with the prompt text before sampling begins. This matches the production prefill behavior in [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c).

### Token-wise Sampling and Grading

A generation loop in [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c) repeatedly calls `ds4_session_step()` to produce one token at a time, applies temperature and ranking logic, appends to the running answer, and finally compares against the expected result. Status codes: `EVAL_PASSED`, `EVAL_FAILED`, or `EVAL_SKIPPED`.

## Building and Running ds4-eval

Compile the binary from the repository root:

```bash
make

```

This produces `./ds4-eval` alongside other ds4 binaries.

### Basic Execution

Run with default model path detection:

```bash
./ds4-eval

```

Specify a checkpoint directory explicitly:

```bash
DS4_MODEL=/path/to/model ./ds4-eval

```

Target a specific GPU in multi-GPU systems:

```bash
DS4_GPU=1 DS4_MODEL=/path/to/model ./ds4-eval

```

### Filtering and Debugging

Run only cases from a specific dataset:

```bash
DS4_EVAL_FILTER=GPQA ./ds4-eval

```

Filter to a single case by ID:

```bash
DS4_EVAL_FILTER=recNu3MXkvWUzHZr9 ./ds4-eval

```

Capture the full ANSI UI for later analysis:

```bash
./ds4-eval | tee ds4-eval.log

```

## Interpreting Results

The harness renders a two-pane terminal UI (defined by color macros at the top of [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c)): prompts on the left, generated answers on the right, updating in-place as tokens stream.

Terminal output follows this pattern:

```

▶️  Question: [prompt text]
🟢  Answer so far: [accumulating generation]
✅  Case recNu3MXkvWUzHZr9 … PASSED
❌  Case 001b51d76b4d4229 … FAILED (got "A", expected "C")

```

Final summary format:

```

🟢  23 / 24 cases passed – regression test SUCCESS

```

### Exit Codes

- `0` – all cases passed
- `1` – one or more cases failed
- `2` – internal error (model loading failure, etc.)

## Integration into Development Workflow

**Run `ds4-eval` after any change touching the inference path.** This includes:

- New CUDA kernels in `ds4_cuda*.cu` files
- KV-store modifications in [`ds4_kv.c`](https://github.com/antirez/ds4/blob/main/ds4_kv.c) or [`ds4_session.c`](https://github.com/antirez/ds4/blob/main/ds4_session.c)
- Sampling logic changes in [`ds4_sampling.c`](https://github.com/antirez/ds4/blob/main/ds4_sampling.c)
- Context window or attention mechanism updates

The harness runs the **identical** code path as [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c), ensuring that optimizations or refactors don't degrade model capabilities. Because the `eval_cases` array is version-controlled, results are reproducible across commits and machines.

## Key Source Files

| File | Purpose |
|------|---------|
| [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c) | Harness implementation, `eval_cases` array, sampling loop, ANSI UI |
| [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) / [`ds4.h`](https://github.com/antirez/ds4/blob/main/ds4.h) | Core inference API: `ds4_load_model()`, `ds4_t` handle definition |
| [`ds4_session.c`](https://github.com/antirez/ds4/blob/main/ds4_session.c) | Session lifecycle: KV-store management, `ds4_session_step()`, `ds4_session_sync()` |
| [`ds4_help.c`](https://github.com/antirez/ds4/blob/main/ds4_help.c) | Unified help text system for all ds4 binaries |
| `Makefile` | Build rules compiling `ds4-eval` with the engine |

## Summary

- **`ds4-eval`** provides end-to-end capability regression testing for the ds4 inference engine
- It exercises the **full production code path** including token generation, KV-store operations, and sampling
- The **`eval_cases`** array contains validated questions from GPQA, SuperGPQA, AIME 2025, and COMPSEC
- Use **`DS4_MODEL`**, **`DS4_GPU`**, and **`DS4_EVAL_FILTER`** environment variables to control execution
- **Exit code `0`** confirms all tests passed; non-zero codes indicate failures or errors
- Integrate into CI/CD pipelines to catch regressions before production deployment

## Frequently Asked Questions

### What makes ds4-eval different from other LLM benchmarks?

`ds4-eval` executes the **exact inference stack** used in production, not a separate evaluation framework. According to the antirez/ds4 source code, it calls `ds4_session_step()` and `ds4_session_sync()` directly—the same functions driving [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c). This catches integration bugs that isolated benchmarks miss.

### How do I add custom test cases to ds4-eval?

Extend the `eval_cases` array in [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c) (around line 95). Each entry requires `source`, `id`, `domain`, `title`, `question`, optional `choice[]` array, and `answer` string. Rebuild with `make` to incorporate new cases.

### Can ds4-eval run without a GPU?

The tool requires CUDA-capable hardware. The `DS4_GPU` environment variable selects among multiple GPUs (`0`, `1`, etc.) but doesn't enable CPU fallback. Check [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) for the `ds4_load_model()` implementation that initializes CUDA contexts.

### Why does ds4-eval show different results than my manual prompting?

The harness uses **fixed sampling parameters** and **chat-formatted prompts** via `ds4_prompt_chat()`. Temperature, top-p, and other sampling settings are controlled within [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c) rather than user-provided. For debugging, run with `DS4_EVAL_FILTER` on a single case and compare token-by-token output against manual `ds4_session` calls.