# How to Use the `ds4-eval` Tool for Capability Regression Testing in the ds4 Inference Engine

> Learn to use ds4-eval for capability regression testing with the ds4 inference engine. Detect regressions in execution output quality and performance using this built-in benchmark harness.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: how-to-guide
- Published: 2026-08-04

---

**`ds4-eval`** is a built‑in benchmark harness that runs a fixed set of prompt‑answer pairs through the **ds4** inference engine to detect regressions in model execution, output quality, and performance.

This guide walks through using `ds4‑eval` for **capability regression testing** after any change that touches the model's execution path—whether that's new CUDA kernels, routing logic, KV‑store implementations, or sampling modifications. The tool exercises the exact end‑to‑end path used in production, including token generation, context window management, and KV‑store syncing.

## What `ds4‑eval` Actually Tests

The harness validates three interconnected layers of the inference stack:

| Layer | Key Functions | Source File |
|-------|-------------|-------------|
| **Model loading & session creation** | `ds4_load_model()`, `ds4_session_new()` | [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) |
| **Prompt preprocessing & pre‑fill** | `ds4_prompt_chat()`, `ds4_session_sync()` | [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c) |
| **Token‑wise sampling & grading** | `ds4_session_step()`, answer comparison | [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c) |

Because `ds4‑eval` calls the same `ds4_session_step()` loop used by [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c), it catches regressions that unit tests would miss.

## The `eval_cases` Benchmark Suite

The test cases live in lines 95–~770 of [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c). Each entry in the `eval_cases` array contains:

- `source` — dataset name (GPQA Diamond, SuperGPQA, AIME 2025, COMPSEC)
- `id` — stable identifier for filtering
- `domain`, `title`, `question` — human‑readable metadata
- `choice[]` — multiple‑choice options where applicable
- `answer` — expected answer token (e.g., `"B"`, `"70"`, `"true"`)

## Building and Running `ds4‑eval`

### Build the Project

```bash
make

```

This produces `./ds4-eval` alongside other binaries.

### Basic Invocation

```bash

# Use default model path

./ds4-eval

# Or specify explicitly

DS4_MODEL=/path/to/checkpoint ./ds4-eval

```

### GPU Selection

```bash
DS4_GPU=1 DS4_MODEL=/path/to/model ./ds4-eval

```

### Filtering Test Cases

The `DS4_EVAL_FILTER` variable accepts dataset names or case IDs:

```bash

# Run only GPQA cases

DS4_EVAL_FILTER=GPQA ./ds4-eval

# Run a specific case by ID

DS4_EVAL_FILTER=recNu3MXkvWUzHZr9 ./ds4-eval

```

### Capturing Output

```bash
./ds4-eval | tee ds4-eval.log

```

## Understanding the Live UI

`ds4‑eval` renders a two‑pane ANSI display (defined by color macros at the top of [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c)):

- **Left pane**: Current prompt text
- **Right pane**: Generated answer, updating token‑by‑token

Sample final output:

```

✅  Case recNu3MXkvWUzHZr9 … PASSED
❌  Case 001b51d76b4d4229 … FAILED (got "A", expected "C")

```

Summary line:

```

🟢  23 / 24 cases passed – regression test SUCCESS

```

## Exit Codes for CI/CD Integration

| Exit Code | Meaning |
|-----------|---------|
| `0` | All cases passed |
| `1` | At least one case failed |
| `2` | Internal error (model loading, etc.) |

Use these in pre‑commit hooks or CI pipelines:

```bash
#!/bin/bash
DS4_EVAL_FILTER=GPQA ./ds4-eval || exit 1

```

## Key Source Files for Deep Dives

| File | Purpose |
|------|---------|
| [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c) | Harness implementation, `eval_cases` array, sampling loop |
| [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) / [`ds4.h`](https://github.com/antirez/ds4/blob/main/ds4.h) | Core inference API |
| [`ds4_session.c`](https://github.com/antirez/ds4/blob/main/ds4_session.c) | KV‑store handling, token generation, sampling |
| [`ds4_help.c`](https://github.com/antirez/ds4/blob/main/ds4_help.c) | `--help` text generation |
| `Makefile` | Build rules for `ds4-eval` |

## When to Run Regression Tests

Trigger `ds4‑eval` after any modification to:

- CUDA kernels in the attention or MLP layers
- KV‑cache memory layout or synchronization
- Token sampling logic (temperature, top‑p, ranking)
- Context window management
- Routing or MoE dispatch code

The harness validates that **output quality** (measured by case pass/fail) and **execution correctness** (measured by crash‑freedom) remain intact.

## Summary

- **`ds4‑eval`** is the canonical regression harness for the ds4 inference engine, shipping in the main repository.
- It runs **24+ curated cases** from GPQA Diamond, SuperGPQA, AIME 2025, and COMPSEC.
- The tool exercises the **full production inference path** via `ds4_session_step()` and `ds4_session_sync()`.
- **Filter with `DS4_EVAL_FILTER`** to debug specific failures without running the full suite.
- **Exit codes 0/1/2** enable seamless CI/CD integration.
- Source references: [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c) (harness), [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c)/[`ds4_session.c`](https://github.com/antirez/ds4/blob/main/ds4_session.c) (engine), `Makefile` (build).

## Frequently Asked Questions

### How do I add my own test cases to `ds4‑eval`?

Append new entries to the `eval_cases` array in [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c) (around line 95). Each entry requires `source`, `id`, `question`, and `answer` fields. Rebuild with `make` and run. The harness will automatically include your case in subsequent runs.

### Can I run `ds4‑eval` without a GPU?

No. The harness requires CUDA initialization through `ds4_load_model()`, which fails gracefully with exit code `2` if no GPU is available. The tool is designed to validate GPU inference paths specifically.

### What's the difference between `ds4‑eval` and `ds4‑server` benchmarking?

`ds4‑server` handles concurrent client requests and measures throughput under load. `ds4‑eval` runs sequential, deterministic cases with known answers to catch **correctness regressions**, not performance variations. Use both: `ds4‑eval` for correctness gates, `ds4‑server` stress tests for performance validation.

### Why does `ds4‑eval` sometimes skip cases?

Cases marked internally (or filtered via `DS4_EVAL_FILTER`) report `SKIPPED` in the UI. The harness also skips cases when the model's vocabulary lacks required tokens, or when context window exceeds force the prompt to truncate below a minimum threshold.