# How to Evaluate Model Capabilities Using ds4-eval: Complete Benchmark Guide

> Learn how to evaluate model capabilities using ds4-eval. Run the binary against GGUF models for a stable inference pipeline across science math and security benchmarks.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: how-to-guide
- Published: 2026-08-08

---

**To evaluate model capabilities using `ds4-eval`, run the compiled binary against a GGUF model file to execute a deterministic regression suite that validates inference pipeline stability across hard-science, mathematics, and security benchmarks.**

The `ds4-eval` tool in the `antirez/ds4` repository provides a lightweight, real-model integration benchmark designed for DeepSeek V4 Flash and GLM 5.2 models. Unlike leaderboard systems that rank model performance, this utility serves as a regression suite that detects subtle drift in quantization, KV-cache handling, and token generation by comparing outputs against ground-truth answers embedded directly in the source code.

## The Six-Stage Evaluation Pipeline

The `ds4-eval` benchmark follows a strict pipeline implemented entirely within [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c). Each stage exercises critical components of the inference stack to ensure deterministic behavior across builds.

### Loading the GGUF Model

The evaluation begins by loading the quantized model weights from a GGUF file. In [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c) (lines 5-15), the tool initializes the model context, allocates the KV cache, and prepares the inference backend (Metal, CUDA, or ROCm) specified via command-line flags.

### Prompt Rendering (lines 30-44)

Each test case is converted into a DeepSeek-compatible chat prompt consisting of system and user messages. The rendering logic in [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c) (lines 30-44) formats questions from the embedded datasets—GPQA Diamond, SuperGPQA, AIME 2025, and COMPSEC—into the precise token structure expected by the model architecture.

### Token Prefilling and Generation (lines 70-95)

The model pre-fills the prompt context, then enters a sampling loop that generates tokens until reaching a budget limit or stop token. This stage, implemented in [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c) (lines 70-95), exercises the full inference path including attention mechanisms and quantization kernels. The deterministic nature of this loop allows the tool to detect regressions in generation quality caused by subtle implementation changes.

### Answer Extraction (lines 210-240)

After generation completes, [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c) (lines 210-240) parses the raw token stream to extract structured answers. The extraction logic handles multiple-choice letters (A, B, C, D) and exact numeric answers, normalizing whitespace and formatting to prepare for comparison against ground truth.

### Scoring Against Ground Truth (lines 250-280)

The extracted answer is compared with the expected value stored in the case definition. The scoring logic in [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c) (lines 250-280) assigns a pass/fail status per question, flagging any deviation from the reference answer that might indicate model drift or inference bugs.

### Result Reporting (lines 300-340)

Finally, [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c) (lines 300-340) prints a concise table showing the status, token counts, model answer, and ground truth for each question. This output format enables quick visual inspection or automated parsing in CI pipelines.

## Running ds4-eval: Practical Examples

The `ds4-eval` binary supports both interactive debugging and automated regression testing through various command-line options.

### Basic Interactive Evaluation

Run the benchmark with the TUI enabled to monitor progress in real time:

```bash
./ds4-eval -m ds4flash.gguf --trace /tmp/ds4-eval.txt

```

This command loads `ds4flash.gguf`, renders all embedded test cases, and writes a complete generation trace to [`/tmp/ds4-eval.txt`](https://github.com/antirez/ds4/blob/main//tmp/ds4-eval.txt) for later analysis. Typical output shows per-question status:

```

# Question 1  PASSED  tokens=2048  model=B  correct=B

# Question 2  PASSED  tokens=438   model=C  correct=C

# Question 3  FAILED  tokens=2048  model=A  correct=C

```

### Headless CI/CD Integration

For automated pipelines, disable the curses interface and fix random seeds to ensure reproducible results:

```bash
./ds4-eval -m ds4flash.gguf --plain \
  --questions 4 --tokens 2048 --temp 0 --seed 1

```

The `--plain` flag streams plain text suitable for log files. Parse the results programmatically:

```bash
./ds4-eval -m ds4flash.gguf --plain | grep PASSED | wc -l

```

### Re-grading Existing Traces

When modifying answer extraction logic, avoid expensive re-generation by re-running the scorer on stored tokens:

```bash
./ds4-eval -m ds4flash.gguf --regrade-trace /tmp/ds4-eval.txt

```

This command re-executes only the extraction and scoring phases defined in [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c) (lines 210-280), enabling rapid iteration on grading algorithms.

### Testing Specific Question Ranges

Limit evaluation to a subset of questions for quick smoke tests:

```bash
./ds4-eval -m ds4flash.gguf --questions 1-10

```

This runs only the first ten cases from the embedded `eval_cases[]` array, reducing validation time from minutes to seconds.

### Cross-Platform Backend Support

Compile and run the benchmark on any supported hardware backend:

```bash

# CUDA multi-GPU

make cuda-spark
./ds4-eval -m ds4flash.gguf --cuda

```

The same [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c) source works on Metal, CUDA, and ROCm without modification, ensuring consistent evaluation metrics across deployment targets.

## Why Use ds4-eval for Model Validation?

**Detects Subtle Drift:** Even minor changes to KV-cache eviction policies or quantization kernel selection can alter token distributions. The benchmark catches these regressions before they reach production.

**Fast Feedback Loop:** A full run completes in seconds on modern GPUs (approximately 10 tokens/second generation), providing immediate validation of inference pipeline changes.

**Zero External Dependencies:** All benchmark data resides in [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c) as static arrays. No network calls, no dataset downloads, and no Python environment required.

**Extensible Architecture:** Add custom regression cases by appending to the `eval_cases[]` array in [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c) and recompiling, allowing teams to validate domain-specific capabilities alongside the standard suite.

## Summary

- **`ds4-eval`** provides deterministic regression testing for GGUF-based models through a six-stage pipeline implemented in [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c).
- The benchmark covers **GPQA Diamond**, **SuperGPQA**, **AIME 2025**, and **COMPSEC** datasets to verify hard-science, math, and security coding capabilities.
- Source code locations: model loading (lines 5-15), prompt rendering (30-44), generation loop (70-95), extraction (210-240), scoring (250-280), and reporting (300-340).
- Run interactively with `--trace` or automate with `--plain`, `--questions`, and `--regrade-trace` flags.
- Supports Metal, CUDA, and ROCm backends through compile-time targets in the repository `Makefile`.

## Frequently Asked Questions

### What model formats does ds4-eval support?

`ds4-eval` exclusively supports **GGUF** files, the quantized format used by DeepSeek V4 Flash and GLM 5.2 models. The loading logic in [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c) (lines 5-15) initializes the model context directly from the GGUF binary without requiring external conversion tools.

### How long does a typical evaluation take?

A complete run of the default benchmark suite finishes in **a few seconds** on modern GPU hardware, generating approximately 10 tokens per second. You can reduce this further by using `--questions` to limit the evaluation to specific subsets for rapid smoke testing.

### Can I add custom benchmark cases to ds4-eval?

Yes. The benchmark is **extensible** by design. Add new cases to the `eval_cases[]` array defined in [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c), then recompile using the standard `Makefile` targets. This allows teams to incorporate proprietary or domain-specific validation questions alongside the standard GPQA and COMPSEC datasets.

### Does ds4-eval require external dependencies?

No. The tool is **self-contained** within the `antirez/ds4` repository. All test questions, ground-truth answers, and scoring logic are embedded as static data structures in [`ds4_eval.c`](https://github.com/antirez/ds4/blob/main/ds4_eval.c), eliminating network dependencies, dataset downloads, or external Python environments.