How to Evaluate Model Capabilities Using ds4-eval: Complete Benchmark Guide

To evaluate model capabilities using ds4-eval, run the compiled binary against a GGUF model file to execute a deterministic regression suite that validates inference pipeline stability across hard-science, mathematics, and security benchmarks.

The ds4-eval tool in the antirez/ds4 repository provides a lightweight, real-model integration benchmark designed for DeepSeek V4 Flash and GLM 5.2 models. Unlike leaderboard systems that rank model performance, this utility serves as a regression suite that detects subtle drift in quantization, KV-cache handling, and token generation by comparing outputs against ground-truth answers embedded directly in the source code.

The Six-Stage Evaluation Pipeline

The ds4-eval benchmark follows a strict pipeline implemented entirely within ds4_eval.c. Each stage exercises critical components of the inference stack to ensure deterministic behavior across builds.

Loading the GGUF Model

The evaluation begins by loading the quantized model weights from a GGUF file. In ds4_eval.c (lines 5-15), the tool initializes the model context, allocates the KV cache, and prepares the inference backend (Metal, CUDA, or ROCm) specified via command-line flags.

Prompt Rendering (lines 30-44)

Each test case is converted into a DeepSeek-compatible chat prompt consisting of system and user messages. The rendering logic in ds4_eval.c (lines 30-44) formats questions from the embedded datasets—GPQA Diamond, SuperGPQA, AIME 2025, and COMPSEC—into the precise token structure expected by the model architecture.

Token Prefilling and Generation (lines 70-95)

The model pre-fills the prompt context, then enters a sampling loop that generates tokens until reaching a budget limit or stop token. This stage, implemented in ds4_eval.c (lines 70-95), exercises the full inference path including attention mechanisms and quantization kernels. The deterministic nature of this loop allows the tool to detect regressions in generation quality caused by subtle implementation changes.

Answer Extraction (lines 210-240)

After generation completes, ds4_eval.c (lines 210-240) parses the raw token stream to extract structured answers. The extraction logic handles multiple-choice letters (A, B, C, D) and exact numeric answers, normalizing whitespace and formatting to prepare for comparison against ground truth.

Scoring Against Ground Truth (lines 250-280)

The extracted answer is compared with the expected value stored in the case definition. The scoring logic in ds4_eval.c (lines 250-280) assigns a pass/fail status per question, flagging any deviation from the reference answer that might indicate model drift or inference bugs.

Result Reporting (lines 300-340)

Finally, ds4_eval.c (lines 300-340) prints a concise table showing the status, token counts, model answer, and ground truth for each question. This output format enables quick visual inspection or automated parsing in CI pipelines.

Running ds4-eval: Practical Examples

The ds4-eval binary supports both interactive debugging and automated regression testing through various command-line options.

Basic Interactive Evaluation

Run the benchmark with the TUI enabled to monitor progress in real time:

./ds4-eval -m ds4flash.gguf --trace /tmp/ds4-eval.txt

This command loads ds4flash.gguf, renders all embedded test cases, and writes a complete generation trace to /tmp/ds4-eval.txt for later analysis. Typical output shows per-question status:


# Question 1  PASSED  tokens=2048  model=B  correct=B

# Question 2  PASSED  tokens=438   model=C  correct=C

# Question 3  FAILED  tokens=2048  model=A  correct=C

Headless CI/CD Integration

For automated pipelines, disable the curses interface and fix random seeds to ensure reproducible results:

./ds4-eval -m ds4flash.gguf --plain \
  --questions 4 --tokens 2048 --temp 0 --seed 1

The --plain flag streams plain text suitable for log files. Parse the results programmatically:

./ds4-eval -m ds4flash.gguf --plain | grep PASSED | wc -l

Re-grading Existing Traces

When modifying answer extraction logic, avoid expensive re-generation by re-running the scorer on stored tokens:

./ds4-eval -m ds4flash.gguf --regrade-trace /tmp/ds4-eval.txt

This command re-executes only the extraction and scoring phases defined in ds4_eval.c (lines 210-280), enabling rapid iteration on grading algorithms.

Testing Specific Question Ranges

Limit evaluation to a subset of questions for quick smoke tests:

./ds4-eval -m ds4flash.gguf --questions 1-10

This runs only the first ten cases from the embedded eval_cases[] array, reducing validation time from minutes to seconds.

Cross-Platform Backend Support

Compile and run the benchmark on any supported hardware backend:


# CUDA multi-GPU

make cuda-spark
./ds4-eval -m ds4flash.gguf --cuda

The same ds4_eval.c source works on Metal, CUDA, and ROCm without modification, ensuring consistent evaluation metrics across deployment targets.

Why Use ds4-eval for Model Validation?

Detects Subtle Drift: Even minor changes to KV-cache eviction policies or quantization kernel selection can alter token distributions. The benchmark catches these regressions before they reach production.

Fast Feedback Loop: A full run completes in seconds on modern GPUs (approximately 10 tokens/second generation), providing immediate validation of inference pipeline changes.

Zero External Dependencies: All benchmark data resides in ds4_eval.c as static arrays. No network calls, no dataset downloads, and no Python environment required.

Extensible Architecture: Add custom regression cases by appending to the eval_cases[] array in ds4_eval.c and recompiling, allowing teams to validate domain-specific capabilities alongside the standard suite.

Summary

  • ds4-eval provides deterministic regression testing for GGUF-based models through a six-stage pipeline implemented in ds4_eval.c.
  • The benchmark covers GPQA Diamond, SuperGPQA, AIME 2025, and COMPSEC datasets to verify hard-science, math, and security coding capabilities.
  • Source code locations: model loading (lines 5-15), prompt rendering (30-44), generation loop (70-95), extraction (210-240), scoring (250-280), and reporting (300-340).
  • Run interactively with --trace or automate with --plain, --questions, and --regrade-trace flags.
  • Supports Metal, CUDA, and ROCm backends through compile-time targets in the repository Makefile.

Frequently Asked Questions

What model formats does ds4-eval support?

ds4-eval exclusively supports GGUF files, the quantized format used by DeepSeek V4 Flash and GLM 5.2 models. The loading logic in ds4_eval.c (lines 5-15) initializes the model context directly from the GGUF binary without requiring external conversion tools.

How long does a typical evaluation take?

A complete run of the default benchmark suite finishes in a few seconds on modern GPU hardware, generating approximately 10 tokens per second. You can reduce this further by using --questions to limit the evaluation to specific subsets for rapid smoke testing.

Can I add custom benchmark cases to ds4-eval?

Yes. The benchmark is extensible by design. Add new cases to the eval_cases[] array defined in ds4_eval.c, then recompile using the standard Makefile targets. This allows teams to incorporate proprietary or domain-specific validation questions alongside the standard GPQA and COMPSEC datasets.

Does ds4-eval require external dependencies?

No. The tool is self-contained within the antirez/ds4 repository. All test questions, ground-truth answers, and scoring logic are embedded as static data structures in ds4_eval.c, eliminating network dependencies, dataset downloads, or external Python environments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →