How to Run ds4-eval and Interpret GPQA, SuperGPQA, and AIME Results

To run ds4-eval, execute ./ds4-eval -m model.gguf against a DeepSeek V4 Flash or compatible GGUF file; the tool grades 92 embedded test cases and outputs a concise PASSED/FAILED report with a fractional score for each GPQA, Super-GPQA, and AIME question.

The ds4-eval utility in the antirez/ds4 repository provides a self-contained benchmark suite for evaluating large language models on graduate-level reasoning tasks. Unlike external evaluation frameworks, this tool loads a local GGUF model, renders prompts from the official evaluation suite defined in ds4_eval.c, and streams token-by-token generation through a split-screen terminal interface before applying rigorous domain-specific grading logic.

Understanding the ds4-eval Benchmark Suite

The evaluator ships with 92 embedded test cases hardcoded as an array of eval_case structs at the top of ds4_eval.c (lines 85-94). These cases are partitioned into four distinct domains designed to stress-test scientific reasoning, specialized knowledge, mathematical precision, and code security analysis.

Test Case Composition

According to the repository documentation in README.md (lines 73-84), the first 75 cases consist of balanced subsets from three primary benchmarks, while the remaining 17 focus on computer security:

  • 25 GPQA Diamond questions: Graduate-level science multiple-choice questions requiring deep domain expertise
  • 25 Super-GPQA questions: Curated specialist-knowledge items audited for key correctness across fields like Agronomy, Physics, and Medicine
  • 25 AIME 2025 problems: Exact-answer mathematics competitions where a single integer answer is required
  • 17 COMPSEC cases: Single-function C/C++ security-bug localization tasks

Each eval_case struct contains fields for source, id, domain, title, question, an array of choice strings (for multiple-choice domains), and the ground-truth answer. For example, a Super-GPQA entry appears as:

{
    .source = "SuperGPQA",
    .id = "001b51d76b4d422988f2c11f104a2c6c",
    .domain = "Agronomy",
    .title = "In production, grass powder is often pressed into pellets , and",
    .question = "In production, grass powder is often pressed into pellets , and crushed hay is processed into ____, grass cake.",
    .choice[0] = "a sheet of paper",
    .choice[1] = "a box of flowers",
    .choice[2] = "a block of grass",
    .choice[3] = "a pot of mud",
    .choice[4] = "a pile of dirt",
    .choice[5] = "a bush of herbs",
    .choice[6] = "a roll of hay",
    .choice[7] = "a stack of leaves",
    .choice[8] = "a sack of grain",
    .choice[9] = "a bundle of sticks",
    .answer = "C",
},

Running ds4-eval from the Command Line

The basic invocation requires specifying a GGUF model path with the -m flag:

./ds4-eval -m ds4flash.gguf

By default, the evaluator runs in thinking mode with a generation budget of 16,000 tokens and displays an interactive TUI (Terminal User Interface) showing the model's reasoning stream alongside the final answer extraction.

Essential Command-Line Flags

For reproducible benchmarking and scripting, the following flags control execution behavior:

Flag Description
--plain Disable the interactive TUI; output plain text suitable for log files
--questions N Run only the first N test cases (useful for validation)
--tokens T Override the default 16,000 token generation budget per case
--temp V Set sampling temperature (use 0 for deterministic output)
--seed S Fix the random seed for reproducible generations
--trace FILE Write detailed execution traces for later replay
--regrade-trace FILE Re-evaluate a previous run without recomputing forward passes
--power P Throttle GPU utilization to P percent

To reproduce the reference baseline shown in the project documentation, execute:

./ds4-eval -m ds4flash.gguf --plain --questions 4 --tokens 2048 --temp 0 --seed 1

This command runs deterministically on the first four cases with limited token budget, matching the output table format documented in the repository's README.

Interpreting Evaluation Results

During execution, ds4-eval prints a per-question status line immediately after grading each case. The output format follows this structure:


[ 1/92] PASSED GPQA_Diamond/recNu3MXkvWUzHZr9: 2048  score 1/1  failed 0/0
[ 2/92] PASSED SuperGPQA/001b51d76b4d422988f2c11f104a2c6c: 438  score 1/1  failed 0/0
[ 3/92] PASSED AIME2025/aime2025-01: 666  score 1/1  failed 0/0
[ 4/92] FAILED SuperGPQA/4a1d1780a93f4093b6fb7d3c314cbea8: 2048  score 0/1  failed 1/1

Decoding the Per-Question Report

Each status line contains critical metadata:

  • PASSED / FAILED: Indicates whether the extracted answer matches the ground truth stored in the eval_case struct
  • score X/Y: The fractional score for that specific question (1/1 for correct, 0/1 for incorrect)
  • failed X/Y: Mirrors the score but used for aggregated failure statistics across the full suite
  • Token count: The number following the case ID shows tokens consumed for that generation

The reference implementation that prints this summary resides in ds4_eval.c around line 2029, where the final evaluation line is formatted and output to the terminal.

Domain-Specific Grading Logic

The answer extraction mechanism, implemented in ds4_eval.c (lines 2666-2671), parses the model's output stream to locate the first token matching a known answer format. However, grading rules differ significantly by domain:

GPQA and Super-GPQA: These multiple-choice domains require the model to output a letter between A and J. The evaluator extracts the first valid letter token and performs exact string comparison against the answer field (e.g., "C"). Any deviation results in a FAILED status with score 0/1.

AIME 2025: Mathematics problems require an exact integer answer. The extractor identifies the first integer token in the output stream and checks for strict equality with the expected numeric string. Because AIME offers no partial credit, even off-by-one errors or extra whitespace cause immediate failure with score 0/1.

Advanced Workflows: Tracing and Re-grading

When debugging grading logic or iterating on answer extraction algorithms, you can avoid recomputing expensive forward passes by utilizing the trace system. First, capture a detailed execution trace:

./ds4-eval -m ds4flash.gguf --trace /tmp/ds4-eval.txt

After modifying the extraction logic in ds4_eval.c, re-evaluate the stored trace without regenerating tokens:

./ds4-eval -m ds4flash.gguf --regrade-trace /tmp/ds4-eval.txt

This workflow is essential for verifying fixes to the parsing functions located near line 2666 without waiting for model inference.

Summary

  • Execution: Run ./ds4-eval -m model.gguf to benchmark a local GGUF model against 92 embedded test cases spanning GPQA Diamond, Super-GPQA, AIME 2025, and COMPSEC domains.
  • Output Format: Each case reports PASSED or FAILED with a score X/Y metric; GPQA and Super-GPQA test multiple-choice letter extraction while AIME requires exact integer matching.
  • Reproducibility: Use --temp 0 --seed N for deterministic outputs and --trace combined with --regrade-trace to iterate on grading logic without re-running model inference.
  • Source References: Test cases are defined in ds4_eval.c (lines 85-94), grading logic resides near line 2029, and answer extraction functions are located around lines 2666-2671.

Frequently Asked Questions

What model format does ds4-eval require?

The tool requires a GGUF format model file, typically produced by quantization tools like llama.cpp. The repository includes a download_model.sh helper script to fetch compatible DeepSeek V4 Flash or GLM 5.2 models. Pass the local path to this file using the -m flag.

How does ds4-eval handle ambiguous answers in multiple-choice questions?

For GPQA and Super-GPQA domains, the evaluator extracts the first valid letter token (A-J) from the model's output stream. If the model generates multiple letters or explanatory text before the answer, the extraction logic in ds4_eval.c (lines 2666-2671) locates the first occurrence. Only exact matches against the ground truth answer field register as correct; there is no tolerance for near-misses or alternative formatting.

Can I run a subset of the benchmark for quick testing?

Yes. Append the --questions N flag to limit execution to the first N cases. For example, ./ds4-eval -m model.gguf --questions 10 runs only the first 10 test cases from the embedded suite, allowing rapid validation of model loading and grading logic without consuming time on the full 92-case battery.

Why does AIME scoring require exact integer matching?

The AIME (American Invitational Mathematics Examination) competition format strictly requires a single three-digit integer answer. The ds4-eval implementation enforces this standard by parsing the first integer token from the output and checking for string equality with the expected answer. This strict validation prevents false positives from nearby numerical values and maintains consistency with the official competition scoring rules where partial credit is not awarded.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →