# How the MTPLX AIME Benchmark Runner Works: Architecture, State Machine, and Prompts

> Explore the MTPLX AIME benchmark runner an OpenAI-compatible state-machine evaluator. Learn how it processes problems and extracts answers using format-contract prompts.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: architecture
- Published: 2026-09-13

---

**The MTPLX AIME benchmark runner is a state-machine-driven evaluator that streams 30 AIME 2026 problems through OpenAI-compatible endpoints, using a family of format-contract prompts to extract answers via `candidate_answer=N` and `\boxed{N}` lines, while ignoring hidden reasoning content for scoring.**

The AIME benchmark runner in the `youssofal/MTPLX` repository provides an end-to-end evaluation framework for testing large language models against the American Invitational Mathematics Examination. Implemented in [`mtplx/benchmarks/runners/aime.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/benchmarks/runners/aime.py), this component orchestrates streaming inference, answer validation, and result persistence through a strict state machine that ensures only one concurrent evaluation runs at a time.

## AIME Benchmark Runner Architecture and Workflow

The runner operates as a self-contained state machine that progresses through distinct lifecycle stages while managing a single active evaluation session.

### State Machine and Concurrency Control

The runner maintains exclusive control over benchmark execution through a simple state machine (`idle → running → paused → done`) with `MAX_PARALLEL = 1`. Only one run may be active at any time; attempts to start additional runs while another is `running` or `paused` return HTTP 409 Conflict.

Finished runs persist in memory for `RETENTION_S` (300 seconds) before garbage collection, allowing clients to query recent results. The registry tracks active run IDs through `list_active_run_id()` and ensures thread-safe state transitions via asyncio primitives defined in [`mtplx/benchmarks/runners/aime.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/benchmarks/runners/aime.py).

### Streaming Evaluation Pipeline

For each of the 30 problems loaded from `mtplx/benchmarks/prompts/aime_2026.jsonl` (default path specified by `DEFAULT_DATASET_PATH`), the runner executes the following stream processing logic:

1. **Problem Loading**: Reads the JSON-Lines dataset into `AIMEProblem` objects containing the mathematical query.
2. **Chat Completion**: Opens an OpenAI-compatible streaming request (`POST /v1/chat/completions?stream=true`) against the local MTPLX server, transmitting a solver prompt combining `SYSTEM_PROMPT` and `USER_PROMPT_SUFFIX`.
3. **Delta Forwarding**: As the model generates tokens, the runner forwards two delta types to an `asyncio.Queue`: `reasoning_content` (hidden scratch space) and `content` (visible output).
4. **Answer Capture**: Only the visible `content` portion is considered the submitted answer; hidden reasoning is explicitly ignored for scoring purposes.
5. **Grading**: The captured text passes to `mtplx.benchmarks.validators.aime` for extraction of `candidate_answer`, verified answers, or `\boxed{}` values, returning a discrete grade status.
6. **Persistence**: Appends a JSON-Lines row to `~/.mtplx/benchmarks/aime/<run_id>.jsonl` for each completed problem.

### Pause, Resume, and Cancel Operations

The state machine supports lifecycle interruptions with specific semantics:

- **Pause**: Aborts the in-flight chat completion, leaves the current question unscored, and transitions to `paused` state awaiting resume.
- **Resume**: Restarts the same question with a fresh two-message prompt, transitioning back to `running`.
- **Cancel**: Terminates the outer asyncio task, propagating `CancelledError` to the streaming response and cleaning up resources.

## Prompt Families and Format Contracts

The runner defines distinct prompt families in [`mtplx/benchmarks/runners/aime.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/benchmarks/runners/aime.py) that establish strict format contracts. All prompts require the model to output two specific lines: `candidate_answer=N` and `\boxed{N}`, where N is the integer answer.

### Standard Solver Prompts

The default evaluation mode uses `SYSTEM_PROMPT` (lines 90-97) and `USER_PROMPT_SUFFIX` (lines 99-102). These instruct the model to place final answers in the exact two-line format while allowing hidden reasoning blocks demarcated by specific tokens. The system message enforces the format contract, while the user suffix reminds the model to close hidden reasoning and follow the line format.

### Fast Mode (Thinking Disabled)

For models operating without chain-of-thought, `FAST_SYSTEM_PROMPT` (lines 103-107) and `FAST_USER_PROMPT_SUFFIX` (lines 109-111) remove hidden reasoning requirements entirely. These prompts instruct the model to terminate output with the two answer lines directly, optimizing for latency over interpretability.

### Verification and Recovery Prompts

Advanced evaluation modes employ specialized verifiers and recovery agents:

- **Verifier (Answer-Only)**: `VERIFIER_ANSWER_ONLY_SYSTEM_PROMPT` and `VERIFIER_ANSWER_ONLY_USER_PROMPT_SUFFIX` engage a second model to check candidate answers, producing compact derivations ending with `verified_answer=N`.
- **Tie-Breaker**: `VERIFIER_TIEBREAKER_SYSTEM_PROMPT` provides stricter verification that never writes `candidate_answer`, resolving ambiguous cases.
- **Adjudicator**: `VERIFIER_ADJUDICATOR_SYSTEM_PROMPT` selects the correct integer when multiple candidate answers exist, or corrects erroneous submissions.
- **Cap Recovery**: `CAP_RECOVERY_FINALIZER_SYSTEM_PROMPT` and `CAP_RECOVERY_FRESH_SYSTEM_PROMPT` handle token-capped attempts by either continuing the existing trajectory or starting fresh visible recovery.
- **Visible Submission**: `VISIBLE_SUBMISSION_SYSTEM_PROMPT` continues reasoning-enabled attempts that previously omitted visible answers.

## Programmatic Control and Testing

The runner exposes a public API for integration and testing:

```python

# Start a new evaluation run

from mtplx.benchmarks.runners.aime import start_run, list_active_run_id

run_id = start_run()
print(f"Initiated AIME benchmark run: {run_id}")

# Poll for completion

import time
while run_id in list_active_run_id():
    time.sleep(1)
print("Benchmark complete")

```

For testing, the constructor accepts an optional `chat_stream_factory` callable, enabling injection of fake event streams without contacting a real LLM server.

## Summary

- The **AIME benchmark runner** implements a strict state machine with `MAX_PARALLEL = 1` to ensure sequential evaluation of mathematical problems.
- **Streaming architecture** separates hidden `reasoning_content` from visible `content`, grading only the latter via [`mtplx/benchmarks/validators/aime.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/benchmarks/validators/aime.py).
- **Format contracts** require exact output lines (`candidate_answer=N` and `\boxed{N}`) enforced through multiple prompt families including standard, fast, verifier, and recovery variants.
- **Lifecycle management** supports pause, resume, and cancel operations with HTTP 409 protection against concurrent runs and 5-minute retention of completed runs.
- **Persistence** writes results to `~/.mtplx/benchmarks/aime/<run_id>.jsonl` while loading problems from `mtplx/benchmarks/prompts/aime_2026.jsonl`.

## Frequently Asked Questions

### How does the runner distinguish between hidden reasoning and visible answers?

The runner parses the streaming response into two distinct delta types: `reasoning_content` (hidden scratch space) and `content` (visible output). As implemented in [`mtplx/benchmarks/runners/aime.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/benchmarks/runners/aime.py), only the visible `content` is captured for grading, while `reasoning_content` is forwarded to the client but excluded from answer validation.

### What happens when I pause and resume a benchmark run?

When paused, the runner aborts the current in-flight chat completion and leaves the active question unscored, transitioning to a `paused` state. Upon resume, it restarts the same question with a fresh two-message prompt, allowing the model to retry without penalizing previous attempts. Cancel operations terminate the outer asyncio task entirely.

### Where are the AIME problems stored and how are results persisted?

The runner loads the 30-problem dataset from `mtplx/benchmarks/prompts/aime_2026.jsonl` (configurable via `DEFAULT_DATASET_PATH`). Results are written as JSON-Lines to `~/.mtplx/benchmarks/aime/<run_id>.jsonl`, with each row containing the problem ID, captured answer, and grade status.

### Can I customize the prompts for different model capabilities?

Yes. The runner selects from multiple prompt families defined in [`mtplx/benchmarks/runners/aime.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/benchmarks/runners/aime.py), including `SYSTEM_PROMPT` for standard reasoning, `FAST_SYSTEM_PROMPT` for thinking-disabled mode, and various verifier prompts. These constants are importable and overridable when instantiating the runner directly, or you can provide a custom `chat_stream_factory` for complete control over the inference pipeline.