How the MTPLX AIME Benchmark Runner Works: Architecture, State Machine, and Prompts
The MTPLX AIME benchmark runner is a state-machine-driven evaluator that streams 30 AIME 2026 problems through OpenAI-compatible endpoints, using a family of format-contract prompts to extract answers via candidate_answer=N and \boxed{N} lines, while ignoring hidden reasoning content for scoring.
The AIME benchmark runner in the youssofal/MTPLX repository provides an end-to-end evaluation framework for testing large language models against the American Invitational Mathematics Examination. Implemented in mtplx/benchmarks/runners/aime.py, this component orchestrates streaming inference, answer validation, and result persistence through a strict state machine that ensures only one concurrent evaluation runs at a time.
AIME Benchmark Runner Architecture and Workflow
The runner operates as a self-contained state machine that progresses through distinct lifecycle stages while managing a single active evaluation session.
State Machine and Concurrency Control
The runner maintains exclusive control over benchmark execution through a simple state machine (idle → running → paused → done) with MAX_PARALLEL = 1. Only one run may be active at any time; attempts to start additional runs while another is running or paused return HTTP 409 Conflict.
Finished runs persist in memory for RETENTION_S (300 seconds) before garbage collection, allowing clients to query recent results. The registry tracks active run IDs through list_active_run_id() and ensures thread-safe state transitions via asyncio primitives defined in mtplx/benchmarks/runners/aime.py.
Streaming Evaluation Pipeline
For each of the 30 problems loaded from mtplx/benchmarks/prompts/aime_2026.jsonl (default path specified by DEFAULT_DATASET_PATH), the runner executes the following stream processing logic:
- Problem Loading: Reads the JSON-Lines dataset into
AIMEProblemobjects containing the mathematical query. - Chat Completion: Opens an OpenAI-compatible streaming request (
POST /v1/chat/completions?stream=true) against the local MTPLX server, transmitting a solver prompt combiningSYSTEM_PROMPTandUSER_PROMPT_SUFFIX. - Delta Forwarding: As the model generates tokens, the runner forwards two delta types to an
asyncio.Queue:reasoning_content(hidden scratch space) andcontent(visible output). - Answer Capture: Only the visible
contentportion is considered the submitted answer; hidden reasoning is explicitly ignored for scoring purposes. - Grading: The captured text passes to
mtplx.benchmarks.validators.aimefor extraction ofcandidate_answer, verified answers, or\boxed{}values, returning a discrete grade status. - Persistence: Appends a JSON-Lines row to
~/.mtplx/benchmarks/aime/<run_id>.jsonlfor each completed problem.
Pause, Resume, and Cancel Operations
The state machine supports lifecycle interruptions with specific semantics:
- Pause: Aborts the in-flight chat completion, leaves the current question unscored, and transitions to
pausedstate awaiting resume. - Resume: Restarts the same question with a fresh two-message prompt, transitioning back to
running. - Cancel: Terminates the outer asyncio task, propagating
CancelledErrorto the streaming response and cleaning up resources.
Prompt Families and Format Contracts
The runner defines distinct prompt families in mtplx/benchmarks/runners/aime.py that establish strict format contracts. All prompts require the model to output two specific lines: candidate_answer=N and \boxed{N}, where N is the integer answer.
Standard Solver Prompts
The default evaluation mode uses SYSTEM_PROMPT (lines 90-97) and USER_PROMPT_SUFFIX (lines 99-102). These instruct the model to place final answers in the exact two-line format while allowing hidden reasoning blocks demarcated by specific tokens. The system message enforces the format contract, while the user suffix reminds the model to close hidden reasoning and follow the line format.
Fast Mode (Thinking Disabled)
For models operating without chain-of-thought, FAST_SYSTEM_PROMPT (lines 103-107) and FAST_USER_PROMPT_SUFFIX (lines 109-111) remove hidden reasoning requirements entirely. These prompts instruct the model to terminate output with the two answer lines directly, optimizing for latency over interpretability.
Verification and Recovery Prompts
Advanced evaluation modes employ specialized verifiers and recovery agents:
- Verifier (Answer-Only):
VERIFIER_ANSWER_ONLY_SYSTEM_PROMPTandVERIFIER_ANSWER_ONLY_USER_PROMPT_SUFFIXengage a second model to check candidate answers, producing compact derivations ending withverified_answer=N. - Tie-Breaker:
VERIFIER_TIEBREAKER_SYSTEM_PROMPTprovides stricter verification that never writescandidate_answer, resolving ambiguous cases. - Adjudicator:
VERIFIER_ADJUDICATOR_SYSTEM_PROMPTselects the correct integer when multiple candidate answers exist, or corrects erroneous submissions. - Cap Recovery:
CAP_RECOVERY_FINALIZER_SYSTEM_PROMPTandCAP_RECOVERY_FRESH_SYSTEM_PROMPThandle token-capped attempts by either continuing the existing trajectory or starting fresh visible recovery. - Visible Submission:
VISIBLE_SUBMISSION_SYSTEM_PROMPTcontinues reasoning-enabled attempts that previously omitted visible answers.
Programmatic Control and Testing
The runner exposes a public API for integration and testing:
# Start a new evaluation run
from mtplx.benchmarks.runners.aime import start_run, list_active_run_id
run_id = start_run()
print(f"Initiated AIME benchmark run: {run_id}")
# Poll for completion
import time
while run_id in list_active_run_id():
time.sleep(1)
print("Benchmark complete")
For testing, the constructor accepts an optional chat_stream_factory callable, enabling injection of fake event streams without contacting a real LLM server.
Summary
- The AIME benchmark runner implements a strict state machine with
MAX_PARALLEL = 1to ensure sequential evaluation of mathematical problems. - Streaming architecture separates hidden
reasoning_contentfrom visiblecontent, grading only the latter viamtplx/benchmarks/validators/aime.py. - Format contracts require exact output lines (
candidate_answer=Nand\boxed{N}) enforced through multiple prompt families including standard, fast, verifier, and recovery variants. - Lifecycle management supports pause, resume, and cancel operations with HTTP 409 protection against concurrent runs and 5-minute retention of completed runs.
- Persistence writes results to
~/.mtplx/benchmarks/aime/<run_id>.jsonlwhile loading problems frommtplx/benchmarks/prompts/aime_2026.jsonl.
Frequently Asked Questions
How does the runner distinguish between hidden reasoning and visible answers?
The runner parses the streaming response into two distinct delta types: reasoning_content (hidden scratch space) and content (visible output). As implemented in mtplx/benchmarks/runners/aime.py, only the visible content is captured for grading, while reasoning_content is forwarded to the client but excluded from answer validation.
What happens when I pause and resume a benchmark run?
When paused, the runner aborts the current in-flight chat completion and leaves the active question unscored, transitioning to a paused state. Upon resume, it restarts the same question with a fresh two-message prompt, allowing the model to retry without penalizing previous attempts. Cancel operations terminate the outer asyncio task entirely.
Where are the AIME problems stored and how are results persisted?
The runner loads the 30-problem dataset from mtplx/benchmarks/prompts/aime_2026.jsonl (configurable via DEFAULT_DATASET_PATH). Results are written as JSON-Lines to ~/.mtplx/benchmarks/aime/<run_id>.jsonl, with each row containing the problem ID, captured answer, and grade status.
Can I customize the prompts for different model capabilities?
Yes. The runner selects from multiple prompt families defined in mtplx/benchmarks/runners/aime.py, including SYSTEM_PROMPT for standard reasoning, FAST_SYSTEM_PROMPT for thinking-disabled mode, and various verifier prompts. These constants are importable and overridable when instantiating the runner directly, or you can provide a custom chat_stream_factory for complete control over the inference pipeline.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →