How to Set Up Reproducible Benchmark Experiments with Maka's Eval Framework
Maka's Eval framework provides a declarative, three-phase workflow for reproducible benchmark experiments: define a validated JSON spec, run via CLI, and collect deterministic results from immutable attempt records.
Setting up reproducible benchmark experiments is critical for engineering teams evaluating AI agents. Maka's Eval framework, located in the apache/maka repository, eliminates hidden state and non-determinism through strict schema validation, pinned executor environments, and append-only result logging. This guide walks through the complete setup based on the actual source implementation.
Understanding the Eval Architecture
Maka's Eval framework operates as a thin semantic layer above the Runtime Host, which serves as the sole execution authority. The @maka/eval package owns experiment semantics only—cells, attempts, result selection, budgets, and verifier configuration—while executor adapters like Harbor and Pier translate generic experiments into containerized runs.
The interaction flow is documented in ARCHITECTURE.md, where the Eval layer feeds cells to the Runtime Host for execution. Each cell represents a unique combination of task × repetition × subject, expanded as a Cartesian product by the CLI.
Core Components
- Runtime Host – Manages sessions, tools, and event logs; the only component with execution authority.
- @maka/eval – Handles experiment semantics without direct execution capability.
- Executor adapters (Harbor, Pier) – Convert abstract experiments into concrete Docker runs.
- Subject types –
makasubjects (built-in benchmark agents) andexternalsubjects (arbitrary CLI programs).
The Three-Phase Reproducibility Workflow
Phase 1: Define and Validate the Experiment Spec
The experiment spec is a JSON document describing every parameter of your benchmark. The packages/eval/src/spec.ts module enforces strict schema validation including required fields, identifier format, and positive integers for numeric values.
A minimal valid spec requires:
schemaVersion– Use"maka.eval.v1"id– Unique experiment identifierbenchmark.idandbenchmark.version– Exact benchmark releaseexecutor.kindandexecutor.config– Fixed Docker image with tagsubjects[]– Subject configurations with credential referencestasks[]– Input prompts or datarepetitions– Integer count (e.g.,3)
The validator in spec.ts rejects malformed specs before execution begins. You can pre-validate without running:
node -e "import('../packages/eval/src/spec.js').parseExperimentSpec(require('./my-spec.json'))"
Phase 2: Execute via the CLI
The public CLI entry point lives in packages/eval/src/cli.ts and exposes the maka eval run command:
maka eval run <spec> --out <directory>
The CLI performs pre-flight checks for paths and Docker availability, then expands your spec into cells. Each cell executes through the Runtime Host, producing immutable attempt records. Usage help is printed from cli.ts line 34.
Example execution:
# Run experiment and store results
maka eval run my-spec.json --out .maka-eval/run-001
Phase 3: Collect Deterministic Results
Each cell produces an immutable attempt record containing score, usage, cost, duration, status, and artifact digests. The earliest valid attempt becomes the authoritative result. Because attempts are append-only, you can re-run experiments and compare kernel digests for exact verification.
Step-by-Step Setup Instructions
Follow these steps to achieve fully reproducible benchmark experiments with Maka's Eval framework:
-
Pin the repository version
git clone https://github.com/apache/maka.git cd maka git checkout v0.1.0 # Use specific tag or commit -
Install dependencies (Node 22+ required)
npm ci -
Prepare executor toolchain (for Harbor/Pier)
Docker images are version-pinned, e.g.,
maka-eval-egress-proxy:12.2.3. Seepackages/eval/README.mdfor Harbor and Pier sections. -
Create a deterministic spec JSON
Include exact versions, fixed image digests, and explicit task inputs. Store actual secret values locally via environment variables referenced in
credentials[], not in version control. -
Validate the spec
Use the spec module validator or rely on CLI validation at launch.
-
Run the experiment
maka eval run my-spec.json --out .maka-eval/run-001 -
Re-run for verification
Execute identical command with fresh output directory. Matching kernel digests confirm reproducibility.
Complete Example: Minimal Experiment Spec
{
"schemaVersion": "maka.eval.v1",
"id": "demo-exp",
"benchmark": {
"id": "terminal-bench",
"version": "2.1",
"config": {}
},
"executor": {
"kind": "harbor",
"config": {
"image": "docker.io/maka/harbor:12.2.3"
}
},
"subjects": [
{
"id": "maka-subject",
"kind": "maka",
"credentials": [],
"config": {}
}
],
"tasks": [
{
"id": "task-1",
"input": "Explain the repository structure.",
"config": {}
}
],
"repetitions": 3,
"budget": {},
"verifier": {}
}
Inspect and Verify Results
After execution, examine the attempt records:
# List all attempt directories
ls .maka-eval/run-001/attempts/
# View formatted result kernel
cat .maka-eval/run-001/attempts/0/result.json | jq .
Each result kernel contains deterministic fields: score, usage, cost, duration, status, and artifact digests. Comparing these across runs proves reproducibility.
Why This Guarantees Reproducibility
- Declarative spec – All parameters explicit in JSON; no hidden state
- Pinned executor images – Docker images identified by digest, preventing accidental upgrades
- Immutable attempt logs – Append-only records; earliest valid attempt is canonical
- Runtime Host sandbox – Tools execute in controlled environment; network access blocked by egress proxy per
README.mdline 70
Key Source Files
| File | Purpose |
|---|---|
packages/eval/src/spec.ts |
JSON schema validation and experiment spec parsing |
packages/eval/src/cli.ts |
CLI implementation with maka eval run command |
ARCHITECTURE.md |
High-level architecture showing Eval-Runtime Host interaction |
packages/eval/README.md |
Framework overview, CLI usage, executor configuration |
Summary
- Maka's Eval framework enables reproducible benchmark experiments through a declarative JSON spec, validated CLI execution, and immutable result logging.
- Pin repository versions, executor images, and task inputs to eliminate variability.
- Validate specs via
packages/eval/src/spec.tsbefore execution. - Run experiments with
maka eval run <spec> --out <dir>and verify by comparing kernel digests across re-runs. - The Runtime Host sandbox and egress proxy block external network access, ensuring identical execution environments.
Frequently Asked Questions
What makes Maka's Eval framework different from other benchmark tools?
Unlike tools that rely on procedural scripts with hidden state, Maka's Eval uses a strictly validated JSON spec where all parameters are explicit. The spec.ts validator enforces schema compliance before execution, and the Runtime Host sandbox guarantees identical execution environments through containerized, network-isolated runs.
How do I prevent credential leakage in version-controlled specs?
Store sensitive values as environment variable names in subjects[].credentials, not literal values. The spec references variable names; actual secrets remain local. This pattern keeps specs shareable and version-controlled while protecting credentials.
Can I reproduce results on a different machine?
Yes, provided you use identical inputs: same repository commit, same Node version (22+), same executor Docker image digest, and same spec. The framework is designed for cross-machine reproducibility—the Runtime Host abstracts hardware differences through containerization.
What happens if an attempt fails or times out?
Failed attempts are logged as immutable records with status fields indicating failure reason. The earliest valid attempt selection logic ignores failed attempts. Re-running the same spec will retry cells, producing new attempts without modifying previous logs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →