How to Run Reproducible Benchmark Experiments with Maka Eval: A Complete Guide
Maka Eval enables deterministic benchmarking through JSON experiment specifications, a CLI driver, and immutable attempt logs stored in the output directory.
Running reproducible benchmark experiments with Maka Eval provides a deterministic way to measure performance across Maka itself and external subjects. The Apache Maka repository implements a three-layer architecture that guarantees repeatability through append-only result logs and containerized execution environments. This guide walks through the end-to-end workflow from specification to result analysis, referencing the actual source implementation.
Understanding the Maka Eval Architecture
The reproducibility of Maka Eval stems from its strict separation of concerns across three logical layers. Each layer is implemented in specific source files within the packages/eval workspace.
The Experiment Specification Layer
At the core of every benchmark is the ExperimentSpec, a JSON-encoded schema that declaratively describes the entire experiment. Defined in packages/eval/src/spec.ts, this specification includes the executor configuration, subject definitions, task lists, repetition counts, and resource constraints.
The specification uses the "schemaVersion": "maka.eval.v1" identifier to ensure forward compatibility. By externalizing all parameters into this immutable JSON file, Maka Eval ensures that the same specification file produces identical experimental conditions across different machines and time periods.
The CLI Driver Layer
The runMakaEvalCli function in packages/eval/src/cli.ts serves as the orchestration entry point. This driver parses the specification, validates environment prerequisites, and prepares the experiment directory structure. Around line 70 in cli.ts, the CLI performs strict validation, refusing to start if required executor binaries, Docker, or environment variables are missing.
The CLI selects between Harbor and Pier executors based on the executor.kind field in the specification, then wires subject adapters and streams the experiment into the Runtime Host.
The Runtime Execution Layer
The runExperiment function in packages/eval/src/runner.ts orchestrates the actual benchmark execution. This layer handles cell creation (the Cartesian product of tasks × repetitions × subjects), manages attempt lifecycles, and collects results.
Crucially, the runtime writes results to an append-only attempt log inside the output directory. This immutability guarantee means that once a cell completes, its result kernel—containing score, usage, cost, duration, and status—cannot be altered, ensuring reproducible benchmarking.
Prerequisites and Environment Setup
Before running experiments, ensure your environment meets the requirements checked by the CLI validation logic.
git clone https://github.com/apache/maka.git
cd maka
npm ci # install dependencies
npm run build # compile all workspaces
The CLI is part of the @maka/cli package. After building, the maka eval commands become available from the repository root. The validation logic in cli.ts (lines 78-80) checks for credential environment variables and executor binaries before execution begins.
Creating the Experiment Specification
Construct a JSON specification file that declares your benchmark parameters. The ExperimentSpec format is documented in the @maka/eval package README at packages/eval/README.md.
Here is a minimal example (experiment.json):
{
"schemaVersion": "maka.eval.v1",
"id": "terminal-bench-2.1",
"executor": { "kind": "harbor", "config": {} },
"subjects": [
{ "kind": "maka", "name": "maka", "credentials": [] },
{ "kind": "external", "name": "deepseek-harness", "credentials": [] }
],
"tasks": [{ "id": "task-1", "benchmark": "terminal-bench" }],
"repetitions": 3
}
This specification defines the experiment matrix: 1 task × 3 repetitions × 2 subjects = 6 total cells. The executor.kind field selects the container runtime, while the subjects array declares both internal Maka instances and external harnesses to benchmark against.
Executing Benchmark Experiments
Run the specification through the CLI to execute the full experiment matrix:
maka eval run experiment.json --out .maka-eval/run-001
The openExperimentDirectory function (in experiment-directory.ts) creates the output structure, including an attempts subdirectory for the immutable logs. The CLI validates the specification schema and environment variables before spawning the Runtime Host.
Upon completion, the CLI prints a JSON summary to stdout:
{"experimentId":"terminal-bench-2.1","cells":18,"incomplete":0}
The cells field reports the total number of Cartesian product combinations processed. If incomplete is greater than zero, indicating cells that did not produce results, the CLI exits with a non-zero status code.
Analyzing and Reproducing Results
All artifacts are stored under the output directory specified by --out. The directory structure follows this pattern:
.maka-eval/run-001/
├─ attempts/
│ ├─ cell-0.json
│ ├─ cell-1.json
│ └─ …
└─ experiment.json # copy of the original spec
Each file in attempts/ contains an immutable result kernel with the complete execution context. Because these logs are append-only and the specification pins all variables (Docker images, subject versions, task definitions), rerunning the exact same command on a machine with identical executor images yields bit-for-bit reproducible results.
Parse results programmatically using Node.js:
import { readFile } from 'node:fs/promises';
import { join } from 'node:path';
const outDir = '.maka-eval/run-001';
const summary = JSON.parse(await readFile(join(outDir, 'summary.json'), 'utf8'));
console.log('Completed cells:', summary.cells);
Advanced Configuration Options
Replaying Failed Cells
When specific cells fail, avoid rerunning the entire matrix by targeting individual cells:
maka eval run experiment.json --out .maka-eval/run-001 --cell cell-3
The --cell argument (parsed around lines 33-35 in cli.ts) instructs the runtime to execute only the specified cell ID, saving time during debugging.
Custom Executors and Credential Injection
Implement a custom HarnessExecutor and reference it in the specification's executor.kind field. The CLI invokes builtinExecutor (lines 11-15 in cli.ts) to load the appropriate implementation. Subject adapters receive secret-free environment variable names, which the CLI pre-validates and passes to the executor's preflight step.
Parallel Execution Control
The Runtime Host respects concurrency limits defined in the specification's task and repetition configuration. This guarantees controlled parallelism while maintaining deterministic ordering of the append-only attempt logs.
For a deeper architectural view of how these components interact, consult ARCHITECTURE.md in the repository root.
Summary
- Maka Eval provides reproducible benchmarking through immutable JSON specifications and append-only result logs.
- The three-layer architecture (specification, CLI driver, runtime) is implemented in
packages/eval/src/spec.ts,packages/eval/src/cli.ts, andpackages/eval/src/runner.ts. - Execute experiments via
maka eval run <spec>.json --out <dir>after building the workspace withnpm run build. - Results are stored in the
attempts/subdirectory as immutable JSON files, enabling exact reproduction when rerun with identical specifications and executor images. - Use
--cell <id>to rerun specific failed cells without reprocessing the entire experiment matrix.
Frequently Asked Questions
What file format does Maka Eval use for experiment specifications?
Maka Eval uses a JSON format conforming to the maka.eval.v1 schema. The ExperimentSpec type is defined in packages/eval/src/spec.ts and includes fields for executor, subjects, tasks, and repetitions. This JSON file serves as the single source of truth for all experimental parameters.
How does Maka Eval ensure reproducibility across different runs?
Reproducibility is enforced through three mechanisms: deterministic ExperimentSpec files that pin all variables, containerized execution environments (Harbor or Pier executors) that ensure consistent runtime conditions, and append-only attempt logs stored in the output directory. The runExperiment function in packages/eval/src/runner.ts writes results immutably, preventing modification of historical data.
Can I rerun only specific failed cells without repeating the entire experiment?
Yes. Pass the --cell <cell-id> flag to the CLI command. The argument parsing logic around lines 33-35 in packages/eval/src/cli.ts handles this flag, instructing the Runtime Host to execute only the specified cell rather than the full Cartesian product matrix.
Where does Maka Eval store benchmark results and logs?
Results are stored in the output directory specified by the --out flag. The openExperimentDirectory function creates an attempts/ subdirectory containing individual JSON files for each cell (e.g., cell-0.json), plus a copy of the original experiment.json specification. Each attempt file contains the complete result kernel including metrics like score, usage, cost, and duration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →