# How to Run Reproducible Benchmark Experiments with Maka Eval: A Complete Guide

> Learn to run reproducible benchmark experiments with Maka Eval. This guide covers deterministic benchmarking using JSON specs, a CLI, and immutable logs for reliable results.

- Repository: [The Apache Software Foundation/maka](https://github.com/apache/maka)
- Tags: how-to-guide
- Published: 2026-08-24

---

**Maka Eval enables deterministic benchmarking through JSON experiment specifications, a CLI driver, and immutable attempt logs stored in the output directory.**

Running reproducible benchmark experiments with Maka Eval provides a deterministic way to measure performance across Maka itself and external subjects. The Apache Maka repository implements a three-layer architecture that guarantees repeatability through append-only result logs and containerized execution environments. This guide walks through the end-to-end workflow from specification to result analysis, referencing the actual source implementation.

## Understanding the Maka Eval Architecture

The reproducibility of Maka Eval stems from its strict separation of concerns across three logical layers. Each layer is implemented in specific source files within the `packages/eval` workspace.

### The Experiment Specification Layer

At the core of every benchmark is the **ExperimentSpec**, a JSON-encoded schema that declaratively describes the entire experiment. Defined in [`packages/eval/src/spec.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/spec.ts), this specification includes the executor configuration, subject definitions, task lists, repetition counts, and resource constraints.

The specification uses the `"schemaVersion": "maka.eval.v1"` identifier to ensure forward compatibility. By externalizing all parameters into this immutable JSON file, Maka Eval ensures that the same specification file produces identical experimental conditions across different machines and time periods.

### The CLI Driver Layer

The `runMakaEvalCli` function in [`packages/eval/src/cli.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/cli.ts) serves as the orchestration entry point. This driver parses the specification, validates environment prerequisites, and prepares the experiment directory structure. Around line 70 in [`cli.ts`](https://github.com/apache/maka/blob/main/cli.ts), the CLI performs strict validation, refusing to start if required executor binaries, Docker, or environment variables are missing.

The CLI selects between **Harbor** and **Pier** executors based on the `executor.kind` field in the specification, then wires subject adapters and streams the experiment into the Runtime Host.

### The Runtime Execution Layer

The `runExperiment` function in [`packages/eval/src/runner.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/runner.ts) orchestrates the actual benchmark execution. This layer handles cell creation (the Cartesian product of tasks × repetitions × subjects), manages attempt lifecycles, and collects results.

Crucially, the runtime writes results to an **append-only attempt log** inside the output directory. This immutability guarantee means that once a cell completes, its result kernel—containing `score`, `usage`, `cost`, `duration`, and `status`—cannot be altered, ensuring reproducible benchmarking.

## Prerequisites and Environment Setup

Before running experiments, ensure your environment meets the requirements checked by the CLI validation logic.

```bash
git clone https://github.com/apache/maka.git
cd maka
npm ci                # install dependencies

npm run build         # compile all workspaces

```

The CLI is part of the `@maka/cli` package. After building, the `maka eval` commands become available from the repository root. The validation logic in [`cli.ts`](https://github.com/apache/maka/blob/main/cli.ts) (lines 78-80) checks for credential environment variables and executor binaries before execution begins.

## Creating the Experiment Specification

Construct a JSON specification file that declares your benchmark parameters. The `ExperimentSpec` format is documented in the `@maka/eval` package README at [`packages/eval/README.md`](https://github.com/apache/maka/blob/main/packages/eval/README.md).

Here is a minimal example ([`experiment.json`](https://github.com/apache/maka/blob/main/experiment.json)):

```json
{
  "schemaVersion": "maka.eval.v1",
  "id": "terminal-bench-2.1",
  "executor": { "kind": "harbor", "config": {} },
  "subjects": [
    { "kind": "maka", "name": "maka", "credentials": [] },
    { "kind": "external", "name": "deepseek-harness", "credentials": [] }
  ],
  "tasks": [{ "id": "task-1", "benchmark": "terminal-bench" }],
  "repetitions": 3
}

```

This specification defines the **experiment matrix**: 1 task × 3 repetitions × 2 subjects = 6 total cells. The `executor.kind` field selects the container runtime, while the `subjects` array declares both internal Maka instances and external harnesses to benchmark against.

## Executing Benchmark Experiments

Run the specification through the CLI to execute the full experiment matrix:

```bash
maka eval run experiment.json --out .maka-eval/run-001

```

The `openExperimentDirectory` function (in [`experiment-directory.ts`](https://github.com/apache/maka/blob/main/experiment-directory.ts)) creates the output structure, including an `attempts` subdirectory for the immutable logs. The CLI validates the specification schema and environment variables before spawning the Runtime Host.

Upon completion, the CLI prints a JSON summary to stdout:

```json
{"experimentId":"terminal-bench-2.1","cells":18,"incomplete":0}

```

The `cells` field reports the total number of Cartesian product combinations processed. If `incomplete` is greater than zero, indicating cells that did not produce results, the CLI exits with a non-zero status code.

## Analyzing and Reproducing Results

All artifacts are stored under the output directory specified by `--out`. The directory structure follows this pattern:

```

.maka-eval/run-001/
├─ attempts/
│  ├─ cell-0.json
│  ├─ cell-1.json
│  └─ …
└─ experiment.json   # copy of the original spec

```

Each file in `attempts/` contains an immutable result kernel with the complete execution context. Because these logs are append-only and the specification pins all variables (Docker images, subject versions, task definitions), rerunning the exact same command on a machine with identical executor images yields bit-for-bit reproducible results.

Parse results programmatically using Node.js:

```javascript
import { readFile } from 'node:fs/promises';
import { join } from 'node:path';

const outDir = '.maka-eval/run-001';
const summary = JSON.parse(await readFile(join(outDir, 'summary.json'), 'utf8'));
console.log('Completed cells:', summary.cells);

```

## Advanced Configuration Options

### Replaying Failed Cells

When specific cells fail, avoid rerunning the entire matrix by targeting individual cells:

```bash
maka eval run experiment.json --out .maka-eval/run-001 --cell cell-3

```

The `--cell` argument (parsed around lines 33-35 in [`cli.ts`](https://github.com/apache/maka/blob/main/cli.ts)) instructs the runtime to execute only the specified cell ID, saving time during debugging.

### Custom Executors and Credential Injection

Implement a custom `HarnessExecutor` and reference it in the specification's `executor.kind` field. The CLI invokes `builtinExecutor` (lines 11-15 in [`cli.ts`](https://github.com/apache/maka/blob/main/cli.ts)) to load the appropriate implementation. Subject adapters receive secret-free environment variable names, which the CLI pre-validates and passes to the executor's `preflight` step.

### Parallel Execution Control

The Runtime Host respects concurrency limits defined in the specification's `task` and `repetition` configuration. This guarantees controlled parallelism while maintaining deterministic ordering of the append-only attempt logs.

For a deeper architectural view of how these components interact, consult [`ARCHITECTURE.md`](https://github.com/apache/maka/blob/main/ARCHITECTURE.md) in the repository root.

## Summary

- **Maka Eval** provides reproducible benchmarking through immutable JSON specifications and append-only result logs.
- The three-layer architecture (specification, CLI driver, runtime) is implemented in [`packages/eval/src/spec.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/spec.ts), [`packages/eval/src/cli.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/cli.ts), and [`packages/eval/src/runner.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/runner.ts).
- Execute experiments via `maka eval run <spec>.json --out <dir>` after building the workspace with `npm run build`.
- Results are stored in the `attempts/` subdirectory as immutable JSON files, enabling exact reproduction when rerun with identical specifications and executor images.
- Use `--cell <id>` to rerun specific failed cells without reprocessing the entire experiment matrix.

## Frequently Asked Questions

### What file format does Maka Eval use for experiment specifications?

Maka Eval uses a JSON format conforming to the `maka.eval.v1` schema. The `ExperimentSpec` type is defined in [`packages/eval/src/spec.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/spec.ts) and includes fields for `executor`, `subjects`, `tasks`, and `repetitions`. This JSON file serves as the single source of truth for all experimental parameters.

### How does Maka Eval ensure reproducibility across different runs?

Reproducibility is enforced through three mechanisms: deterministic **ExperimentSpec** files that pin all variables, containerized execution environments (Harbor or Pier executors) that ensure consistent runtime conditions, and append-only attempt logs stored in the output directory. The `runExperiment` function in [`packages/eval/src/runner.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/runner.ts) writes results immutably, preventing modification of historical data.

### Can I rerun only specific failed cells without repeating the entire experiment?

Yes. Pass the `--cell <cell-id>` flag to the CLI command. The argument parsing logic around lines 33-35 in [`packages/eval/src/cli.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/cli.ts) handles this flag, instructing the Runtime Host to execute only the specified cell rather than the full Cartesian product matrix.

### Where does Maka Eval store benchmark results and logs?

Results are stored in the output directory specified by the `--out` flag. The `openExperimentDirectory` function creates an `attempts/` subdirectory containing individual JSON files for each cell (e.g., [`cell-0.json`](https://github.com/apache/maka/blob/main/cell-0.json)), plus a copy of the original [`experiment.json`](https://github.com/apache/maka/blob/main/experiment.json) specification. Each attempt file contains the complete result kernel including metrics like `score`, `usage`, `cost`, and `duration`.