# How to Set Up Reproducible Benchmark Experiments with Maka's Eval Framework

> Learn to set up reproducible benchmark experiments with Maka's Eval framework. Define specs, run CLI commands, and collect deterministic results for reliable testing.

- Repository: [The Apache Software Foundation/maka](https://github.com/apache/maka)
- Tags: tutorial
- Published: 2026-08-29

---

**Maka's Eval framework provides a declarative, three-phase workflow for reproducible benchmark experiments: define a validated JSON spec, run via CLI, and collect deterministic results from immutable attempt records.**

Setting up reproducible benchmark experiments is critical for engineering teams evaluating AI agents. Maka's Eval framework, located in the `apache/maka` repository, eliminates hidden state and non-determinism through strict schema validation, pinned executor environments, and append-only result logging. This guide walks through the complete setup based on the actual source implementation.

## Understanding the Eval Architecture

Maka's Eval framework operates as a thin semantic layer above the **Runtime Host**, which serves as the sole execution authority. The `@maka/eval` package owns experiment semantics only—cells, attempts, result selection, budgets, and verifier configuration—while executor adapters like Harbor and Pier translate generic experiments into containerized runs.

The interaction flow is documented in [`ARCHITECTURE.md`](https://github.com/apache/maka/blob/main/ARCHITECTURE.md), where the Eval layer feeds cells to the Runtime Host for execution. Each cell represents a unique combination of task × repetition × subject, expanded as a Cartesian product by the CLI.

### Core Components

- **Runtime Host** – Manages sessions, tools, and event logs; the only component with execution authority.
- **@maka/eval** – Handles experiment semantics without direct execution capability.
- **Executor adapters** (Harbor, Pier) – Convert abstract experiments into concrete Docker runs.
- **Subject types** – `maka` subjects (built-in benchmark agents) and `external` subjects (arbitrary CLI programs).

## The Three-Phase Reproducibility Workflow

### Phase 1: Define and Validate the Experiment Spec

The experiment spec is a JSON document describing every parameter of your benchmark. The [`packages/eval/src/spec.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/spec.ts) module enforces strict schema validation including required fields, identifier format, and positive integers for numeric values.

A minimal valid spec requires:

- `schemaVersion` – Use `"maka.eval.v1"`
- `id` – Unique experiment identifier
- `benchmark.id` and `benchmark.version` – Exact benchmark release
- `executor.kind` and `executor.config` – Fixed Docker image with tag
- `subjects[]` – Subject configurations with credential references
- `tasks[]` – Input prompts or data
- `repetitions` – Integer count (e.g., `3`)

The **validator** in [`spec.ts`](https://github.com/apache/maka/blob/main/spec.ts) rejects malformed specs before execution begins. You can pre-validate without running:

```bash
node -e "import('../packages/eval/src/spec.js').parseExperimentSpec(require('./my-spec.json'))"

```

### Phase 2: Execute via the CLI

The public CLI entry point lives in [`packages/eval/src/cli.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/cli.ts) and exposes the `maka eval run` command:

```bash
maka eval run <spec> --out <directory>

```

The CLI performs pre-flight checks for paths and Docker availability, then expands your spec into cells. Each cell executes through the Runtime Host, producing immutable attempt records. Usage help is printed from [`cli.ts`](https://github.com/apache/maka/blob/main/cli.ts) line 34.

Example execution:

```bash

# Run experiment and store results

maka eval run my-spec.json --out .maka-eval/run-001

```

### Phase 3: Collect Deterministic Results

Each cell produces an **immutable attempt record** containing score, usage, cost, duration, status, and artifact digests. The earliest valid attempt becomes the authoritative result. Because attempts are append-only, you can re-run experiments and compare kernel digests for exact verification.

## Step-by-Step Setup Instructions

Follow these steps to achieve fully reproducible benchmark experiments with Maka's Eval framework:

1. **Pin the repository version**

   ```bash
   git clone https://github.com/apache/maka.git
   cd maka
   git checkout v0.1.0  # Use specific tag or commit

   ```

2. **Install dependencies** (Node 22+ required)

   ```bash
   npm ci
   ```

3. **Prepare executor toolchain** (for Harbor/Pier)

   Docker images are version-pinned, e.g., `maka-eval-egress-proxy:12.2.3`. See [`packages/eval/README.md`](https://github.com/apache/maka/blob/main/packages/eval/README.md) for Harbor and Pier sections.

4. **Create a deterministic spec JSON**

   Include exact versions, fixed image digests, and explicit task inputs. Store actual secret values locally via environment variables referenced in `credentials[]`, not in version control.

5. **Validate the spec**

   Use the spec module validator or rely on CLI validation at launch.

6. **Run the experiment**

   ```bash
   maka eval run my-spec.json --out .maka-eval/run-001
   ```

7. **Re-run for verification**

   Execute identical command with fresh output directory. Matching kernel digests confirm reproducibility.

## Complete Example: Minimal Experiment Spec

```json
{
  "schemaVersion": "maka.eval.v1",
  "id": "demo-exp",
  "benchmark": {
    "id": "terminal-bench",
    "version": "2.1",
    "config": {}
  },
  "executor": {
    "kind": "harbor",
    "config": {
      "image": "docker.io/maka/harbor:12.2.3"
    }
  },
  "subjects": [
    {
      "id": "maka-subject",
      "kind": "maka",
      "credentials": [],
      "config": {}
    }
  ],
  "tasks": [
    {
      "id": "task-1",
      "input": "Explain the repository structure.",
      "config": {}
    }
  ],
  "repetitions": 3,
  "budget": {},
  "verifier": {}
}

```

## Inspect and Verify Results

After execution, examine the attempt records:

```bash

# List all attempt directories

ls .maka-eval/run-001/attempts/

# View formatted result kernel

cat .maka-eval/run-001/attempts/0/result.json | jq .

```

Each result kernel contains deterministic fields: **score**, **usage**, **cost**, **duration**, **status**, and **artifact digests**. Comparing these across runs proves reproducibility.

## Why This Guarantees Reproducibility

- **Declarative spec** – All parameters explicit in JSON; no hidden state
- **Pinned executor images** – Docker images identified by digest, preventing accidental upgrades
- **Immutable attempt logs** – Append-only records; earliest valid attempt is canonical
- **Runtime Host sandbox** – Tools execute in controlled environment; network access blocked by egress proxy per [`README.md`](https://github.com/apache/maka/blob/main/README.md) line 70

## Key Source Files

| File | Purpose |
|------|---------|
| [`packages/eval/src/spec.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/spec.ts) | JSON schema validation and experiment spec parsing |
| [`packages/eval/src/cli.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/cli.ts) | CLI implementation with `maka eval run` command |
| [`ARCHITECTURE.md`](https://github.com/apache/maka/blob/main/ARCHITECTURE.md) | High-level architecture showing Eval-Runtime Host interaction |
| [`packages/eval/README.md`](https://github.com/apache/maka/blob/main/packages/eval/README.md) | Framework overview, CLI usage, executor configuration |

## Summary

- Maka's Eval framework enables reproducible benchmark experiments through a **declarative JSON spec**, **validated CLI execution**, and **immutable result logging**.
- Pin repository versions, executor images, and task inputs to eliminate variability.
- Validate specs via [`packages/eval/src/spec.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/spec.ts) before execution.
- Run experiments with `maka eval run <spec> --out <dir>` and verify by comparing kernel digests across re-runs.
- The Runtime Host sandbox and egress proxy block external network access, ensuring identical execution environments.

## Frequently Asked Questions

### What makes Maka's Eval framework different from other benchmark tools?

Unlike tools that rely on procedural scripts with hidden state, Maka's Eval uses a **strictly validated JSON spec** where all parameters are explicit. The [`spec.ts`](https://github.com/apache/maka/blob/main/spec.ts) validator enforces schema compliance before execution, and the Runtime Host sandbox guarantees identical execution environments through containerized, network-isolated runs.

### How do I prevent credential leakage in version-controlled specs?

Store sensitive values as **environment variable names** in `subjects[].credentials`, not literal values. The spec references variable names; actual secrets remain local. This pattern keeps specs shareable and version-controlled while protecting credentials.

### Can I reproduce results on a different machine?

Yes, provided you use **identical inputs**: same repository commit, same Node version (22+), same executor Docker image digest, and same spec. The framework is designed for cross-machine reproducibility—the Runtime Host abstracts hardware differences through containerization.

### What happens if an attempt fails or times out?

Failed attempts are logged as immutable records with status fields indicating failure reason. The **earliest valid attempt** selection logic ignores failed attempts. Re-running the same spec will retry cells, producing new attempts without modifying previous logs.