How to Set Up Reproducible Benchmark Experiments with Maka's Eval Framework

Maka's Eval framework provides a declarative, three-phase workflow for reproducible benchmark experiments: define a validated JSON spec, run via CLI, and collect deterministic results from immutable attempt records.

Setting up reproducible benchmark experiments is critical for engineering teams evaluating AI agents. Maka's Eval framework, located in the apache/maka repository, eliminates hidden state and non-determinism through strict schema validation, pinned executor environments, and append-only result logging. This guide walks through the complete setup based on the actual source implementation.

Understanding the Eval Architecture

Maka's Eval framework operates as a thin semantic layer above the Runtime Host, which serves as the sole execution authority. The @maka/eval package owns experiment semantics only—cells, attempts, result selection, budgets, and verifier configuration—while executor adapters like Harbor and Pier translate generic experiments into containerized runs.

The interaction flow is documented in ARCHITECTURE.md, where the Eval layer feeds cells to the Runtime Host for execution. Each cell represents a unique combination of task × repetition × subject, expanded as a Cartesian product by the CLI.

Core Components

  • Runtime Host – Manages sessions, tools, and event logs; the only component with execution authority.
  • @maka/eval – Handles experiment semantics without direct execution capability.
  • Executor adapters (Harbor, Pier) – Convert abstract experiments into concrete Docker runs.
  • Subject types – maka subjects (built-in benchmark agents) and external subjects (arbitrary CLI programs).

The Three-Phase Reproducibility Workflow

Phase 1: Define and Validate the Experiment Spec

The experiment spec is a JSON document describing every parameter of your benchmark. The packages/eval/src/spec.ts module enforces strict schema validation including required fields, identifier format, and positive integers for numeric values.

A minimal valid spec requires:

  • schemaVersion – Use "maka.eval.v1"
  • id – Unique experiment identifier
  • benchmark.id and benchmark.version – Exact benchmark release
  • executor.kind and executor.config – Fixed Docker image with tag
  • subjects[] – Subject configurations with credential references
  • tasks[] – Input prompts or data
  • repetitions – Integer count (e.g., 3)

The validator in spec.ts rejects malformed specs before execution begins. You can pre-validate without running:

node -e "import('../packages/eval/src/spec.js').parseExperimentSpec(require('./my-spec.json'))"

Phase 2: Execute via the CLI

The public CLI entry point lives in packages/eval/src/cli.ts and exposes the maka eval run command:

maka eval run <spec> --out <directory>

The CLI performs pre-flight checks for paths and Docker availability, then expands your spec into cells. Each cell executes through the Runtime Host, producing immutable attempt records. Usage help is printed from cli.ts line 34.

Example execution:


# Run experiment and store results

maka eval run my-spec.json --out .maka-eval/run-001

Phase 3: Collect Deterministic Results

Each cell produces an immutable attempt record containing score, usage, cost, duration, status, and artifact digests. The earliest valid attempt becomes the authoritative result. Because attempts are append-only, you can re-run experiments and compare kernel digests for exact verification.

Step-by-Step Setup Instructions

Follow these steps to achieve fully reproducible benchmark experiments with Maka's Eval framework:

  1. Pin the repository version

    git clone https://github.com/apache/maka.git
    cd maka
    git checkout v0.1.0  # Use specific tag or commit
    
  2. Install dependencies (Node 22+ required)

    npm ci
  3. Prepare executor toolchain (for Harbor/Pier)

    Docker images are version-pinned, e.g., maka-eval-egress-proxy:12.2.3. See packages/eval/README.md for Harbor and Pier sections.

  4. Create a deterministic spec JSON

    Include exact versions, fixed image digests, and explicit task inputs. Store actual secret values locally via environment variables referenced in credentials[], not in version control.

  5. Validate the spec

    Use the spec module validator or rely on CLI validation at launch.

  6. Run the experiment

    maka eval run my-spec.json --out .maka-eval/run-001
  7. Re-run for verification

    Execute identical command with fresh output directory. Matching kernel digests confirm reproducibility.

Complete Example: Minimal Experiment Spec

{
  "schemaVersion": "maka.eval.v1",
  "id": "demo-exp",
  "benchmark": {
    "id": "terminal-bench",
    "version": "2.1",
    "config": {}
  },
  "executor": {
    "kind": "harbor",
    "config": {
      "image": "docker.io/maka/harbor:12.2.3"
    }
  },
  "subjects": [
    {
      "id": "maka-subject",
      "kind": "maka",
      "credentials": [],
      "config": {}
    }
  ],
  "tasks": [
    {
      "id": "task-1",
      "input": "Explain the repository structure.",
      "config": {}
    }
  ],
  "repetitions": 3,
  "budget": {},
  "verifier": {}
}

Inspect and Verify Results

After execution, examine the attempt records:


# List all attempt directories

ls .maka-eval/run-001/attempts/

# View formatted result kernel

cat .maka-eval/run-001/attempts/0/result.json | jq .

Each result kernel contains deterministic fields: score, usage, cost, duration, status, and artifact digests. Comparing these across runs proves reproducibility.

Why This Guarantees Reproducibility

  • Declarative spec – All parameters explicit in JSON; no hidden state
  • Pinned executor images – Docker images identified by digest, preventing accidental upgrades
  • Immutable attempt logs – Append-only records; earliest valid attempt is canonical
  • Runtime Host sandbox – Tools execute in controlled environment; network access blocked by egress proxy per README.md line 70

Key Source Files

File Purpose
packages/eval/src/spec.ts JSON schema validation and experiment spec parsing
packages/eval/src/cli.ts CLI implementation with maka eval run command
ARCHITECTURE.md High-level architecture showing Eval-Runtime Host interaction
packages/eval/README.md Framework overview, CLI usage, executor configuration

Summary

  • Maka's Eval framework enables reproducible benchmark experiments through a declarative JSON spec, validated CLI execution, and immutable result logging.
  • Pin repository versions, executor images, and task inputs to eliminate variability.
  • Validate specs via packages/eval/src/spec.ts before execution.
  • Run experiments with maka eval run <spec> --out <dir> and verify by comparing kernel digests across re-runs.
  • The Runtime Host sandbox and egress proxy block external network access, ensuring identical execution environments.

Frequently Asked Questions

What makes Maka's Eval framework different from other benchmark tools?

Unlike tools that rely on procedural scripts with hidden state, Maka's Eval uses a strictly validated JSON spec where all parameters are explicit. The spec.ts validator enforces schema compliance before execution, and the Runtime Host sandbox guarantees identical execution environments through containerized, network-isolated runs.

How do I prevent credential leakage in version-controlled specs?

Store sensitive values as environment variable names in subjects[].credentials, not literal values. The spec references variable names; actual secrets remain local. This pattern keeps specs shareable and version-controlled while protecting credentials.

Can I reproduce results on a different machine?

Yes, provided you use identical inputs: same repository commit, same Node version (22+), same executor Docker image digest, and same spec. The framework is designed for cross-machine reproducibility—the Runtime Host abstracts hardware differences through containerization.

What happens if an attempt fails or times out?

Failed attempts are logged as immutable records with status fields indicating failure reason. The earliest valid attempt selection logic ignores failed attempts. Re-running the same spec will retry cells, producing new attempts without modifying previous logs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →