How to Run Multi-Arm Benchmark Experiments with Eval Cells and Attempts in Apache Maka

Multi-arm benchmark experiments in Apache Maka execute multiple subjects (arms) simultaneously through a hierarchy of Experiments → Cells → Attempts, using the maka eval run command to expand the Cartesian product of tasks, repetitions, and subjects while maintaining isolated attempt logs for reproducible infrastructure retries.

Apache Maka provides a declarative framework for running systematic benchmarks across multiple AI agents or configurations. The system treats each combination of task, repetition, and subject as an isolated Cell, allowing you to run complex multi-arm studies where different agents compete on identical tasks without interference.

Understanding the Experiment Hierarchy

Maka’s evaluation framework is built on three nested abstractions defined in packages/eval/README.md. Understanding this hierarchy is essential for designing reproducible benchmark studies.

Experiments as Declarative Specifications

An Experiment is a JSON specification that declares the complete benchmark configuration. According to the source code in packages/eval/src/experiment.ts, the experiment parser validates and expands this specification into discrete execution units.

The experiment schema requires:

  • Benchmark configuration (id, version, config)
  • Executor reference (e.g., "harbor")
  • Subjects array (each representing one arm)
  • Tasks array defining the evaluation prompts
  • Repetition count for statistical significance
  • Budget constraints (max tokens)
  • Verifier for result validation
  • Concurrency limits
{
  "benchmark": { "id": "terminal-bench", "version": "2.1", "config": {} },
  "executor": "harbor",
  "subjects": [
    { "id": "deepseek-harness", "type": "harness", "arm": "deepseek" },
    { "id": "maka", "type": "maka", "arm": "maka" }
  ],
  "tasks": [{ "id": "task-001", "prompt": "Complete the benchmark task." }],
  "repetition": 16,
  "budget": { "maxTokens": 100000 },
  "verifier": "terminal-bench",
  "concurrency": 8
}

Cells and the Cartesian Product

A Cell represents the Cartesian product of task × repetition × subject. As documented in packages/eval/README.md (lines 30-31), each cell is an independent trial that the Runtime Host executes in isolation.

For multi-arm experiments, each subject adds its own container. An eight-arm experiment with 16 repetitions generates 128 concurrent trials (8 arms × 16 repetitions). The packages/runtime-host/src/server/goal-coordinator.ts orchestrates these cells, bridging the CLI and containerized execution environments.

When the subject's Agent.run() completes inside the Harbor or Pier container, the host records the Result Kernel containing score, usage, cost, duration, status, and artifacts.

Attempts and Infrastructure Retry Logic

Attempts provide append-only retry logs per cell (lines 38-39 of packages/eval/README.md). If a trial fails due to transient infrastructure errors, Maka does not overwrite the failed result; instead, it appends a new attempt to the same cell's log.

The system uses earliest valid attempt wins semantics. This guarantees reproducibility—earlier attempts remain visible for debugging while successful retries provide the final data. The attempt logs are stored under .maka-eval/<run-id>/attempts/<cell-id>/.

Running Multi-Arm Experiments via CLI

The CLI entry point in packages/cli/src/cli-core.ts (line 136) drives the full experiment lifecycle. To execute a multi-arm benchmark:

maka eval run experiments/terminal-bench-2.1-deepseek-v4-flash-four-arm.json \
    --out ./run-004

The command performs the following operations:

  1. Validates the experiment JSON schema
  2. Expands the specification into individual cells
  3. Validates executor prerequisites (machine paths, relay files, Docker daemon)
  4. Dispatches cells to the Runtime Host up to the concurrency limit
  5. Collects results and pairs them by task.id for statistical analysis

Results are not mixed across arms during execution; they remain isolated by cell ID and are paired only during post-processing for comparison.

Targeting Specific Cells for Retry

When infrastructure failures occur, target specific cells without re-running the entire experiment. The --cell flag accepts the cell identifier format task::<index>::<subject>:

maka eval run experiments/terminal-bench-2.1-deepseek-v4-flash-four-arm.json \
    --out ./run-004 \
    --cell task::5::deepseek-harness

Only the failed or indeterminate cell is re-executed. The new attempt is appended to the existing attempt log for that cell, preserving the history of transient failures while updating the final result.

Pre-Execution Validation

Before launching trials, Maka validates execution prerequisites to prevent half-started runs. The framework checks:

  • Executor machine paths accessibility
  • Bundled relay files integrity
  • Pinned Harbor or Pier Python distributions availability
  • Docker daemon connectivity

Missing prerequisites abort the CLI immediately with clear error messages, ensuring resources are not wasted on invalid configurations.

Summary

  • Experiments are declarative JSON specs that define multi-arm benchmarks with subjects, tasks, and repetition counts.
  • Cells instantiate the Cartesian product of task × repetition × subject, enabling isolated parallel execution of up to 128+ concurrent trials.
  • Attempts provide append-only retry logs per cell, using "earliest valid wins" semantics for reproducible infrastructure recovery.
  • The maka eval run command expands experiments and orchestrates execution through the Runtime Host.
  • Use --cell flags to retry specific failed cells without re-running successful trials.
  • Pre-execution validation in packages/cli/src/cli-core.ts ensures Docker, paths, and distributions are ready before any trials start.

Frequently Asked Questions

What is the difference between a Cell and an Attempt in Maka?

A Cell is the fundamental unit of work representing one specific combination of task, repetition, and subject. An Attempt is a single execution trial within that cell. If a cell fails due to infrastructure errors, Maka creates a new attempt rather than overwriting the failure, maintaining an append-only log where the earliest successful attempt provides the final result.

How does Maka handle concurrent execution in multi-arm experiments?

Maka calculates the Cartesian product of all tasks, repetitions, and subjects to generate cells, then dispatches them concurrently up to the specified limit. Each arm runs in its own container (Harbor or Pier), ensuring isolation. An eight-arm experiment with 16 repetitions produces 128 cells that execute in parallel batches according to the concurrency parameter.

Can I mix different executor types in a single multi-arm experiment?

No. The experiment specification defines a single executor (e.g., "harbor") for the entire experiment. However, you can define multiple subjects of different types (harness, maka, etc.) within that executor. Each subject arm must be compatible with the specified executor, as validated in packages/eval/src/experiment.ts.

Where does Maka store the results and logs from benchmark runs?

Results are stored in the directory specified by the --out flag (default: .maka-eval/run-<timestamp>/). The structure includes cell logs and attempt logs under attempts/<cell-id>/, containing result kernels with scores, usage metrics, costs, and artifacts. This directory serves as the workspace for the specific run and enables targeted retries via the --cell parameter.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →