# How to Run Multi-Arm Benchmark Experiments with Eval Cells and Attempts in Apache Maka

> Learn to run multi-arm benchmark experiments in Apache Maka. Execute multiple subjects simultaneously using Eval Cells and Attempts for reproducible results.

- Repository: [The Apache Software Foundation/maka](https://github.com/apache/maka)
- Tags: how-to-guide
- Published: 2026-08-27

---

**Multi-arm benchmark experiments in Apache Maka execute multiple subjects (arms) simultaneously through a hierarchy of Experiments → Cells → Attempts, using the `maka eval run` command to expand the Cartesian product of tasks, repetitions, and subjects while maintaining isolated attempt logs for reproducible infrastructure retries.**

Apache Maka provides a declarative framework for running systematic benchmarks across multiple AI agents or configurations. The system treats each combination of task, repetition, and subject as an isolated **Cell**, allowing you to run complex multi-arm studies where different agents compete on identical tasks without interference.

## Understanding the Experiment Hierarchy

Maka’s evaluation framework is built on three nested abstractions defined in [`packages/eval/README.md`](https://github.com/apache/maka/blob/main/packages/eval/README.md). Understanding this hierarchy is essential for designing reproducible benchmark studies.

### Experiments as Declarative Specifications

An **Experiment** is a JSON specification that declares the complete benchmark configuration. According to the source code in [`packages/eval/src/experiment.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/experiment.ts), the experiment parser validates and expands this specification into discrete execution units.

The experiment schema requires:
- **Benchmark** configuration (id, version, config)
- **Executor** reference (e.g., "harbor")
- **Subjects** array (each representing one arm)
- **Tasks** array defining the evaluation prompts
- **Repetition** count for statistical significance
- **Budget** constraints (max tokens)
- **Verifier** for result validation
- **Concurrency** limits

```json
{
  "benchmark": { "id": "terminal-bench", "version": "2.1", "config": {} },
  "executor": "harbor",
  "subjects": [
    { "id": "deepseek-harness", "type": "harness", "arm": "deepseek" },
    { "id": "maka", "type": "maka", "arm": "maka" }
  ],
  "tasks": [{ "id": "task-001", "prompt": "Complete the benchmark task." }],
  "repetition": 16,
  "budget": { "maxTokens": 100000 },
  "verifier": "terminal-bench",
  "concurrency": 8
}

```

### Cells and the Cartesian Product

A **Cell** represents the Cartesian product of `task × repetition × subject`. As documented in [`packages/eval/README.md`](https://github.com/apache/maka/blob/main/packages/eval/README.md) (lines 30-31), each cell is an independent trial that the Runtime Host executes in isolation.

For multi-arm experiments, each subject adds its own container. An eight-arm experiment with 16 repetitions generates 128 concurrent trials (8 arms × 16 repetitions). The [`packages/runtime-host/src/server/goal-coordinator.ts`](https://github.com/apache/maka/blob/main/packages/runtime-host/src/server/goal-coordinator.ts) orchestrates these cells, bridging the CLI and containerized execution environments.

When the subject's `Agent.run()` completes inside the Harbor or Pier container, the host records the **Result Kernel** containing score, usage, cost, duration, status, and artifacts.

### Attempts and Infrastructure Retry Logic

**Attempts** provide append-only retry logs per cell (lines 38-39 of [`packages/eval/README.md`](https://github.com/apache/maka/blob/main/packages/eval/README.md)). If a trial fails due to transient infrastructure errors, Maka does not overwrite the failed result; instead, it appends a new attempt to the same cell's log.

The system uses **earliest valid attempt wins** semantics. This guarantees reproducibility—earlier attempts remain visible for debugging while successful retries provide the final data. The attempt logs are stored under `.maka-eval/<run-id>/attempts/<cell-id>/`.

## Running Multi-Arm Experiments via CLI

The CLI entry point in [`packages/cli/src/cli-core.ts`](https://github.com/apache/maka/blob/main/packages/cli/src/cli-core.ts) (line 136) drives the full experiment lifecycle. To execute a multi-arm benchmark:

```bash
maka eval run experiments/terminal-bench-2.1-deepseek-v4-flash-four-arm.json \
    --out ./run-004

```

The command performs the following operations:
1. Validates the experiment JSON schema
2. Expands the specification into individual cells
3. Validates executor prerequisites (machine paths, relay files, Docker daemon)
4. Dispatches cells to the Runtime Host up to the concurrency limit
5. Collects results and pairs them by `task.id` for statistical analysis

Results are **not** mixed across arms during execution; they remain isolated by cell ID and are paired only during post-processing for comparison.

## Targeting Specific Cells for Retry

When infrastructure failures occur, target specific cells without re-running the entire experiment. The `--cell` flag accepts the cell identifier format `task::<index>::<subject>`:

```bash
maka eval run experiments/terminal-bench-2.1-deepseek-v4-flash-four-arm.json \
    --out ./run-004 \
    --cell task::5::deepseek-harness

```

Only the failed or indeterminate cell is re-executed. The new attempt is appended to the existing attempt log for that cell, preserving the history of transient failures while updating the final result.

## Pre-Execution Validation

Before launching trials, Maka validates execution prerequisites to prevent half-started runs. The framework checks:
- **Executor machine paths** accessibility
- **Bundled relay files** integrity
- **Pinned Harbor or Pier** Python distributions availability
- **Docker daemon** connectivity

Missing prerequisites abort the CLI immediately with clear error messages, ensuring resources are not wasted on invalid configurations.

## Summary

- **Experiments** are declarative JSON specs that define multi-arm benchmarks with subjects, tasks, and repetition counts.
- **Cells** instantiate the Cartesian product of task × repetition × subject, enabling isolated parallel execution of up to 128+ concurrent trials.
- **Attempts** provide append-only retry logs per cell, using "earliest valid wins" semantics for reproducible infrastructure recovery.
- The **`maka eval run`** command expands experiments and orchestrates execution through the Runtime Host.
- Use **`--cell`** flags to retry specific failed cells without re-running successful trials.
- Pre-execution validation in [`packages/cli/src/cli-core.ts`](https://github.com/apache/maka/blob/main/packages/cli/src/cli-core.ts) ensures Docker, paths, and distributions are ready before any trials start.

## Frequently Asked Questions

### What is the difference between a Cell and an Attempt in Maka?

A **Cell** is the fundamental unit of work representing one specific combination of task, repetition, and subject. An **Attempt** is a single execution trial within that cell. If a cell fails due to infrastructure errors, Maka creates a new attempt rather than overwriting the failure, maintaining an append-only log where the earliest successful attempt provides the final result.

### How does Maka handle concurrent execution in multi-arm experiments?

Maka calculates the Cartesian product of all tasks, repetitions, and subjects to generate cells, then dispatches them concurrently up to the specified limit. Each arm runs in its own container (Harbor or Pier), ensuring isolation. An eight-arm experiment with 16 repetitions produces 128 cells that execute in parallel batches according to the concurrency parameter.

### Can I mix different executor types in a single multi-arm experiment?

No. The experiment specification defines a single **executor** (e.g., "harbor") for the entire experiment. However, you can define multiple **subjects** of different types (harness, maka, etc.) within that executor. Each subject arm must be compatible with the specified executor, as validated in [`packages/eval/src/experiment.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/experiment.ts).

### Where does Maka store the results and logs from benchmark runs?

Results are stored in the directory specified by the `--out` flag (default: `.maka-eval/run-<timestamp>/`). The structure includes cell logs and attempt logs under `attempts/<cell-id>/`, containing result kernels with scores, usage metrics, costs, and artifacts. This directory serves as the workspace for the specific run and enables targeted retries via the `--cell` parameter.