# How Apache Maka Handles Multi-Arm Experiments: Eval System Architecture

> Discover how Apache Maka handles multi-arm experiments by transforming them into discrete ExperimentCell objects orchestrated by the harbor executor. Learn the core of its eval system architecture.

- Repository: [The Apache Software Foundation/maka](https://github.com/apache/maka)
- Tags: architecture
- Published: 2026-08-25

---

**Apache Maka’s eval subsystem treats multi-arm experiments as a Cartesian product of tasks, repetitions, and subjects, generating discrete `ExperimentCell` objects that the runner orchestrates through the harbor executor.**

Apache Maka provides a sophisticated evaluation framework for benchmarking AI systems across diverse configurations. The **multi-arm experiments** capability allows researchers to test multiple task variations, subjects, and repetition counts within a single experiment definition. This article examines how the eval system in `apache/maka` expands experiment specifications into executable arms and manages their lifecycle from specification to result aggregation.

## Experiment Specification and the ExperimentSpec Type

Every multi-arm experiment begins with an [`experiment.json`](https://github.com/apache/maka/blob/main/experiment.json) file that conforms to the `ExperimentSpec` type defined in [`packages/eval/src/experiment.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/experiment.ts). This specification declares the experimental variables that the system will combine into individual arms.

The spec contains four critical arrays that define the experiment matrix:
- **`tasks`** – The distinct task configurations to evaluate
- **`subjects`** – The systems or models under test  
- **`repetitions`** – The number of times each task-subject pair should run
- **`benchmark`**, **`executor`**, **`budget`**, and **`verifier`** – Shared configuration applied to every generated arm

The eval system treats these arrays as dimensions in a Cartesian product, ensuring comprehensive coverage of the experimental design space.

## Generating Arms with the Cartesian Product Expansion

The transformation from specification to executable units occurs in the `expandExperiment` function within [`packages/eval/src/experiment.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/experiment.ts). This generator creates an `ExperimentCell` for every combination of task, repetition, and subject (lines 84–100).

```typescript
// packages/eval/src/experiment.ts
export function expandExperiment(spec: ExperimentSpec): ExperimentCell[] {
  return spec.tasks.flatMap(task =>
    Array.from({ length: spec.repetitions }, (_, i) => i + 1).flatMap(repetition =>
      spec.subjects.map(subject => ({
        id: `${task.id}::${repetition}::${subject.id}`,
        experimentId: spec.id,
        benchmark: spec.benchmark,
        executor: spec.executor,
        subject,
        task,
        repetition,
        budget: spec.budget,
        verifier: spec.verifier,
      })),
    ),
  );
}

```

Each generated cell receives a unique identifier following the pattern **`<task-id>::<repetition-number>::<subject-id>`**. This naming convention ensures that every experimental arm is addressable independently, even when multiple arms share the same task or subject.

The `ExperimentCell` object binds the specific task payload, subject information, and repetition number to the shared benchmark and executor configuration, creating a self-contained unit of work ready for execution.

## Cell Selection and Validation in the Runner

The orchestration logic resides in [`packages/eval/src/runner.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/runner.ts), which loads the experiment specification and invokes `expandExperiment` to obtain the complete list of cells. The runner supports selective execution through a filtering mechanism that validates user-provided cell IDs against the generated set.

When specific arms are requested via the `--select` flag, the runner verifies each ID exists in the known cells map:

```typescript
// packages/eval/src/runner.ts (partial)
const cells = expandExperiment(spec);
if (selected.length) {
  for (const id of selected) {
    if (!known.has(id)) {
      throw new Error(`unknown experiment cell: ${id}`);
    }
  }
}

```

If a requested cell ID does not match the pattern `<task-id>::<repetition>::<subject-id>` or references a combination not present in the spec, the runner immediately throws an **`unknown experiment cell`** error. This strict validation prevents partial experiment runs due to typos or configuration mismatches.

## Executing Individual Arms Through the Harbor Executor

For each selected `ExperimentCell`, the runner creates an **attempt** object that encapsulates the execution context. The attempt contains the bound benchmark configuration, executor settings, subject metadata, task payload, repetition number, budget constraints, and verification rules.

The runner dispatches these attempts to the **harbor executor**, which handles the actual execution environment setup, process isolation, and resource management. This separation of concerns allows the eval system to maintain a declarative experiment definition while the executor handles the imperative execution details.

Each arm runs independently, enabling parallel execution across the experiment matrix. The system tracks results at the cell level, preserving the granularity needed for statistical analysis across repetitions and comparative analysis across subjects.

## CLI Interface for Running Multi-Arm Experiments

Researchers interact with the multi-arm system through the `maka eval` command exposed in [`packages/cli/src/cli-core.ts`](https://github.com/apache/maka/blob/main/packages/cli/src/cli-core.ts). The CLI accepts an experiment specification and optional cell selection parameters.

```bash

# Run a multi-arm experiment, optionally limiting to specific arms

maka eval run --spec path/to/experiment.json \
               --select taskA::1::subjectX,taskB::2::subjectY

```

When the `--select` flag is omitted, the runner executes all arms in the Cartesian product. This flexibility supports both full factorial experiments and targeted re-runs of specific problematic configurations.

## Summary

- **Cartesian product expansion** – The `expandExperiment` function in [`packages/eval/src/experiment.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/experiment.ts) generates arms by combining tasks, repetitions, and subjects into discrete `ExperimentCell` objects.
- **Structured cell identifiers** – Each arm receives a unique ID formatted as `<task-id>::<repetition-number>::<subject-id>`, enabling precise selection and debugging.
- **Strict validation** – The runner in [`packages/eval/src/runner.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/runner.ts) validates selected cell IDs against the expanded set, throwing an error for unknown cells to prevent execution failures.
- **Attempt-based execution** – The system wraps each cell in an attempt object containing full execution context before dispatching to the harbor executor for actual running.
- **Flexible CLI targeting** – The `maka eval run` command supports both full matrix execution and selective arm running through comma-separated cell IDs.

## Frequently Asked Questions

### What constitutes a "multi-arm" experiment in Apache Maka?

A multi-arm experiment is any evaluation that tests multiple configurations through the Cartesian product of tasks, subjects, and repetitions defined in an [`experiment.json`](https://github.com/apache/maka/blob/main/experiment.json) file. Each unique combination generates one arm (an `ExperimentCell`), allowing systematic testing of how different subjects perform across various tasks with statistical repetition.

### How does the eval system handle invalid cell IDs?

The runner validates all user-provided cell IDs against the set generated by `expandExperiment` before execution begins. If a requested ID does not exist, the system throws an `unknown experiment cell` error immediately, preventing partial experiment execution due to configuration errors.

### Can I run only specific arms of a multi-arm experiment?

Yes. Use the `--select` flag with the `maka eval run` command, providing comma-separated cell IDs in the format `task-id::repetition::subject-id`. This allows targeted execution of specific task-subject combinations or rerunning failed arms without executing the entire experiment matrix.

### Where are multi-arm experiment results stored?

Results from each `ExperimentCell` are stored independently based on the cell ID, maintaining separation between different tasks, subjects, and repetitions. This structure enables per-arm performance analysis, variance calculation across repetitions, and comparative studies between subjects while preserving the experimental provenance of each data point.