How Apache Maka Handles Multi-Arm Experiments: Eval System Architecture
Apache Maka’s eval subsystem treats multi-arm experiments as a Cartesian product of tasks, repetitions, and subjects, generating discrete ExperimentCell objects that the runner orchestrates through the harbor executor.
Apache Maka provides a sophisticated evaluation framework for benchmarking AI systems across diverse configurations. The multi-arm experiments capability allows researchers to test multiple task variations, subjects, and repetition counts within a single experiment definition. This article examines how the eval system in apache/maka expands experiment specifications into executable arms and manages their lifecycle from specification to result aggregation.
Experiment Specification and the ExperimentSpec Type
Every multi-arm experiment begins with an experiment.json file that conforms to the ExperimentSpec type defined in packages/eval/src/experiment.ts. This specification declares the experimental variables that the system will combine into individual arms.
The spec contains four critical arrays that define the experiment matrix:
tasks– The distinct task configurations to evaluatesubjects– The systems or models under testrepetitions– The number of times each task-subject pair should runbenchmark,executor,budget, andverifier– Shared configuration applied to every generated arm
The eval system treats these arrays as dimensions in a Cartesian product, ensuring comprehensive coverage of the experimental design space.
Generating Arms with the Cartesian Product Expansion
The transformation from specification to executable units occurs in the expandExperiment function within packages/eval/src/experiment.ts. This generator creates an ExperimentCell for every combination of task, repetition, and subject (lines 84–100).
// packages/eval/src/experiment.ts
export function expandExperiment(spec: ExperimentSpec): ExperimentCell[] {
return spec.tasks.flatMap(task =>
Array.from({ length: spec.repetitions }, (_, i) => i + 1).flatMap(repetition =>
spec.subjects.map(subject => ({
id: `${task.id}::${repetition}::${subject.id}`,
experimentId: spec.id,
benchmark: spec.benchmark,
executor: spec.executor,
subject,
task,
repetition,
budget: spec.budget,
verifier: spec.verifier,
})),
),
);
}
Each generated cell receives a unique identifier following the pattern <task-id>::<repetition-number>::<subject-id>. This naming convention ensures that every experimental arm is addressable independently, even when multiple arms share the same task or subject.
The ExperimentCell object binds the specific task payload, subject information, and repetition number to the shared benchmark and executor configuration, creating a self-contained unit of work ready for execution.
Cell Selection and Validation in the Runner
The orchestration logic resides in packages/eval/src/runner.ts, which loads the experiment specification and invokes expandExperiment to obtain the complete list of cells. The runner supports selective execution through a filtering mechanism that validates user-provided cell IDs against the generated set.
When specific arms are requested via the --select flag, the runner verifies each ID exists in the known cells map:
// packages/eval/src/runner.ts (partial)
const cells = expandExperiment(spec);
if (selected.length) {
for (const id of selected) {
if (!known.has(id)) {
throw new Error(`unknown experiment cell: ${id}`);
}
}
}
If a requested cell ID does not match the pattern <task-id>::<repetition>::<subject-id> or references a combination not present in the spec, the runner immediately throws an unknown experiment cell error. This strict validation prevents partial experiment runs due to typos or configuration mismatches.
Executing Individual Arms Through the Harbor Executor
For each selected ExperimentCell, the runner creates an attempt object that encapsulates the execution context. The attempt contains the bound benchmark configuration, executor settings, subject metadata, task payload, repetition number, budget constraints, and verification rules.
The runner dispatches these attempts to the harbor executor, which handles the actual execution environment setup, process isolation, and resource management. This separation of concerns allows the eval system to maintain a declarative experiment definition while the executor handles the imperative execution details.
Each arm runs independently, enabling parallel execution across the experiment matrix. The system tracks results at the cell level, preserving the granularity needed for statistical analysis across repetitions and comparative analysis across subjects.
CLI Interface for Running Multi-Arm Experiments
Researchers interact with the multi-arm system through the maka eval command exposed in packages/cli/src/cli-core.ts. The CLI accepts an experiment specification and optional cell selection parameters.
# Run a multi-arm experiment, optionally limiting to specific arms
maka eval run --spec path/to/experiment.json \
--select taskA::1::subjectX,taskB::2::subjectY
When the --select flag is omitted, the runner executes all arms in the Cartesian product. This flexibility supports both full factorial experiments and targeted re-runs of specific problematic configurations.
Summary
- Cartesian product expansion – The
expandExperimentfunction inpackages/eval/src/experiment.tsgenerates arms by combining tasks, repetitions, and subjects into discreteExperimentCellobjects. - Structured cell identifiers – Each arm receives a unique ID formatted as
<task-id>::<repetition-number>::<subject-id>, enabling precise selection and debugging. - Strict validation – The runner in
packages/eval/src/runner.tsvalidates selected cell IDs against the expanded set, throwing an error for unknown cells to prevent execution failures. - Attempt-based execution – The system wraps each cell in an attempt object containing full execution context before dispatching to the harbor executor for actual running.
- Flexible CLI targeting – The
maka eval runcommand supports both full matrix execution and selective arm running through comma-separated cell IDs.
Frequently Asked Questions
What constitutes a "multi-arm" experiment in Apache Maka?
A multi-arm experiment is any evaluation that tests multiple configurations through the Cartesian product of tasks, subjects, and repetitions defined in an experiment.json file. Each unique combination generates one arm (an ExperimentCell), allowing systematic testing of how different subjects perform across various tasks with statistical repetition.
How does the eval system handle invalid cell IDs?
The runner validates all user-provided cell IDs against the set generated by expandExperiment before execution begins. If a requested ID does not exist, the system throws an unknown experiment cell error immediately, preventing partial experiment execution due to configuration errors.
Can I run only specific arms of a multi-arm experiment?
Yes. Use the --select flag with the maka eval run command, providing comma-separated cell IDs in the format task-id::repetition::subject-id. This allows targeted execution of specific task-subject combinations or rerunning failed arms without executing the entire experiment matrix.
Where are multi-arm experiment results stored?
Results from each ExperimentCell are stored independently based on the cell ID, maintaining separation between different tasks, subjects, and repetitions. This structure enables per-arm performance analysis, variance calculation across repetitions, and comparative studies between subjects while preserving the experimental provenance of each data point.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →