How Maka's Evaluation Uses Declarative Multi-Arm Experiments: Task × Repetition × Subject Expansion

Maka's evaluation framework expands declarative JSON experiment specifications into a Cartesian grid of task × repetition × subject cells using the expandExperiment function, enabling reproducible, parallel execution across diverse runtime environments.

Apache Maka's evaluation system treats experiments as declarative specifications that define tasks, repetition counts, and execution subjects. The framework automatically expands these multi-arm experiments into discrete, immutable cells that can be executed and recorded independently, as implemented in the apache/maka repository under packages/eval/src.

Anatomy of a Declarative Multi-Arm Experiment

At the core of Maka's evaluation model is the ExperimentSpec interface, defined in packages/eval/src/experiment.ts. This specification acts as a blueprint containing three critical dimensions:

  • Tasks — An array of logical work-units, each with a unique identifier and input configuration.
  • Repetitions — An integer specifying how many times each task must be executed to ensure statistical significance.
  • Subjects — An array of execution environments, which may include native Maka subjects (kind: "maka") or external adapters (kind: "external").

When parsed via parseExperimentSpec in packages/eval/src/spec.ts, a raw JSON configuration transforms into a strongly-typed object ready for expansion. The declarative nature ensures that experiments are version-controlled, reproducible, and independent of execution infrastructure.

Cartesian Expansion into Task × Repetition × Subject Cells

The expansion logic resides in the expandExperiment function (packages/eval/src/experiment.ts, lines 84–100). This function computes the Cartesian product of the three experimental dimensions, generating an ExperimentCell for every unique combination:

export function expandExperiment(spec: ExperimentSpec): ExperimentCell[] {
  return spec.tasks.flatMap(task =>
    Array.from({ length: spec.repetitions }, (_, i) => i + 1).flatMap(repetition =>
      spec.subjects.map(subject => ({
        id: `${task.id}::${repetition}::${subject.id}`,
        experimentId: spec.id,
        benchmark: spec.benchmark,
        executor: spec.executor,
        subject,
        task,
        repetition,
        budget: spec.budget,
        verifier: spec.verifier,
      }))
    )
  );
}

This algorithm produces a deterministic cell identifier using the format taskId::repetition::subjectId. Because the identifier is derived from the spec's content rather than runtime state, cells remain immutable and idempotent. Each ExperimentCell encapsulates all data required for a single execution attempt, including the specific subject adapter configuration, task input, and verification rules.

Cell Execution and Concurrent Grouping

Once expanded, cells enter the execution phase via runExperiment in packages/eval/src/runner.ts (lines 38–50). The orchestration process follows a strict pipeline:

  1. Generation — All cells are materialized upfront via expandExperiment.
  2. Selection — Optional filtering allows targeting specific cell IDs for partial re-runs.
  3. Grouping — The groupTaskCells utility (lines 40–48) organizes cells by task and repetition, creating logical groups where all subjects for the same task repetition execute concurrently.
  4. Throttling — The maxConcurrentTaskGroups limit controls parallelism to prevent resource exhaustion.

During execution, the runner delegates to the appropriate subject adapter. Native Maka execution flows through packages/eval/src/maka-subject.ts, while external environments use packages/eval/src/external-subject.ts. Both adapters implement a common interface that the harness executor (packages/eval/src/harness-executor.ts) invokes to run the actual task code. Upon completion, results are persisted as CellAttempt records, storing scores, duration metrics, and artifacts.

Complete Implementation Example

Defining the Experiment Specification

A minimal experiment requires only a JSON definition specifying the Cartesian dimensions:

{
  "schemaVersion": "maka.eval.v1",
  "id": "bench-01",
  "benchmark": { "id": "my-bench", "version": "v1.0", "config": {} },
  "executor": { "kind": "harness", "config": {} },
  "subjects": [
    { "id": "maka-local", "kind": "maka", "credentials": [], "config": {} },
    { "id": "external-api", "kind": "external", "credentials": ["API_KEY"], "config": {} }
  ],
  "tasks": [
    { "id": "task-a", "input": "prompt-a", "config": {} },
    { "id": "task-b", "input": "prompt-b", "config": {} }
  ],
  "repetitions": 3,
  "budget": {},
  "verifier": { "reward": "score" }
}

This configuration generates 12 cells (2 tasks × 3 repetitions × 2 subjects).

Executing the Experiment Programmatically

Import the evaluation primitives and run the full matrix:

import { parseExperimentSpec } from '@maka/eval/src/spec.js';
import { runExperiment } from '@maka/eval/src/runner.js';
import { inMemoryStore } from '@maka/eval/src/attempt-store.js';
import { harnessExecutor } from '@maka/eval/src/harness-executor.js';
import { makaSubject, externalSubject } from '@maka/eval/src/maka-subject.js';

// 1️⃣ Parse the JSON spec
const spec = parseExperimentSpec(JSON.parse(mySpecJson));

// 2️⃣ Prepare adapters
const subjects = [makaSubject, externalSubject];

// 3️⃣ Execute the multi-arm experiment
const results = await runExperiment({
  spec,
  store: inMemoryStore(),
  executor: harnessExecutor,
  subjects,
});

// `results` is a Map keyed by cell ID → final `CellAttempt`
console.log('All cells executed:', Array.from(results.keys()));

Querying Specific Cell Results

Retrieve individual cell outcomes using the deterministic identifier format:

const cellId = 'task-a::2::maka-local'; // task-a, repetition 2, maka-local subject
const attempt = results.get(cellId);
if (attempt) {
  console.log('Score:', attempt.result.score);
  console.log('Duration (ms):', attempt.result.durationMs);
}

Summary

  • Declarative specifications in Maka define experiments through tasks, repetitions, and subjects, enabling version-controlled experimental design.
  • The expandExperiment function in packages/eval/src/experiment.ts generates the complete Cartesian product of task × repetition × subject dimensions into immutable ExperimentCell objects.
  • Deterministic cell identifiers ensure that execution attempts are idempotent and reproducible across runs.
  • The runExperiment orchestrator groups cells by task and repetition to enable parallel execution across subjects while respecting concurrency limits.
  • Subject adapters in maka-subject.ts and external-subject.ts provide a unified interface for executing tasks across heterogeneous environments.

Frequently Asked Questions

What is a declarative multi-arm experiment in Maka?

A declarative multi-arm experiment is a JSON specification that defines multiple "arms" or variants of a task to be tested across different subjects (execution environments) and statistical repetitions. Rather than imperatively scripting each run, researchers declare the experimental matrix, and Maka's evaluation framework handles the expansion and execution automatically.

How does Maka handle repetitions across different subjects?

Maka treats repetitions as a discrete dimension in the Cartesian expansion. For each repetition index (1 through N), the expandExperiment function creates distinct cells for every subject. During execution, the groupTaskCells logic ensures that all subjects for a specific task and repetition run concurrently, allowing for fair A/B comparisons under identical conditions.

What is the purpose of the ExperimentCell ID format?

The cell ID format taskId::repetition::subjectId serves as a deterministic, immutable primary key for a specific experimental observation. This structure guarantees that re-running the same specification produces identical cell identifiers, enabling result caching, idempotent retries, and longitudinal comparisons across experiment versions.

Can I execute specific cells without running the full experiment?

Yes. The runExperiment function accepts an optional filter parameter that allows targeting specific cell IDs. This capability supports debugging individual task-subject combinations or re-running failed cells without re-executing the entire Cartesian grid, significantly reducing computational costs during iterative development.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →