# Structure of an Experiment in the Maka Evaluation Framework

> Understand the structure of an experiment in the Maka evaluation framework. Learn how Maka models experiments as typed JSON objects conforming to the ExperimentSpec schema.

- Repository: [The Apache Software Foundation/maka](https://github.com/apache/maka)
- Tags: architecture
- Published: 2026-09-04

---

**The Maka evaluation framework models an experiment as a typed JSON object conforming to the `ExperimentSpec` schema, which is validated by `parseExperimentSpec` and expanded into executable `ExperimentCell` units.**

The Apache Maka project provides a rigorous evaluation framework for benchmarking distributed systems and query engines. Understanding the structure of an experiment in the Maka evaluation framework enables precise control over workloads, execution environments, and result verification. At its core, the framework enforces a strict contract defined in TypeScript types and validated at runtime.

## Core Schema and Top-Level Fields

The `ExperimentSpec` type defined in [`packages/eval/src/experiment.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/experiment.ts) establishes the canonical structure for all experiment definitions. This typed JSON object contains mandatory fields that define the benchmark, subjects, and execution context, plus optional fields for concurrency control.

### Mandatory Fields

Every experiment must include these top-level properties:

- **schemaVersion** — Must be the literal string `"maka.eval.v1"` to identify the schema version used by the parser.
- **id** — A unique string identifier that must match the regular expression `/^[A-Za-z0-9][A-Za-z0-9._-]*$/`.
- **benchmark** — An object specifying `{ id, version, config }` to determine which benchmark suite runs and its parameters.
- **executor** — Describes the execution environment with `kind` (e.g., `"docker"`) and an arbitrary `config` object.
- **subjects** — An array of entities that execute tasks. Each subject requires `{ id, kind, credentials, config }`, where `kind` distinguishes between Maka instances and external services.
- **tasks** — An array of workload definitions, each containing `{ id, input, config }` to specify what operations to perform.
- **repetitions** — A positive integer indicating how many times each subject-task combination executes.
- **budget** — A `JsonObject` defining constraints such as time limits or resource consumption caps.
- **verifier** — A `JsonObject` containing validation rules applied to execution results.

### Optional Execution Control

- **execution** — An optional object containing `{ maxConcurrentTaskGroups }`. When omitted, the framework defaults this value to `1`, enforcing sequential execution of task groups.

## Validation and Type Safety

The `parseExperimentSpec` function in [`packages/eval/src/spec.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/spec.ts) transforms raw JSON into a validated `ExperimentSpec` instance. This function performs rigorous checks for field presence, type correctness, and identifier uniqueness according to the regex pattern defined above. 

Crucially, `parseExperimentSpec` returns a **deep-frozen** object, guaranteeing immutability at runtime. This prevents accidental mutation of the experiment configuration during the evaluation lifecycle.

## Expansion into Executable Cells

Before runtime execution, the framework expands a single `ExperimentSpec` into a flat list of `ExperimentCell` objects. The `expandExperiment` function in [`packages/eval/src/experiment.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/experiment.ts) generates these cells by computing the Cartesian product of tasks, repetitions, and subjects:

```typescript
export function expandExperiment(spec: ExperimentSpec): ExperimentCell[] {
  return spec.tasks.flatMap(task =>
    Array.from({ length: spec.repetitions }, (_, r) => r + 1).flatMap(repetition =>
      spec.subjects.map(subject => ({
        id: `${task.id}::${repetition}::${subject.id}`,
        experimentId: spec.id,
        benchmark: spec.benchmark,
        executor: spec.executor,
        subject,
        task,
        repetition,
        budget: spec.budget,
        verifier: spec.verifier,
      }))
    )
  );
}

```

Each `ExperimentCell` receives a unique composite ID formatted as `${task.id}::${repetition}::${subject.id}`. These cells are stored in an attempt store ([`packages/eval/src/attempt-store.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/attempt-store.ts)) and fed to the executor runtime for isolated, trackable execution.

## Practical Implementation Examples

### Defining an Experiment JSON

```json
{
  "schemaVersion": "maka.eval.v1",
  "id": "example‑exp‑001",
  "benchmark": {
    "id": "tpch",
    "version": "1.0",
    "config": { "scaleFactor": 0.1 }
  },
  "executor": {
    "kind": "docker",
    "config": { "image": "maka‑executor:latest" }
  },
  "subjects": [
    {
      "id": "subject‑a",
      "kind": "maka",
      "credentials": ["cred‑a"],
      "config": { "host": "localhost", "port": 8080 }
    }
  ],
  "tasks": [
    { "id": "q1", "input": "SELECT * FROM lineitem", "config": {} }
  ],
  "repetitions": 3,
  "budget": { "time": "5m" },
  "verifier": { "type": "none" }
}

```

### Parsing and Executing in Code

```typescript
import { readFile } from 'node:fs/promises';
import { parseExperimentSpec } from './spec.js';
import { expandExperiment } from './experiment.js';

// Load a JSON file containing the experiment definition
const raw = JSON.parse(await readFile('my-experiment.json', 'utf8'));

// Validate and obtain a typed spec
const spec = parseExperimentSpec(raw);

// Produce the list of executable cells
const cells = expandExperiment(spec);

console.log(`Experiment "${spec.id}" expands to ${cells.length} cells`);
cells.forEach(cell => console.log(cell.id));

```

### Persisting Experiment State

```typescript
import { openExperimentDirectory } from './experiment-directory.js';

// Choose a directory where all attempts for this experiment will be stored
const dir = await openExperimentDirectory('/tmp/maka-experiment', spec);

console.log('Experiment stored at', dir.root);

```

## Summary

- The Maka evaluation framework enforces a strict JSON schema through the `ExperimentSpec` type in [`packages/eval/src/experiment.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/experiment.ts).
- `parseExperimentSpec` in [`packages/eval/src/spec.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/spec.ts) validates identifiers against `/^[A-Za-z0-9][A-Za-z0-9._-]*$/` and returns deep-frozen, immutable objects.
- Experiments expand into discrete `ExperimentCell` units via `expandExperiment`, creating a runnable matrix of task-subject-repetition combinations.
- Optional `execution` controls allow configuration of `maxConcurrentTaskGroups` for parallelism tuning.
- The framework integrates with [`packages/eval/src/experiment-directory.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/experiment-directory.ts) for persistent state management across evaluation attempts.

## Frequently Asked Questions

### What is the required schema version for Maka experiments?

The `schemaVersion` field must be set to the literal string `"maka.eval.v1"`. This identifier signals to `parseExperimentSpec` that the JSON conforms to the expected structure defined in [`packages/eval/src/experiment.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/experiment.ts).

### How does the Maka framework validate experiment identifiers?

According to the validation logic in [`packages/eval/src/spec.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/spec.ts), all identifiers must match the regular expression `/^[A-Za-z0-9][A-Za-z0-9._-]*$/`. The parser checks for uniqueness across subjects and tasks, throwing an error if duplicates or invalid characters are detected.

### What happens if the execution field is omitted from the experiment specification?

When the optional `execution` field is absent, `parseExperimentSpec` canonicalizes the value of `maxConcurrentTaskGroups` to `1`. This default ensures sequential execution of task groups unless explicitly configured for parallelism.

### How are experiment cells generated from the specification?

The `expandExperiment` function performs a Cartesian expansion across three dimensions: tasks, repetitions, and subjects. For each combination, it creates an `ExperimentCell` with a unique composite ID (`${task.id}::${repetition}::${subject.id}`), enabling the executor runtime to process each unit independently while tracking its origin within the broader experiment structure.