# What Is the Responsibility of the @maka/eval Package in Apache Maka?

> Discover the core responsibility of the @maka/eval package in Apache Maka. Learn how it acts as the semantic engine for experiment schemas, execution, and validation.

- Repository: [The Apache Software Foundation/maka](https://github.com/apache/maka)
- Tags: deep-dive
- Published: 2026-08-22

---

**The `@maka/eval` package serves as the semantic engine for the Maka platform, responsible for defining experiment schemas, expanding specifications into executable cells, validating execution environments, and orchestrating trial runs while delegating actual subject execution to Runtime hosts.**

The `@maka/eval` package in the Apache Maka repository governs how experiments are structured, validated, and launched. It establishes the JSON schema for experiment definitions and manages the complete lifecycle from specification to result aggregation. According to the source code in [`packages/eval/src/experiment.ts`](https://github.com/apache/maka/blob/main/packages/eval/src/experiment.ts), this package acts as the central coordinator that translates high-level experiment configurations into discrete, executable units without directly executing Maka subjects itself.

## Defining Experiment Semantics with ExperimentSpec

At its core, the responsibility of the `@maka/eval` package begins with **modeling experiments** through strict type definitions. The package defines `ExperimentSpec`, the canonical JSON schema that describes an entire experiment configuration including tasks, subjects, repetitions, and executor requirements.

As documented in [`packages/eval/README.md`](https://github.com/apache/maka/blob/main/packages/eval/README.md), the package "owns experiment semantics." This ownership manifests in [`src/experiment.ts`](https://github.com/apache/maka/blob/main/src/experiment.ts), where the TypeScript interfaces establish the contract for valid experiment configurations. The specification includes definitions for benchmarks, executors, and subject mappings that downstream components consume.

## Expanding Specifications into Executable Cells

One critical responsibility is **expanding abstract specifications into concrete work units**. The `expandExperiment()` function in [`src/experiment.ts`](https://github.com/apache/maka/blob/main/src/experiment.ts) (lines 84-100) takes an `ExperimentSpec` and generates the complete set of `ExperimentCell` objects representing every combination of task, repetition, and subject.

This expansion creates the **Cartesian product** of tasks × repetitions × subjects. Each resulting cell contains a unique identifier and explicit references to its benchmark, executor, subject, task, and repetition index. This deterministic enumeration ensures reproducible experiment execution across distributed Runtime hosts.

```typescript
// Expand the spec into individual cells for execution
import { expandExperiment } from '@maka/eval';

const cells = expandExperiment(spec);
console.log('Total cells:', cells.length);
// Each cell = { id, experimentId, benchmark, executor, subject, task, repetition, … }

```

## CLI Interface for Experiment Orchestration

The package exposes a public **command-line interface** (`maka eval run …`) implemented in [`src/cli.ts`](https://github.com/apache/maka/blob/main/src/cli.ts). This CLI serves as the primary entry point for researchers executing experiments, but it never runs Maka subjects directly. Instead, it validates the experiment spec, checks executor prerequisites including machine paths and Docker availability, and delegates execution to Runtime hosts.

The programmatic equivalent `runExperiment()` function provides the same orchestration capabilities within TypeScript applications:

```typescript
// Run the experiment via the programmatic API
import { runExperiment } from '@maka/eval';

await runExperiment(spec, {
  outDir: '.maka-eval/run-001',
});
// Validates executor, verifies toolchains, starts trials
// Writes per-cell attempt logs and aggregated results

```

## Environment Verification and Toolchain Validation

Before any trial begins, `@maka/eval` performs rigorous **environment verification** through [`src/toolchain-verification.ts`](https://github.com/apache/maka/blob/main/src/toolchain-verification.ts). This module checks executor prerequisites including:

- Toolchain versions and compatibility
- Harbor/Pier Python distribution availability  
- Docker environment status
- Bundled toolchain integrity

These pre-flight checks ensure experiments run on known, reproducible environments. The package also surface-validates external subjects to guarantee they meet the experiment's execution requirements before delegation to Runtime hosts.

## Managing Experiment Results and Artifacts

The final responsibility encompasses **result handling** defined in [`src/result.ts`](https://github.com/apache/maka/blob/main/src/result.ts). This module specifies the shape of experiment outcomes including score metrics, resource usage, cost tracking, execution status, and artifact locations. The package provides utilities for persisting individual attempt logs and aggregating final outcomes across all cells.

The [`src/runner.ts`](https://github.com/apache/maka/blob/main/src/runner.ts) module coordinates the actual trial launches, interfacing between the expanded cell specifications and the verification systems to ensure each experiment unit completes with properly structured output data.

```typescript
// Build an ExperimentSpec from JSON configuration
import { ExperimentSpec } from '@maka/eval';
import fs from 'fs';

const spec: ExperimentSpec = JSON.parse(
  fs.readFileSync('experiment.json', 'utf-8')
);

```

## Summary

- **`@maka/eval`** defines the semantic model for Maka experiments through `ExperimentSpec` in [`src/experiment.ts`](https://github.com/apache/maka/blob/main/src/experiment.ts).
- **`expandExperiment()`** generates the complete Cartesian product of tasks, repetitions, and subjects as `ExperimentCell` objects.
- The **CLI and programmatic API** orchestrate trials without executing subjects directly, delegating to Runtime hosts instead.
- **[`src/toolchain-verification.ts`](https://github.com/apache/maka/blob/main/src/toolchain-verification.ts)** validates execution environments including Docker, Python distributions, and toolchains before trials begin.
- Result structures and persistence utilities in [`src/result.ts`](https://github.com/apache/maka/blob/main/src/result.ts) standardize outcome aggregation across distributed experiment runs.

## Frequently Asked Questions

### What distinguishes @maka/eval from the Maka Runtime host?

The `@maka/eval` package defines *what* experiments run and *how* they are structured, while the Runtime host handles the actual execution of Maka subjects. The eval package validates specifications, expands them into cells, and orchestrates trials, but intentionally delegates subject execution to Runtime hosts to maintain separation between experiment semantics and execution environments.

### How does the expandExperiment function determine the number of cells?

The `expandExperiment()` function calculates the Cartesian product of three dimensions: tasks, repetitions, and subjects. For an experiment with *T* tasks, *R* repetitions, and *S* subjects, it generates exactly *T × R × S* `ExperimentCell` objects, as implemented in [`src/experiment.ts`](https://github.com/apache/maka/blob/main/src/experiment.ts) lines 84-100.

### What prerequisites does @maka/eval verify before running trials?

According to [`src/toolchain-verification.ts`](https://github.com/apache/maka/blob/main/src/toolchain-verification.ts), the package verifies executor machine paths, Docker availability, bundled toolchains, and Harbor/Pier Python distributions. These checks ensure the execution environment matches the experiment requirements and can reproducibly run the specified subjects.

### Where are experiment results structured and stored?

Result definitions reside in [`src/result.ts`](https://github.com/apache/maka/blob/main/src/result.ts), which exports types for scores, usage metrics, costs, statuses, and artifacts. The package writes per-cell attempt logs to the specified output directory and aggregates final results after all cells complete, providing a standardized interface for experiment outcomes.