What Is the Responsibility of the @maka/eval Package in Apache Maka?

The @maka/eval package serves as the semantic engine for the Maka platform, responsible for defining experiment schemas, expanding specifications into executable cells, validating execution environments, and orchestrating trial runs while delegating actual subject execution to Runtime hosts.

The @maka/eval package in the Apache Maka repository governs how experiments are structured, validated, and launched. It establishes the JSON schema for experiment definitions and manages the complete lifecycle from specification to result aggregation. According to the source code in packages/eval/src/experiment.ts, this package acts as the central coordinator that translates high-level experiment configurations into discrete, executable units without directly executing Maka subjects itself.

Defining Experiment Semantics with ExperimentSpec

At its core, the responsibility of the @maka/eval package begins with modeling experiments through strict type definitions. The package defines ExperimentSpec, the canonical JSON schema that describes an entire experiment configuration including tasks, subjects, repetitions, and executor requirements.

As documented in packages/eval/README.md, the package "owns experiment semantics." This ownership manifests in src/experiment.ts, where the TypeScript interfaces establish the contract for valid experiment configurations. The specification includes definitions for benchmarks, executors, and subject mappings that downstream components consume.

Expanding Specifications into Executable Cells

One critical responsibility is expanding abstract specifications into concrete work units. The expandExperiment() function in src/experiment.ts (lines 84-100) takes an ExperimentSpec and generates the complete set of ExperimentCell objects representing every combination of task, repetition, and subject.

This expansion creates the Cartesian product of tasks × repetitions × subjects. Each resulting cell contains a unique identifier and explicit references to its benchmark, executor, subject, task, and repetition index. This deterministic enumeration ensures reproducible experiment execution across distributed Runtime hosts.

// Expand the spec into individual cells for execution
import { expandExperiment } from '@maka/eval';

const cells = expandExperiment(spec);
console.log('Total cells:', cells.length);
// Each cell = { id, experimentId, benchmark, executor, subject, task, repetition, … }

CLI Interface for Experiment Orchestration

The package exposes a public command-line interface (maka eval run …) implemented in src/cli.ts. This CLI serves as the primary entry point for researchers executing experiments, but it never runs Maka subjects directly. Instead, it validates the experiment spec, checks executor prerequisites including machine paths and Docker availability, and delegates execution to Runtime hosts.

The programmatic equivalent runExperiment() function provides the same orchestration capabilities within TypeScript applications:

// Run the experiment via the programmatic API
import { runExperiment } from '@maka/eval';

await runExperiment(spec, {
  outDir: '.maka-eval/run-001',
});
// Validates executor, verifies toolchains, starts trials
// Writes per-cell attempt logs and aggregated results

Environment Verification and Toolchain Validation

Before any trial begins, @maka/eval performs rigorous environment verification through src/toolchain-verification.ts. This module checks executor prerequisites including:

  • Toolchain versions and compatibility
  • Harbor/Pier Python distribution availability
  • Docker environment status
  • Bundled toolchain integrity

These pre-flight checks ensure experiments run on known, reproducible environments. The package also surface-validates external subjects to guarantee they meet the experiment's execution requirements before delegation to Runtime hosts.

Managing Experiment Results and Artifacts

The final responsibility encompasses result handling defined in src/result.ts. This module specifies the shape of experiment outcomes including score metrics, resource usage, cost tracking, execution status, and artifact locations. The package provides utilities for persisting individual attempt logs and aggregating final outcomes across all cells.

The src/runner.ts module coordinates the actual trial launches, interfacing between the expanded cell specifications and the verification systems to ensure each experiment unit completes with properly structured output data.

// Build an ExperimentSpec from JSON configuration
import { ExperimentSpec } from '@maka/eval';
import fs from 'fs';

const spec: ExperimentSpec = JSON.parse(
  fs.readFileSync('experiment.json', 'utf-8')
);

Summary

  • @maka/eval defines the semantic model for Maka experiments through ExperimentSpec in src/experiment.ts.
  • expandExperiment() generates the complete Cartesian product of tasks, repetitions, and subjects as ExperimentCell objects.
  • The CLI and programmatic API orchestrate trials without executing subjects directly, delegating to Runtime hosts instead.
  • src/toolchain-verification.ts validates execution environments including Docker, Python distributions, and toolchains before trials begin.
  • Result structures and persistence utilities in src/result.ts standardize outcome aggregation across distributed experiment runs.

Frequently Asked Questions

What distinguishes @maka/eval from the Maka Runtime host?

The @maka/eval package defines what experiments run and how they are structured, while the Runtime host handles the actual execution of Maka subjects. The eval package validates specifications, expands them into cells, and orchestrates trials, but intentionally delegates subject execution to Runtime hosts to maintain separation between experiment semantics and execution environments.

How does the expandExperiment function determine the number of cells?

The expandExperiment() function calculates the Cartesian product of three dimensions: tasks, repetitions, and subjects. For an experiment with T tasks, R repetitions, and S subjects, it generates exactly T × R × S ExperimentCell objects, as implemented in src/experiment.ts lines 84-100.

What prerequisites does @maka/eval verify before running trials?

According to src/toolchain-verification.ts, the package verifies executor machine paths, Docker availability, bundled toolchains, and Harbor/Pier Python distributions. These checks ensure the execution environment matches the experiment requirements and can reproducibly run the specified subjects.

Where are experiment results structured and stored?

Result definitions reside in src/result.ts, which exports types for scores, usage metrics, costs, statuses, and artifacts. The package writes per-cell attempt logs to the specified output directory and aggregates final results after all cells complete, providing a standardized interface for experiment outcomes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →