Apache Maka `packages/eval` Directory Responsibilities: Experiment Evaluation Framework

The packages/eval directory defines Apache Maka's core evaluation framework, responsible for experiment semantics, specification validation, trial orchestration, and result normalization.

The packages/eval directory serves as the central evaluation engine for the Apache Maka project, transforming declarative experiment definitions into executable trial matrices. This package establishes the hierarchy of Experiment → Cells → Attempts → Results, validates execution prerequisites, and delegates actual subject execution to external Runtime Hosts. Understanding the responsibilities of the packages/eval directory is essential for extending Maka's benchmarking capabilities or integrating custom evaluation pipelines.

Experiment Model and Specification Layer

The evaluation framework begins with a strict semantic model defined in the package documentation and implemented across several core modules.

Defining the Experiment Hierarchy

According to the package README, the packages/eval directory declares the high-level semantics governing how tasks, repetitions, subjects, and budgets combine into executable units. This hierarchy structures evaluation workflows as a Cartesian product of task × repetition × subject combinations referred to as cells, with each cell containing multiple attempts that yield results.

Parsing and Validating Specifications

The src/spec.ts module handles the ingestion of JSON experiment specifications, expanding them into the full matrix of execution cells. Before any trial begins, the specification parser verifies that the selected executor's prerequisites are satisfied, including path accessibility, relay file presence, and Docker availability. This validation ensures that the experiment model aligns with available infrastructure before resource allocation occurs.

Trial Execution Orchestration

While packages/eval does not execute subjects directly, it coordinates all aspects of trial preparation and launch through delegated Runtime Hosts.

The Runner and HarnessExecutor

In src/runner.ts, the Runner class creates a HarnessExecutor, prepares each cell's environment, and launches subjects via the Runtime Host. This orchestration layer manages the lifecycle of individual attempts, ensuring that execution contexts are properly initialized before subject code runs. The runner abstracts the complexity of cross-platform execution while maintaining strict contract adherence.

Pre-flight Environment Validation

Before launching trials, src/install-preflight.ts performs comprehensive environment checks to validate that required Docker images, Python environments, and other dependencies are present. These pre-flight checks prevent resource waste by catching configuration errors before the experiment matrix begins execution.

Toolchain Verification

The src/toolchain-verification.ts module ensures that bundled toolchains—such as the DeepSeek Harness—match pinned manifest versions before trials commence. This verification prevents version drift in evaluation substrates that could compromise result comparability across experiment runs.

Subject Management and Runtime Policies

The packages/eval directory abstracts diverse subject types behind unified interfaces while enforcing security constraints.

Subject Type Abstractions

Three primary modules model different subject categories:

This abstraction allows the evaluation framework to treat local benchmarks, containerized Harbor subjects, and external CLI tools as interchangeable execution targets.

Runtime Policy Enforcement

The src/maka-runtime-policy.ts module restricts available tools for benchmark subjects, typically enforcing a limited Bash subset and explicitly removing web-fetch capabilities. This policy enforcement ensures that benchmark results reflect algorithmic capability rather than external resource access, maintaining experimental integrity.

Provider Integration and Metering

For subjects requiring external Large Language Model (LLM) access, packages/eval implements mediation layers that handle authentication and usage tracking.

Provider Adapters

Two critical adapters bridge subjects and external services:

These adapters normalize disparate provider interfaces into consistent evaluation contracts while managing authentication and request routing.

Accurate Usage Tracking

The src/metering-checkpoint.ts module records usage snapshots from the proxy layer, ensuring accurate cost attribution even when subjects crash or terminate unexpectedly. This metering system captures resource consumption at checkpoints throughout execution, preventing data loss during partial failures.

Result Collection and Artifact Management

After trial completion, packages/eval normalizes raw outputs into standardized contracts and manages generated files.

Normalized Result Contracts

The src/result.ts module defines the Result kernel, which contains structured fields for score, usage, cost, duration, status, and artifacts. This module discards unstructured stdout data, presenting only a consistent result contract that downstream analysis tools can consume reliably.

Experiment Artifacts

The src/maka-artifacts.ts module records files generated during trials—including logs, checkpoints, and usage JSON—and computes SHA-256 digests for integrity verification. This artifact management ensures that experimental outputs are tamper-evident and reproducible.

Command-Line Interface

The src/cli.ts module exposes the public CLI interface for the evaluation framework, implementing commands such as:


# Execute a full experiment defined in experiment.json

maka eval run experiment.json --out ./maka-eval/run-001

# Replace a single failed cell

maka eval run experiment.json --out ./run-002 --cell cell-42

The CLI supports flags like --cell for targeted re-execution of specific experiment components, facilitating iterative debugging of failed trials without rerunning entire matrices.

Programmatic Usage Examples

Beyond CLI interaction, packages/eval exposes a programmatic API for custom evaluation pipelines:

import { loadSpec, runExperiment } from '@maka/eval';

// Load a spec file and run it programmatically
const spec = await loadSpec('experiment.json');
const results = await runExperiment(spec, {
  outDir: './run-programmatic',
});
console.log('All attempts completed:', results);

Accessing and analyzing attempt results:

import { readResults } from '@maka/eval/result';

const attempts = await readResults('./run-001');
attempts.forEach(a => {
  console.log(`Cell ${a.cellId} – status: ${a.result.status}, score: ${a.result.score}`);
});

Summary

  • packages/eval defines the Experiment → Cells → Attempts → Results hierarchy that structures all Maka evaluation workflows.
  • Specification parsing in src/spec.ts validates JSON experiments and expands them into executable Cartesian products of task, repetition, and subject combinations.
  • The Runner class in src/runner.ts orchestrates trial execution by preparing environments and delegating to Runtime Hosts rather than executing subjects directly.
  • Subject abstractions in maka-subject.ts, harbor-maka-subject.ts, and external-subject.ts unify diverse execution targets behind common interfaces.
  • Provider adapters and metering checkpoints ensure accurate cost attribution and normalized access to external LLM services.
  • Result normalization in src/result.ts enforces consistent score-cost-duration contracts while discarding unstructured output data.

Frequently Asked Questions

What is the relationship between packages/eval and Runtime Hosts?

The packages/eval directory owns experiment semantics and orchestration but delegates actual subject execution to Runtime Hosts such as Harbor, Pier, or external wrappers. The Runner class in src/runner.ts prepares execution environments and invokes these hosts, which handle the low-level process isolation and resource management.

How does packages/eval handle metering when a subject crashes?

The src/metering-checkpoint.ts module records usage snapshots from the proxy layer throughout execution, ensuring accurate cost attribution even when subjects crash or terminate unexpectedly. This checkpoint system captures resource consumption data at intervals rather than relying solely on final process exit codes.

Can I use packages/eval without the Maka CLI?

Yes. While src/cli.ts provides the maka eval run command-line interface, you can import @maka/eval programmatically using loadSpec() and runExperiment() to execute experiments within custom Node.js applications or automated testing pipelines.

How are toolchains verified before trial execution?

The src/toolchain-verification.ts module validates that bundled toolchains—such as the DeepSeek Harness—match the pinned manifest versions specified in the experiment configuration. This verification occurs after pre-flight checks but before subject launch, preventing version mismatches that could invalidate benchmark results.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →