# Apache Maka `packages/eval` Directory Responsibilities: Experiment Evaluation Framework

> Explore the Apache Maka packages eval directory, its core evaluation framework, experiment semantics, validation, trial orchestration, and result normalization responsibilities.

- Repository: [The Apache Software Foundation/maka](https://github.com/apache/maka)
- Tags: internals
- Published: 2026-08-24

---

**The `packages/eval` directory defines Apache Maka's core evaluation framework, responsible for experiment semantics, specification validation, trial orchestration, and result normalization.**

The `packages/eval` directory serves as the central evaluation engine for the Apache Maka project, transforming declarative experiment definitions into executable trial matrices. This package establishes the hierarchy of **Experiment → Cells → Attempts → Results**, validates execution prerequisites, and delegates actual subject execution to external Runtime Hosts. Understanding the responsibilities of the `packages/eval` directory is essential for extending Maka's benchmarking capabilities or integrating custom evaluation pipelines.

## Experiment Model and Specification Layer

The evaluation framework begins with a strict semantic model defined in the package documentation and implemented across several core modules.

### Defining the Experiment Hierarchy

According to the package README, the `packages/eval` directory declares the high-level semantics governing how tasks, repetitions, subjects, and budgets combine into executable units. This hierarchy structures evaluation workflows as a Cartesian product of *task × repetition × subject* combinations referred to as **cells**, with each cell containing multiple **attempts** that yield **results**.

### Parsing and Validating Specifications

The [`src/spec.ts`](https://github.com/apache/maka/blob/main/src/spec.ts) module handles the ingestion of JSON experiment specifications, expanding them into the full matrix of execution cells. Before any trial begins, the specification parser verifies that the selected executor's prerequisites are satisfied, including path accessibility, relay file presence, and Docker availability. This validation ensures that the experiment model aligns with available infrastructure before resource allocation occurs.

## Trial Execution Orchestration

While `packages/eval` does not execute subjects directly, it coordinates all aspects of trial preparation and launch through delegated Runtime Hosts.

### The Runner and HarnessExecutor

In [`src/runner.ts`](https://github.com/apache/maka/blob/main/src/runner.ts), the `Runner` class creates a `HarnessExecutor`, prepares each cell's environment, and launches subjects via the Runtime Host. This orchestration layer manages the lifecycle of individual attempts, ensuring that execution contexts are properly initialized before subject code runs. The runner abstracts the complexity of cross-platform execution while maintaining strict contract adherence.

### Pre-flight Environment Validation

Before launching trials, [`src/install-preflight.ts`](https://github.com/apache/maka/blob/main/src/install-preflight.ts) performs comprehensive environment checks to validate that required Docker images, Python environments, and other dependencies are present. These pre-flight checks prevent resource waste by catching configuration errors before the experiment matrix begins execution.

### Toolchain Verification

The [`src/toolchain-verification.ts`](https://github.com/apache/maka/blob/main/src/toolchain-verification.ts) module ensures that bundled toolchains—such as the DeepSeek Harness—match pinned manifest versions before trials commence. This verification prevents version drift in evaluation substrates that could compromise result comparability across experiment runs.

## Subject Management and Runtime Policies

The `packages/eval` directory abstracts diverse subject types behind unified interfaces while enforcing security constraints.

### Subject Type Abstractions

Three primary modules model different subject categories:

- **[`src/maka-subject.ts`](https://github.com/apache/maka/blob/main/src/maka-subject.ts)**: Handles internal Maka benchmark subjects with specific credential bindings
- **[`src/harbor-maka-subject.ts`](https://github.com/apache/maka/blob/main/src/harbor-maka-subject.ts)**: Manages Harbor-hosted subjects within containerized environments
- **[`src/external-subject.ts`](https://github.com/apache/maka/blob/main/src/external-subject.ts)**: Wraps arbitrary external commands as evaluable subjects

This abstraction allows the evaluation framework to treat local benchmarks, containerized Harbor subjects, and external CLI tools as interchangeable execution targets.

### Runtime Policy Enforcement

The [`src/maka-runtime-policy.ts`](https://github.com/apache/maka/blob/main/src/maka-runtime-policy.ts) module restricts available tools for benchmark subjects, typically enforcing a limited Bash subset and explicitly removing web-fetch capabilities. This policy enforcement ensures that benchmark results reflect algorithmic capability rather than external resource access, maintaining experimental integrity.

## Provider Integration and Metering

For subjects requiring external Large Language Model (LLM) access, `packages/eval` implements mediation layers that handle authentication and usage tracking.

### Provider Adapters

Two critical adapters bridge subjects and external services:

- **[`src/provider-web-tool-surface.ts`](https://github.com/apache/maka/blob/main/src/provider-web-tool-surface.ts)**: Implements the web-tool surface for LLM provider interaction
- **[`src/provider-admission.ts`](https://github.com/apache/maka/blob/main/src/provider-admission.ts)**: Handles admission events and result framing from provider APIs

These adapters normalize disparate provider interfaces into consistent evaluation contracts while managing authentication and request routing.

### Accurate Usage Tracking

The [`src/metering-checkpoint.ts`](https://github.com/apache/maka/blob/main/src/metering-checkpoint.ts) module records usage snapshots from the proxy layer, ensuring accurate cost attribution even when subjects crash or terminate unexpectedly. This metering system captures resource consumption at checkpoints throughout execution, preventing data loss during partial failures.

## Result Collection and Artifact Management

After trial completion, `packages/eval` normalizes raw outputs into standardized contracts and manages generated files.

### Normalized Result Contracts

The [`src/result.ts`](https://github.com/apache/maka/blob/main/src/result.ts) module defines the `Result` kernel, which contains structured fields for **score**, **usage**, **cost**, **duration**, **status**, and **artifacts**. This module discards unstructured stdout data, presenting only a consistent result contract that downstream analysis tools can consume reliably.

### Experiment Artifacts

The [`src/maka-artifacts.ts`](https://github.com/apache/maka/blob/main/src/maka-artifacts.ts) module records files generated during trials—including logs, checkpoints, and usage JSON—and computes SHA-256 digests for integrity verification. This artifact management ensures that experimental outputs are tamper-evident and reproducible.

## Command-Line Interface

The [`src/cli.ts`](https://github.com/apache/maka/blob/main/src/cli.ts) module exposes the public CLI interface for the evaluation framework, implementing commands such as:

```bash

# Execute a full experiment defined in experiment.json

maka eval run experiment.json --out ./maka-eval/run-001

# Replace a single failed cell

maka eval run experiment.json --out ./run-002 --cell cell-42

```

The CLI supports flags like `--cell` for targeted re-execution of specific experiment components, facilitating iterative debugging of failed trials without rerunning entire matrices.

## Programmatic Usage Examples

Beyond CLI interaction, `packages/eval` exposes a programmatic API for custom evaluation pipelines:

```typescript
import { loadSpec, runExperiment } from '@maka/eval';

// Load a spec file and run it programmatically
const spec = await loadSpec('experiment.json');
const results = await runExperiment(spec, {
  outDir: './run-programmatic',
});
console.log('All attempts completed:', results);

```

Accessing and analyzing attempt results:

```typescript
import { readResults } from '@maka/eval/result';

const attempts = await readResults('./run-001');
attempts.forEach(a => {
  console.log(`Cell ${a.cellId} – status: ${a.result.status}, score: ${a.result.score}`);
});

```

## Summary

- **`packages/eval` defines the Experiment → Cells → Attempts → Results hierarchy** that structures all Maka evaluation workflows.
- **Specification parsing in [`src/spec.ts`](https://github.com/apache/maka/blob/main/src/spec.ts)** validates JSON experiments and expands them into executable Cartesian products of task, repetition, and subject combinations.
- **The `Runner` class in [`src/runner.ts`](https://github.com/apache/maka/blob/main/src/runner.ts)** orchestrates trial execution by preparing environments and delegating to Runtime Hosts rather than executing subjects directly.
- **Subject abstractions** in [`maka-subject.ts`](https://github.com/apache/maka/blob/main/maka-subject.ts), [`harbor-maka-subject.ts`](https://github.com/apache/maka/blob/main/harbor-maka-subject.ts), and [`external-subject.ts`](https://github.com/apache/maka/blob/main/external-subject.ts) unify diverse execution targets behind common interfaces.
- **Provider adapters and metering checkpoints** ensure accurate cost attribution and normalized access to external LLM services.
- **Result normalization in [`src/result.ts`](https://github.com/apache/maka/blob/main/src/result.ts)** enforces consistent score-cost-duration contracts while discarding unstructured output data.

## Frequently Asked Questions

### What is the relationship between `packages/eval` and Runtime Hosts?

The `packages/eval` directory owns experiment semantics and orchestration but delegates actual subject execution to Runtime Hosts such as Harbor, Pier, or external wrappers. The `Runner` class in [`src/runner.ts`](https://github.com/apache/maka/blob/main/src/runner.ts) prepares execution environments and invokes these hosts, which handle the low-level process isolation and resource management.

### How does `packages/eval` handle metering when a subject crashes?

The [`src/metering-checkpoint.ts`](https://github.com/apache/maka/blob/main/src/metering-checkpoint.ts) module records usage snapshots from the proxy layer throughout execution, ensuring accurate cost attribution even when subjects crash or terminate unexpectedly. This checkpoint system captures resource consumption data at intervals rather than relying solely on final process exit codes.

### Can I use `packages/eval` without the Maka CLI?

Yes. While [`src/cli.ts`](https://github.com/apache/maka/blob/main/src/cli.ts) provides the `maka eval run` command-line interface, you can import `@maka/eval` programmatically using `loadSpec()` and `runExperiment()` to execute experiments within custom Node.js applications or automated testing pipelines.

### How are toolchains verified before trial execution?

The [`src/toolchain-verification.ts`](https://github.com/apache/maka/blob/main/src/toolchain-verification.ts) module validates that bundled toolchains—such as the DeepSeek Harness—match the pinned manifest versions specified in the experiment configuration. This verification occurs after pre-flight checks but before subject launch, preventing version mismatches that could invalidate benchmark results.