Apache Maka Experiment Cells, Attempts, and Result Kernels: Eval System Architecture

Apache Maka's evaluation subsystem isolates experiments inside Docker-based cells, executes them as discrete retryable attempts, and captures deterministic outcomes as serializable result kernels.

Apache Maka's evaluation (Eval) subsystem provides a robust framework for running isolated experiments on subject code. The architecture centers on three foundational concepts: experiment cells, attempts, and result kernels, which together ensure reproducible, policy-driven execution environments. This system is implemented primarily within the harbor and pier eval frameworks in the Apache Maka repository.

Core Architecture Components

Experiment Cells

An experiment cell is an isolated execution environment provisioned when a trial starts. According to the implementation in packages/eval/harbor/run_trial.py, cells are Docker-based containers that encapsulate the subject binary, the selected Eval framework (harbor or pier), and any required policy artefacts. This isolation ensures that each trial runs independently of the host system and other experiments.

Attempts

An attempt represents a single execution of the subject code under a specific policy or configuration. The relay agent (packages/eval/harbor/relay_agent.py) drives the subject inside the cell, spawning a new attempt for every policy change or retry. Each attempt runs the subject under the current policy, watches for cancellation signals, and records any runtime exceptions. This design allows for fine-grained retries and policy variations within a single experiment.

Result Kernels

A result kernel is a lightweight, serialisable data structure that captures the outcome of an attempt. After an attempt finishes, the relay agent collects execution metadata—including exit status, stdout/stderr logs, resource metrics, and any artefacts—and stores them in a kernel object. These kernels are passed back to the trial controller, which aggregates them from all attempts to form the final experiment report.

Execution Flow

The evaluation process follows a strict three-phase lifecycle that ensures deterministic, reproducible results.

Cell Provisioning

When run_trial() is invoked, the system initializes the framework using install() and selected() from packages/eval/harbor/eval_framework.py, then provisions a new cell. This containerized environment contains all dependencies required for the subject code execution.

Attempt Execution

Inside the cell, the relay agent manages attempt lifecycle through run_attempt(). The agent executes the subject via cell.exec_subject(policy), monitors for asyncio.CancelledError to handle premature termination, and ensures proper cleanup regardless of outcome.

Kernel Aggregation

Upon attempt completion, the relay agent constructs a ResultKernel instance containing the attempt ID, return code, output streams, and collected metrics. The trial controller accumulates these kernels across all attempts, enabling downstream analysis, comparison, or persistence of experiment results.

Implementation Details

Framework Selection

The eval framework selector in packages/eval/harbor/eval_framework.py registers the chosen framework (harbor or pier), determining the specific cell implementation and execution policies for the experiment.

Relay Agent Operations

The packages/eval/harbor/relay_agent.py file implements the core attempt-handling logic. It manages subject execution inside the cell, handles cancellation scenarios by raising RuntimeError when execution is aborted, and constructs the result kernel through cell.collect_metrics() and execution metadata.

Test Coverage

The test suite validates these components across multiple scenarios:

Code Examples

Running a Trial

The following example demonstrates initializing the framework and executing a trial that returns result kernels:


# packages/eval/harbor/run_trial.py – high‑level entry point

from maka.eval.harbor.run_trial import run_trial

# Initialise the framework (harbor or pier)

from maka.eval.harbor.eval_framework import install, selected
install("harbor")

# Execute a trial; `subject_argv` is the command to evaluate

kernels = run_trial(subject_argv=["/app/my_program", "--mode", "test"])
for kernel in kernels:
    print(f"Attempt {kernel.attempt_id}: exit={kernel.exit_code}")

Relay Agent Attempt Handling

This simplified illustration shows how the relay agent executes attempts and constructs result kernels:


# packages/eval/harbor/relay_agent.py – simplified illustration

async def run_attempt(cell, policy):
    try:
        execution = await cell.exec_subject(policy)
        await execution.wait()
    except asyncio.CancelledError:
        raise RuntimeError("Maka Eval subject execution was cancelled")
    finally:
        # Build the result kernel from execution metadata

        kernel = ResultKernel(
            attempt_id=cell.next_attempt_id(),
            exit_code=execution.returncode,
            stdout=await execution.stdout.read(),
            stderr=await execution.stderr.read(),
            metrics=cell.collect_metrics(),
        )
        return kernel

Aggregating Results

After all attempts complete, the trial controller aggregates the kernels:


# Trial controller aggregation pattern

all_kernels = []
while not trial.complete():
    kernel = await relay_agent.run_attempt(current_cell, current_policy)
    all_kernels.append(kernel)

# `all_kernels` now holds the result kernels for the experiment

Summary

  • Experiment cells provide Docker-based isolation for trials, ensuring that subject code execution cannot affect the host system or other experiments.
  • Attempts represent individual executions under specific policies, enabling retry logic and policy variations within a single experiment lifecycle.
  • Result kernels offer deterministic, serializable snapshots of attempt outcomes, capturing exit codes, logs, metrics, and artefacts for downstream analysis.
  • The architecture separates concerns between cell provisioning (run_trial.py), attempt execution (relay_agent.py), and framework selection (eval_framework.py).

Frequently Asked Questions

What is the difference between a cell and an attempt in Apache Maka?

An experiment cell is the isolated container environment (Docker-based) that hosts the execution, while an attempt is a single run of the subject code within that cell under a specific policy. A cell may contain multiple attempts if retries or policy variations are required.

How does the relay agent handle execution failures?

The relay agent (relay_agent.py) wraps subject execution in try-catch blocks to handle asyncio.CancelledError and other exceptions. When cancellation occurs, it raises a RuntimeError with the message "Maka Eval subject execution was cancelled". Regardless of success or failure, the agent always constructs a result kernel in the finally block to ensure outcome capture.

What data structure stores experiment outcomes?

Outcomes are stored in result kernels, lightweight serialisable objects created at the end of each attempt. These kernels contain the attempt_id, exit_code, stdout, stderr, and metrics collected during execution, enabling deterministic comparison across different attempts.

Where is the result kernel defined in the source code?

While ResultKernel is constructed in packages/eval/harbor/relay_agent.py, its structure emerges from the kernel assembly logic in that file rather than a separate class definition file. The constructor receives fields including attempt_id, exit_code, output streams, and metrics gathered via cell.collect_metrics().

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →