Maka Eval Framework: Understanding Experiment, Cell, and Repetition

Maka's Eval framework uses three core abstractions—Experiment, Cell, and Repetition—to create isolated, repeatable test executions with automatic retry policies and network filtering.

The Maka Eval framework, located in packages/eval/harbor, provides a lightweight runtime for reproducible testing. Built for Apache Maka, it separates test orchestration from execution isolation, enabling deterministic retries and statistical validation across multiple trials. Each abstraction serves a distinct purpose: Experiment defines the scenario, Cell contains the execution, and Repetition governs retry behavior through configurable policies.

What Is an Experiment in Maka Eval?

An Experiment is the top-level container that describes an entire test scenario. It groups related cells, defines global policies, and orchestrates the complete lifecycle from launch to completion.

The implementation lives in eval_framework.py as the EvalFramework class:


# From packages/eval/harbor/eval_framework.py

class EvalFramework:
    """Parses experiment specs and coordinates cell execution."""

Key responsibilities include:

  • Parsing YAML/JSON specifications
  • Instantiating CellEnvironment objects for each defined cell
  • Applying a RunTrialPolicy to govern failure handling
  • Coordinating start/stop operations across the entire run

Experiments are typically defined declaratively:


# experiment.yaml

name: demo-experiment
policy:
  maxAttempts: 3
  retryOnFailure: true
cells:
  - id: shell-cell
    command: "node -e \"console.log('hello')\""
    env:
      NODE_ENV: test

Or created programmatically in TypeScript:

import { createEvalFramework } from '@maka/eval';

const spec = {
  name: 'demo-experiment',
  cells: [{ id: 'shell-cell', command: 'node -e "console.log(\'hello\')"' }],
  policy: { maxAttempts: 3, retryOnFailure: true }
};

const framework = await createEvalFramework(spec);
const result = await framework.run();

What Is a Cell in Maka Eval?

A Cell is the smallest executable unit within an experiment. It encapsulates a single isolated workload—whether a shell command, Python script, or Docker container—and guarantees that side effects do not leak between executions.

The cell implementation resides in cell_egress_namespace.py:


# From packages/eval/harbor/test_cell_egress_namespace.py

class CellEnvironment:
    """Sets up sandboxed execution environment for a cell."""

Each cell receives:

  • Network namespace isolation — separate network stack prevents cross-cell interference
  • Temporary filesystem — ephemeral storage that is cleaned between trials
  • Managed I/O streams — stdout, stderr, and result payloads captured and reported

The CellEnvironment constructs these sandboxes dynamically, ensuring that repeating a cell produces identical starting conditions regardless of previous trial outcomes.

What Is Repetition in Maka Eval?

Repetition (also called a Trial) represents a single execution attempt of a cell. The framework automatically repeats cells based on configurable policies to handle flaky operations or gather statistically significant results.

Repetition logic is implemented in run_trial.py:


# From packages/eval/harbor/run_trial.py

class RunTrialPolicy:
    """Decides whether to retry, accept, or abort based on trial results."""
    
class FrameworkVersionMismatch(Exception):
    """Raised when cell framework version conflicts with runner."""

The RunTrialPolicy tracks:

  • Attempt counters and timing per cell
  • Success/failure criteria from result payloads
  • Back-off strategies and maximum attempt limits

After each trial, the policy examines the cell's exit code, stdout, and structured JSON payload to determine whether to accept the result, retry with the same configuration, or abort the entire experiment.

Supporting Components: Relay Agent and Egress Filter

Two additional components enable the framework's network isolation and observability capabilities.

Relay Agent

The RelayAgent in relay_agent.py acts as a lightweight proxy between the cell and host:


# From packages/eval/harbor/relay_agent.py

class RelayAgent:
    """Forwards traffic between cell and host, injecting filters."""

It opens bidirectional streams, applies EgressFilter policies transparently, and handles graceful cell shutdown without modifying the cell's code.

Egress Filter

Network-level filtering is implemented in egress_filter.py:


# From packages/eval/harbor/egress_filter.py

class CloseRawLayer:
    """Intercepts raw TCP before Mitmproxy classification."""

Filters can block, rewrite, or log outbound requests. When traffic cannot be parsed, the filter emits classification errors for debugging.

How Experiment, Cell, and Repetition Work Together

The execution flow follows four distinct phases:

  1. Experiment creation — EvalFramework validates the spec and builds CellEnvironment instances
  2. Cell launch — Each cell spawns in a sandbox with a dedicated RelayAgent
  3. Trial loop — RunTrialPolicy governs repetition based on returned payloads
  4. Telemetry collection — RelayAgent forwards traffic through EgressFilter for metrics and policy enforcement

This architecture ensures that filesystem mutations, network state, and environment variables remain isolated per trial. Flaky tests become deterministic, and statistical data collection scales reliably across hundreds of repetitions.

Key Source Files in Maka Eval

File Purpose
eval_framework.py Core orchestration logic for experiments
run_trial.py RunTrialPolicy and trial-control exceptions
relay_agent.py Proxy connecting cells to host with filter injection
egress_filter.py Network egress filtering implementations
test_cell_egress_namespace.py Unit tests for CellEnvironment construction
test_eval_framework.py End-to-end experiment lifecycle tests

Running Maka Eval Experiments

CLI execution uses the published npm package:

npx @maka/eval run experiment.yaml

The CLI parses the YAML spec, instantiates the framework, and executes the full trial loop automatically.

Summary

  • Experiment — Defines test scenarios, groups cells, and applies global policies via EvalFramework
  • Cell — Provides isolated execution environments through CellEnvironment with sandboxed network and filesystem
  • Repetition — Enables deterministic retries and statistical validation through RunTrialPolicy
  • Relay Agent and Egress Filter — Enable transparent network interception and policy enforcement without cell modification

Frequently Asked Questions

How does Maka Eval handle flaky test failures?

The RunTrialPolicy in run_trial.py automatically retries cells based on configurable criteria including max attempts, retry-on-failure flags, and back-off strategies. Each retry starts with a fresh CellEnvironment, eliminating state corruption from previous attempts.

Can cells in Maka Eval run Docker containers?

Yes. Cells support multiple execution types including shell commands, Python scripts, and container images. The CellEnvironment abstracts the runtime, applying the same isolation guarantees regardless of execution method.

What file formats does Maka Eval accept for experiment specifications?

Maka Eval accepts both YAML and JSON formats. The EvalFramework class in eval_framework.py parses these specifications to instantiate cells and configure trial policies programmatically.

How does network isolation work between repeated cell trials?

Each trial receives a dedicated network namespace created by CellEnvironment. The RelayAgent forwards all traffic through configurable EgressFilter instances, ensuring that network state from one trial cannot affect subsequent repetitions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →