Maka Eval Framework: Understanding Experiment, Cell, and Repetition
Maka's Eval framework uses three core abstractions—Experiment, Cell, and Repetition—to create isolated, repeatable test executions with automatic retry policies and network filtering.
The Maka Eval framework, located in packages/eval/harbor, provides a lightweight runtime for reproducible testing. Built for Apache Maka, it separates test orchestration from execution isolation, enabling deterministic retries and statistical validation across multiple trials. Each abstraction serves a distinct purpose: Experiment defines the scenario, Cell contains the execution, and Repetition governs retry behavior through configurable policies.
What Is an Experiment in Maka Eval?
An Experiment is the top-level container that describes an entire test scenario. It groups related cells, defines global policies, and orchestrates the complete lifecycle from launch to completion.
The implementation lives in eval_framework.py as the EvalFramework class:
# From packages/eval/harbor/eval_framework.py
class EvalFramework:
"""Parses experiment specs and coordinates cell execution."""
Key responsibilities include:
- Parsing YAML/JSON specifications
- Instantiating
CellEnvironmentobjects for each defined cell - Applying a
RunTrialPolicyto govern failure handling - Coordinating start/stop operations across the entire run
Experiments are typically defined declaratively:
# experiment.yaml
name: demo-experiment
policy:
maxAttempts: 3
retryOnFailure: true
cells:
- id: shell-cell
command: "node -e \"console.log('hello')\""
env:
NODE_ENV: test
Or created programmatically in TypeScript:
import { createEvalFramework } from '@maka/eval';
const spec = {
name: 'demo-experiment',
cells: [{ id: 'shell-cell', command: 'node -e "console.log(\'hello\')"' }],
policy: { maxAttempts: 3, retryOnFailure: true }
};
const framework = await createEvalFramework(spec);
const result = await framework.run();
What Is a Cell in Maka Eval?
A Cell is the smallest executable unit within an experiment. It encapsulates a single isolated workload—whether a shell command, Python script, or Docker container—and guarantees that side effects do not leak between executions.
The cell implementation resides in cell_egress_namespace.py:
# From packages/eval/harbor/test_cell_egress_namespace.py
class CellEnvironment:
"""Sets up sandboxed execution environment for a cell."""
Each cell receives:
- Network namespace isolation — separate network stack prevents cross-cell interference
- Temporary filesystem — ephemeral storage that is cleaned between trials
- Managed I/O streams — stdout, stderr, and result payloads captured and reported
The CellEnvironment constructs these sandboxes dynamically, ensuring that repeating a cell produces identical starting conditions regardless of previous trial outcomes.
What Is Repetition in Maka Eval?
Repetition (also called a Trial) represents a single execution attempt of a cell. The framework automatically repeats cells based on configurable policies to handle flaky operations or gather statistically significant results.
Repetition logic is implemented in run_trial.py:
# From packages/eval/harbor/run_trial.py
class RunTrialPolicy:
"""Decides whether to retry, accept, or abort based on trial results."""
class FrameworkVersionMismatch(Exception):
"""Raised when cell framework version conflicts with runner."""
The RunTrialPolicy tracks:
- Attempt counters and timing per cell
- Success/failure criteria from result payloads
- Back-off strategies and maximum attempt limits
After each trial, the policy examines the cell's exit code, stdout, and structured JSON payload to determine whether to accept the result, retry with the same configuration, or abort the entire experiment.
Supporting Components: Relay Agent and Egress Filter
Two additional components enable the framework's network isolation and observability capabilities.
Relay Agent
The RelayAgent in relay_agent.py acts as a lightweight proxy between the cell and host:
# From packages/eval/harbor/relay_agent.py
class RelayAgent:
"""Forwards traffic between cell and host, injecting filters."""
It opens bidirectional streams, applies EgressFilter policies transparently, and handles graceful cell shutdown without modifying the cell's code.
Egress Filter
Network-level filtering is implemented in egress_filter.py:
# From packages/eval/harbor/egress_filter.py
class CloseRawLayer:
"""Intercepts raw TCP before Mitmproxy classification."""
Filters can block, rewrite, or log outbound requests. When traffic cannot be parsed, the filter emits classification errors for debugging.
How Experiment, Cell, and Repetition Work Together
The execution flow follows four distinct phases:
- Experiment creation —
EvalFrameworkvalidates the spec and buildsCellEnvironmentinstances - Cell launch — Each cell spawns in a sandbox with a dedicated
RelayAgent - Trial loop —
RunTrialPolicygoverns repetition based on returned payloads - Telemetry collection —
RelayAgentforwards traffic throughEgressFilterfor metrics and policy enforcement
This architecture ensures that filesystem mutations, network state, and environment variables remain isolated per trial. Flaky tests become deterministic, and statistical data collection scales reliably across hundreds of repetitions.
Key Source Files in Maka Eval
| File | Purpose |
|---|---|
eval_framework.py |
Core orchestration logic for experiments |
run_trial.py |
RunTrialPolicy and trial-control exceptions |
relay_agent.py |
Proxy connecting cells to host with filter injection |
egress_filter.py |
Network egress filtering implementations |
test_cell_egress_namespace.py |
Unit tests for CellEnvironment construction |
test_eval_framework.py |
End-to-end experiment lifecycle tests |
Running Maka Eval Experiments
CLI execution uses the published npm package:
npx @maka/eval run experiment.yaml
The CLI parses the YAML spec, instantiates the framework, and executes the full trial loop automatically.
Summary
- Experiment — Defines test scenarios, groups cells, and applies global policies via
EvalFramework - Cell — Provides isolated execution environments through
CellEnvironmentwith sandboxed network and filesystem - Repetition — Enables deterministic retries and statistical validation through
RunTrialPolicy - Relay Agent and Egress Filter — Enable transparent network interception and policy enforcement without cell modification
Frequently Asked Questions
How does Maka Eval handle flaky test failures?
The RunTrialPolicy in run_trial.py automatically retries cells based on configurable criteria including max attempts, retry-on-failure flags, and back-off strategies. Each retry starts with a fresh CellEnvironment, eliminating state corruption from previous attempts.
Can cells in Maka Eval run Docker containers?
Yes. Cells support multiple execution types including shell commands, Python scripts, and container images. The CellEnvironment abstracts the runtime, applying the same isolation guarantees regardless of execution method.
What file formats does Maka Eval accept for experiment specifications?
Maka Eval accepts both YAML and JSON formats. The EvalFramework class in eval_framework.py parses these specifications to instantiate cells and configure trial policies programmatically.
How does network isolation work between repeated cell trials?
Each trial receives a dedicated network namespace created by CellEnvironment. The RelayAgent forwards all traffic through configurable EgressFilter instances, ensuring that network state from one trial cannot affect subsequent repetitions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →