How the Apache Maka Harbor/Pier Eval Harness Executes Multi-Arm Experiments

Apache Maka's Harbor/Pier framework runs multi-arm experiments by spawning isolated Docker containers per arm, applying customizable relay policies for early termination, and aggregating results through an egress-filtered pipeline.

The Eval component in Apache Maka provides a production-grade harness for reproducible A/B-style experimentation. Codenamed Harbor (the runtime orchestrator) and Pier (the client library), this framework enables engineers to define, execute, and analyze multi-arm trials with isolated workloads and policy-driven lifecycle management. This article examines the three-layer architecture and step-by-step execution flow based on the source code in apache/maka.

Three-Layer Architecture

The Harbor/Pier framework separates concerns across definition, orchestration, and result handling.

Experiment Definition Layer (Pier)

Users describe experiments through JSON or YAML manifests that declare arms, parameters, and trial policies. Pier parses these definitions in pier/experiment.py, constructing a Trial object that encapsulates one complete experiment execution. This abstraction allows experiment designers to version control their trials as declarative configuration rather than imperative code.

Trial Orchestration Layer (Harbor)

Harbor receives Trial objects from Pier and manages their execution. The core components are:

Harbor spawns trial workers—one isolated Docker container per arm—ensuring complete namespace separation between competing variants.

Result Collection & Egress Filtering

After arm completion, packages/eval/harbor/egress_filter.py validates all artifacts against a configurable whitelist before persistence. This safety mechanism prevents accidental data leakage, particularly important when experiments process sensitive production data.

Multi-Arm Experiment Execution Flow

The following six-step process governs every Harbor/Pier experiment:

  1. Define the experiment — Create a manifest specifying each arm's container image, command, and resource constraints
  2. Create a trial — Pier builds a Trial object and transmits it to Harbor via the eval_framework.run_trial RPC
  3. Spawn arm workers — Harbor iterates through arms, launching Docker containers via the relay_agent with full namespace isolation
  4. Apply trial policy — Concurrently evaluate relay policies (such as early stopping thresholds) against live metrics
  5. Collect results — Pull output files; enforce egress filtering before storage
  6. Aggregate and report — Harbor consolidates metrics into a TrialResult returned to Pier for presentation

Relay Policy Implementation

The relay policy system enables dynamic experiment management without code redeployment. Harbor continuously evaluates policies against streaming metrics, allowing arms to be:

  • Terminated early when performance degrades below thresholds
  • Promoted automatically when statistical significance is achieved
  • Throttled to control resource expenditure

Policy customization occurs in relay_agent.py. The framework ships with common policies and accepts user-defined implementations.

Code Examples

Defining a Multi-Arm Experiment Manifest


# experiment.yaml

name: search-rank-test
arms:
  - name: control
    image: myapp:v1.0
    cmd: ["python", "rank.py", "--model=baseline"]
    resources:
      memory: "4Gi"
      cpu: "2"

  - name: candidate
    image: myapp:v1.1
    cmd: ["python", "rank.py", "--model=transformer"]
    resources:
      memory: "4Gi"
      cpu: "2"

policy:
  type: early_stop
  metric: click_through_rate
  threshold: 0.02
  evaluation_interval_seconds: 60

Launching from Pier

from eval.harbor.eval_framework import run_trial

# Trial object created from validated manifest

experiment = load_experiment("experiment.yaml")

# Harbor spawns Docker workers, applies policy, filters egress

result = run_trial(experiment)

print(f"Best arm: {result.best_arm}")
print(f"Winner metrics: {result.arms[result.best_arm].metrics}")
print(f"Total duration: {result.duration_seconds}s")

Custom Relay Policy


# In packages/eval/harbor/relay_agent.py

class StatisticalSignificancePolicy:
    """
    Promote an arm when p-value against control 
    drops below alpha, with minimum sample size guardrails.
    """
    
    def __init__(self, alpha: float = 0.05, min_samples: int = 10000):
        self.alpha = alpha
        self.min_samples = min_samples
    
    def evaluate(self, arm_metrics: dict, control_metrics: dict) -> Action:
        if arm_metrics["sample_size"] < self.min_samples:
            return Action.CONTINUE
            
        p_value = compute_t_test(
            arm_metrics["conversions"], 
            control_metrics["conversions"]
        )
        
        if p_value < self.alpha and arm_metrics["conversion_rate"] > control_metrics["conversion_rate"]:
            return Action.PROMOTE
            
        if p_value < self.alpha and arm_metrics["conversion_rate"] < control_metrics["conversion_rate"]:
            return Action.TERMINATE
            
        return Action.CONTINUE

Concurrency and Resource Model

Because each arm executes in an independent Docker container, Harbor imposes no artificial concurrency limits. The practical maximum depends on:

  • Available host/container orchestrator resources
  • Egress filter throughput for artifact validation
  • Policy evaluation frequency (configurable per experiment)

Resource quotas declared in arm definitions are enforced at the container runtime level, preventing resource starvation across concurrent trials.

Key Source Files

File Responsibility
packages/eval/harbor/eval_framework.py Trial orchestration and RPC handling
packages/eval/harbor/relay_agent.py Policy engine for arm lifecycle decisions
packages/eval/harbor/egress_filter.py Artifact validation and data loss prevention
packages/eval/harbor/run_trial.py Single-arm execution and result collection
packages/eval/harbor/test_eval_framework.py Unit tests demonstrating usage patterns

Summary

  • Harbor/Pier separates experiment definition (Pier) from execution (Harbor) for reproducible multi-arm experimentation
  • Each arm runs in an isolated Docker container with configurable resource limits
  • Relay policies in relay_agent.py enable dynamic termination, promotion, and throttling without code changes
  • Egress filtering prevents data leakage by whitelist-validating all output artifacts
  • The architecture scales horizontally based on host resources rather than framework constraints

Frequently Asked Questions

How does Harbor ensure isolation between experiment arms?

Harbor launches each arm as a separate Docker container with its own network namespace, process tree, and filesystem. This container-level isolation prevents interference between arms, ensures clean teardown, and enables reproducible execution regardless of what runs in other arms.

Can I customize when an arm gets terminated early?

Yes. The relay_agent.py policy system accepts custom implementations. Define a class with an evaluate() method that receives current metrics and returns Action.CONTINUE, Action.TERMINATE, or Action.PROMOTE. Reference your policy class in the experiment manifest's policy.type field.

What prevents sensitive data from leaking through experiment artifacts?

The egress_filter.py component validates every artifact against a configured whitelist of allowed file types, paths, and content patterns. Only matching artifacts persist to the experiment store. This enforcement occurs before any Pier-level access, providing defense in depth for production data handling.

How does Pier communicate with Harbor during a running trial?

Pier initiates a single RPC call to eval_framework.run_trial, passing the serialized Trial object. Harbor then assumes control, executing arms and applying policies without further Pier interaction until completion. This fire-and-forget model simplifies client implementations and allows Harbor to optimize internally for throughput and fault tolerance.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →