How the Apache Maka Harbor/Pier Eval Harness Executes Multi-Arm Experiments
Apache Maka's Harbor/Pier framework runs multi-arm experiments by spawning isolated Docker containers per arm, applying customizable relay policies for early termination, and aggregating results through an egress-filtered pipeline.
The Eval component in Apache Maka provides a production-grade harness for reproducible A/B-style experimentation. Codenamed Harbor (the runtime orchestrator) and Pier (the client library), this framework enables engineers to define, execute, and analyze multi-arm trials with isolated workloads and policy-driven lifecycle management. This article examines the three-layer architecture and step-by-step execution flow based on the source code in apache/maka.
Three-Layer Architecture
The Harbor/Pier framework separates concerns across definition, orchestration, and result handling.
Experiment Definition Layer (Pier)
Users describe experiments through JSON or YAML manifests that declare arms, parameters, and trial policies. Pier parses these definitions in pier/experiment.py, constructing a Trial object that encapsulates one complete experiment execution. This abstraction allows experiment designers to version control their trials as declarative configuration rather than imperative code.
Trial Orchestration Layer (Harbor)
Harbor receives Trial objects from Pier and manages their execution. The core components are:
packages/eval/harbor/eval_framework.py— Central entry point that creates and managesTrialobjects, coordinating the entire lifecyclepackages/eval/harbor/relay_agent.py— Implements relay policies that determine when arms terminate or receive additional trafficpackages/eval/harbor/run_trial.py— Executes individual trials and returns raw results
Harbor spawns trial workers—one isolated Docker container per arm—ensuring complete namespace separation between competing variants.
Result Collection & Egress Filtering
After arm completion, packages/eval/harbor/egress_filter.py validates all artifacts against a configurable whitelist before persistence. This safety mechanism prevents accidental data leakage, particularly important when experiments process sensitive production data.
Multi-Arm Experiment Execution Flow
The following six-step process governs every Harbor/Pier experiment:
- Define the experiment — Create a manifest specifying each arm's container image, command, and resource constraints
- Create a trial — Pier builds a
Trialobject and transmits it to Harbor via theeval_framework.run_trialRPC - Spawn arm workers — Harbor iterates through arms, launching Docker containers via the
relay_agentwith full namespace isolation - Apply trial policy — Concurrently evaluate relay policies (such as early stopping thresholds) against live metrics
- Collect results — Pull output files; enforce egress filtering before storage
- Aggregate and report — Harbor consolidates metrics into a
TrialResultreturned to Pier for presentation
Relay Policy Implementation
The relay policy system enables dynamic experiment management without code redeployment. Harbor continuously evaluates policies against streaming metrics, allowing arms to be:
- Terminated early when performance degrades below thresholds
- Promoted automatically when statistical significance is achieved
- Throttled to control resource expenditure
Policy customization occurs in relay_agent.py. The framework ships with common policies and accepts user-defined implementations.
Code Examples
Defining a Multi-Arm Experiment Manifest
# experiment.yaml
name: search-rank-test
arms:
- name: control
image: myapp:v1.0
cmd: ["python", "rank.py", "--model=baseline"]
resources:
memory: "4Gi"
cpu: "2"
- name: candidate
image: myapp:v1.1
cmd: ["python", "rank.py", "--model=transformer"]
resources:
memory: "4Gi"
cpu: "2"
policy:
type: early_stop
metric: click_through_rate
threshold: 0.02
evaluation_interval_seconds: 60
Launching from Pier
from eval.harbor.eval_framework import run_trial
# Trial object created from validated manifest
experiment = load_experiment("experiment.yaml")
# Harbor spawns Docker workers, applies policy, filters egress
result = run_trial(experiment)
print(f"Best arm: {result.best_arm}")
print(f"Winner metrics: {result.arms[result.best_arm].metrics}")
print(f"Total duration: {result.duration_seconds}s")
Custom Relay Policy
# In packages/eval/harbor/relay_agent.py
class StatisticalSignificancePolicy:
"""
Promote an arm when p-value against control
drops below alpha, with minimum sample size guardrails.
"""
def __init__(self, alpha: float = 0.05, min_samples: int = 10000):
self.alpha = alpha
self.min_samples = min_samples
def evaluate(self, arm_metrics: dict, control_metrics: dict) -> Action:
if arm_metrics["sample_size"] < self.min_samples:
return Action.CONTINUE
p_value = compute_t_test(
arm_metrics["conversions"],
control_metrics["conversions"]
)
if p_value < self.alpha and arm_metrics["conversion_rate"] > control_metrics["conversion_rate"]:
return Action.PROMOTE
if p_value < self.alpha and arm_metrics["conversion_rate"] < control_metrics["conversion_rate"]:
return Action.TERMINATE
return Action.CONTINUE
Concurrency and Resource Model
Because each arm executes in an independent Docker container, Harbor imposes no artificial concurrency limits. The practical maximum depends on:
- Available host/container orchestrator resources
- Egress filter throughput for artifact validation
- Policy evaluation frequency (configurable per experiment)
Resource quotas declared in arm definitions are enforced at the container runtime level, preventing resource starvation across concurrent trials.
Key Source Files
| File | Responsibility |
|---|---|
packages/eval/harbor/eval_framework.py |
Trial orchestration and RPC handling |
packages/eval/harbor/relay_agent.py |
Policy engine for arm lifecycle decisions |
packages/eval/harbor/egress_filter.py |
Artifact validation and data loss prevention |
packages/eval/harbor/run_trial.py |
Single-arm execution and result collection |
packages/eval/harbor/test_eval_framework.py |
Unit tests demonstrating usage patterns |
Summary
- Harbor/Pier separates experiment definition (Pier) from execution (Harbor) for reproducible multi-arm experimentation
- Each arm runs in an isolated Docker container with configurable resource limits
- Relay policies in
relay_agent.pyenable dynamic termination, promotion, and throttling without code changes - Egress filtering prevents data leakage by whitelist-validating all output artifacts
- The architecture scales horizontally based on host resources rather than framework constraints
Frequently Asked Questions
How does Harbor ensure isolation between experiment arms?
Harbor launches each arm as a separate Docker container with its own network namespace, process tree, and filesystem. This container-level isolation prevents interference between arms, ensures clean teardown, and enables reproducible execution regardless of what runs in other arms.
Can I customize when an arm gets terminated early?
Yes. The relay_agent.py policy system accepts custom implementations. Define a class with an evaluate() method that receives current metrics and returns Action.CONTINUE, Action.TERMINATE, or Action.PROMOTE. Reference your policy class in the experiment manifest's policy.type field.
What prevents sensitive data from leaking through experiment artifacts?
The egress_filter.py component validates every artifact against a configured whitelist of allowed file types, paths, and content patterns. Only matching artifacts persist to the experiment store. This enforcement occurs before any Pier-level access, providing defense in depth for production data handling.
How does Pier communicate with Harbor during a running trial?
Pier initiates a single RPC call to eval_framework.run_trial, passing the serialized Trial object. Harbor then assumes control, executing arms and applying policies without further Pier interaction until completion. This fire-and-forget model simplifies client implementations and allows Harbor to optimize internally for throughput and fault tolerance.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →