# How the Apache Maka Harbor/Pier Eval Harness Executes Multi-Arm Experiments

> Discover how Apache Maka's Harbor Pier framework executes multi-arm experiments. Learn about Docker containers, relay policies for early termination, and result aggregation.

- Repository: [The Apache Software Foundation/maka](https://github.com/apache/maka)
- Tags: how-to-guide
- Published: 2026-09-02

---

**Apache Maka's Harbor/Pier framework runs multi-arm experiments by spawning isolated Docker containers per arm, applying customizable relay policies for early termination, and aggregating results through an egress-filtered pipeline.**

The **Eval** component in Apache Maka provides a production-grade harness for reproducible A/B-style experimentation. Codenamed **Harbor** (the runtime orchestrator) and **Pier** (the client library), this framework enables engineers to define, execute, and analyze multi-arm trials with isolated workloads and policy-driven lifecycle management. This article examines the three-layer architecture and step-by-step execution flow based on the source code in `apache/maka`.

## Three-Layer Architecture

The Harbor/Pier framework separates concerns across definition, orchestration, and result handling.

### Experiment Definition Layer (Pier)

Users describe experiments through JSON or YAML manifests that declare **arms**, **parameters**, and **trial policies**. Pier parses these definitions in [`pier/experiment.py`](https://github.com/apache/maka/blob/main/pier/experiment.py), constructing a `Trial` object that encapsulates one complete experiment execution. This abstraction allows experiment designers to version control their trials as declarative configuration rather than imperative code.

### Trial Orchestration Layer (Harbor)

Harbor receives `Trial` objects from Pier and manages their execution. The core components are:

- **[`packages/eval/harbor/eval_framework.py`](https://github.com/apache/maka/blob/main/packages/eval/harbor/eval_framework.py)** — Central entry point that creates and manages `Trial` objects, coordinating the entire lifecycle
- **[`packages/eval/harbor/relay_agent.py`](https://github.com/apache/maka/blob/main/packages/eval/harbor/relay_agent.py)** — Implements **relay policies** that determine when arms terminate or receive additional traffic
- **[`packages/eval/harbor/run_trial.py`](https://github.com/apache/maka/blob/main/packages/eval/harbor/run_trial.py)** — Executes individual trials and returns raw results

Harbor spawns **trial workers**—one isolated Docker container per arm—ensuring complete namespace separation between competing variants.

### Result Collection & Egress Filtering

After arm completion, **[`packages/eval/harbor/egress_filter.py`](https://github.com/apache/maka/blob/main/packages/eval/harbor/egress_filter.py)** validates all artifacts against a configurable whitelist before persistence. This safety mechanism prevents accidental data leakage, particularly important when experiments process sensitive production data.

## Multi-Arm Experiment Execution Flow

The following six-step process governs every Harbor/Pier experiment:

1. **Define the experiment** — Create a manifest specifying each arm's container image, command, and resource constraints
2. **Create a trial** — Pier builds a `Trial` object and transmits it to Harbor via the `eval_framework.run_trial` RPC
3. **Spawn arm workers** — Harbor iterates through arms, launching Docker containers via the `relay_agent` with full namespace isolation
4. **Apply trial policy** — Concurrently evaluate relay policies (such as early stopping thresholds) against live metrics
5. **Collect results** — Pull output files; enforce egress filtering before storage
6. **Aggregate and report** — Harbor consolidates metrics into a `TrialResult` returned to Pier for presentation

## Relay Policy Implementation

The **relay policy** system enables dynamic experiment management without code redeployment. Harbor continuously evaluates policies against streaming metrics, allowing arms to be:

- **Terminated early** when performance degrades below thresholds
- **Promoted** automatically when statistical significance is achieved
- **Throttled** to control resource expenditure

Policy customization occurs in [`relay_agent.py`](https://github.com/apache/maka/blob/main/relay_agent.py). The framework ships with common policies and accepts user-defined implementations.

## Code Examples

### Defining a Multi-Arm Experiment Manifest

```yaml

# experiment.yaml

name: search-rank-test
arms:
  - name: control
    image: myapp:v1.0
    cmd: ["python", "rank.py", "--model=baseline"]
    resources:
      memory: "4Gi"
      cpu: "2"

  - name: candidate
    image: myapp:v1.1
    cmd: ["python", "rank.py", "--model=transformer"]
    resources:
      memory: "4Gi"
      cpu: "2"

policy:
  type: early_stop
  metric: click_through_rate
  threshold: 0.02
  evaluation_interval_seconds: 60

```

### Launching from Pier

```python
from eval.harbor.eval_framework import run_trial

# Trial object created from validated manifest

experiment = load_experiment("experiment.yaml")

# Harbor spawns Docker workers, applies policy, filters egress

result = run_trial(experiment)

print(f"Best arm: {result.best_arm}")
print(f"Winner metrics: {result.arms[result.best_arm].metrics}")
print(f"Total duration: {result.duration_seconds}s")

```

### Custom Relay Policy

```python

# In packages/eval/harbor/relay_agent.py

class StatisticalSignificancePolicy:
    """
    Promote an arm when p-value against control 
    drops below alpha, with minimum sample size guardrails.
    """
    
    def __init__(self, alpha: float = 0.05, min_samples: int = 10000):
        self.alpha = alpha
        self.min_samples = min_samples
    
    def evaluate(self, arm_metrics: dict, control_metrics: dict) -> Action:
        if arm_metrics["sample_size"] < self.min_samples:
            return Action.CONTINUE
            
        p_value = compute_t_test(
            arm_metrics["conversions"], 
            control_metrics["conversions"]
        )
        
        if p_value < self.alpha and arm_metrics["conversion_rate"] > control_metrics["conversion_rate"]:
            return Action.PROMOTE
            
        if p_value < self.alpha and arm_metrics["conversion_rate"] < control_metrics["conversion_rate"]:
            return Action.TERMINATE
            
        return Action.CONTINUE

```

## Concurrency and Resource Model

Because each arm executes in an independent Docker container, Harbor imposes **no artificial concurrency limits**. The practical maximum depends on:

- Available host/container orchestrator resources
- Egress filter throughput for artifact validation
- Policy evaluation frequency (configurable per experiment)

Resource quotas declared in arm definitions are enforced at the container runtime level, preventing resource starvation across concurrent trials.

## Key Source Files

| File | Responsibility |
|------|--------------|
| [`packages/eval/harbor/eval_framework.py`](https://github.com/apache/maka/blob/main/packages/eval/harbor/eval_framework.py) | Trial orchestration and RPC handling |
| [`packages/eval/harbor/relay_agent.py`](https://github.com/apache/maka/blob/main/packages/eval/harbor/relay_agent.py) | Policy engine for arm lifecycle decisions |
| [`packages/eval/harbor/egress_filter.py`](https://github.com/apache/maka/blob/main/packages/eval/harbor/egress_filter.py) | Artifact validation and data loss prevention |
| [`packages/eval/harbor/run_trial.py`](https://github.com/apache/maka/blob/main/packages/eval/harbor/run_trial.py) | Single-arm execution and result collection |
| [`packages/eval/harbor/test_eval_framework.py`](https://github.com/apache/maka/blob/main/packages/eval/harbor/test_eval_framework.py) | Unit tests demonstrating usage patterns |

## Summary

- **Harbor/Pier** separates experiment definition (Pier) from execution (Harbor) for reproducible multi-arm experimentation
- Each arm runs in an **isolated Docker container** with configurable resource limits
- **Relay policies** in [`relay_agent.py`](https://github.com/apache/maka/blob/main/relay_agent.py) enable dynamic termination, promotion, and throttling without code changes
- **Egress filtering** prevents data leakage by whitelist-validating all output artifacts
- The architecture scales horizontally based on host resources rather than framework constraints

## Frequently Asked Questions

### How does Harbor ensure isolation between experiment arms?

Harbor launches each arm as a separate Docker container with its own network namespace, process tree, and filesystem. This container-level isolation prevents interference between arms, ensures clean teardown, and enables reproducible execution regardless of what runs in other arms.

### Can I customize when an arm gets terminated early?

Yes. The [`relay_agent.py`](https://github.com/apache/maka/blob/main/relay_agent.py) policy system accepts custom implementations. Define a class with an `evaluate()` method that receives current metrics and returns `Action.CONTINUE`, `Action.TERMINATE`, or `Action.PROMOTE`. Reference your policy class in the experiment manifest's `policy.type` field.

### What prevents sensitive data from leaking through experiment artifacts?

The [`egress_filter.py`](https://github.com/apache/maka/blob/main/egress_filter.py) component validates every artifact against a configured whitelist of allowed file types, paths, and content patterns. Only matching artifacts persist to the experiment store. This enforcement occurs before any Pier-level access, providing defense in depth for production data handling.

### How does Pier communicate with Harbor during a running trial?

Pier initiates a single RPC call to `eval_framework.run_trial`, passing the serialized `Trial` object. Harbor then assumes control, executing arms and applying policies without further Pier interaction until completion. This fire-and-forget model simplifies client implementations and allows Harbor to optimize internally for throughput and fault tolerance.