# How Results Are Managed and Selected in Apache Maka’s Evaluation Framework

> Discover how Apache Maka manages and selects results in its evaluation framework. Learn about backend selection, relay agents, and structured JSON output for efficient data handling.

- Repository: [The Apache Software Foundation/maka](https://github.com/apache/maka)
- Tags: internals
- Published: 2026-09-04

---

**Apache Maka’s evaluation harness uses a global registration pattern to abstract backend selection between Harbor and Pier, storing the chosen framework in a private module-level variable while delegating result generation to backend-specific relay agents that write structured JSON files.**

Apache Maka’s evaluation framework provides a unified interface for executing trials across multiple evaluation backends. Understanding how results are managed and selected in Maka's evaluation framework reveals a design that decouples trial orchestration from backend-specific implementations through global state management and dynamic dispatch.

## Framework Registration and Validation

When a trial initiates, the system validates and registers the evaluation backend through a strict registration process. In [`packages/eval/harbor/run_trial.py`](https://github.com/apache/maka/blob/main/packages/eval/harbor/run_trial.py), the command-line argument specifying the framework is parsed and validated before calling the registration function.

The `install(framework)` function in [`packages/eval/harbor/eval_framework.py`](https://github.com/apache/maka/blob/main/packages/eval/harbor/eval_framework.py) accepts only two supported values: `"harbor"` and `"pier"`. This function stores the selection in a private module-level variable `_framework`. If an unsupported framework is specified, the system raises a `RuntimeError` immediately, preventing accidental misconfiguration.

```python

# From run_trial.py - Framework selection and registration

from packages.eval.harbor.eval_framework import install

# Command-line validation (simplified)

framework = argv[1]  # Must be "harbor" or "pier"

install(framework)   # Registers the backend globally

```

## Global Framework Selection

Once installed, any component can retrieve the active framework using the `selected()` function from the same module. This function checks that a framework has been previously installed and returns the stored string value.

According to the apache/maka source code, calling `selected()` before `install()` raises a `RuntimeError`, ensuring that backend-dependent code never executes with an uninitialized framework. The trial runner, relay agents, and test utilities all import this function to determine which backend implementation to invoke.

```python

# From relay_agent.py - Dynamic backend dispatch

from packages.eval.harbor.eval_framework import selected

def start_relay():
    backend = selected()  # Returns "harbor" or "pier"

    if backend == "harbor":
        from .harbor_impl import HarborRelay
        relay = HarborRelay()
    else:
        from .pier_impl import PierRelay
        relay = PierRelay()
    relay.run()

```

## Backend-Specific Result Generation

The actual trial results are produced by the backend-specific relay implementation determined by the global selector. The selected framework dictates which relay class executes the trial logic and writes the outcome.

As implemented in apache/maka, the Harbor or Pier relay agent writes results to a structured JSON file located in the trial’s working directory. This decouples result formatting from the evaluation harness—the framework only knows that *some* backend will produce a [`result.json`](https://github.com/apache/maka/blob/main/result.json) file, not how that file is generated.

## Result Retrieval and Validation

The test harness reads these backend-agnostic JSON outputs to aggregate metrics and surface results. In [`packages/eval/harbor/test_eval_framework.py`](https://github.com/apache/maka/blob/main/packages/eval/harbor/test_eval_framework.py), utilities parse the structured JSON to verify trial outcomes.

Because the harness calls `selected()` only once per process, the result-management code remains completely decoupled from the backend choice. Swapping Harbor for Pier requires changing only the initial configuration line, with all downstream components automatically adapting to the new result format.

```python

# From test_eval_framework.py - Reading trial results

import json
from pathlib import Path

def load_results(trial_dir: Path) -> dict:
    result_file = trial_dir / "result.json"
    with result_file.open() as f:
        return json.load(f)

# Example assertion pattern

def test_trial_success(tmp_path):
    results = load_results(tmp_path)
    assert results["status"] == "passed"
    assert results["score"] > 0.8

```

## Summary

- **Strict Registration**: The `install()` function in [`packages/eval/harbor/eval_framework.py`](https://github.com/apache/maka/blob/main/packages/eval/harbor/eval_framework.py) validates and stores the framework choice in a private `_framework` variable, rejecting unsupported backends with a `RuntimeError`.
- **Global Access Pattern**: Components retrieve the active backend via `selected()`, which ensures initialization before returning either `"harbor"` or `"pier"`.
- **Dynamic Dispatch**: Relay agents use the selected framework value to dynamically import and instantiate the correct backend implementation (`HarborRelay` or `PierRelay`).
- **Decoupled Results**: Backend-specific relays write structured JSON files that the test harness reads agnostically, allowing single-line backend swaps without modifying result-parsing logic.

## Frequently Asked Questions

### What happens if I try to use the evaluation framework without installing a backend first?

The `selected()` function raises a `RuntimeError` if called before `install()`. This guard ensures that all components verify the framework is initialized before attempting to communicate with backend-specific implementations, preventing silent failures in the evaluation pipeline.

### Can I add custom evaluation backends beyond Harbor and Pier?

The current implementation in [`packages/eval/harbor/eval_framework.py`](https://github.com/apache/maka/blob/main/packages/eval/harbor/eval_framework.py) explicitly restricts the `_framework` variable to only `"harbor"` or `"pier"` values. Adding new backends would require modifying the validation logic in the `install()` function and implementing corresponding relay classes that conform to the expected result JSON schema.

### Where are trial results physically stored in the Maka evaluation framework?

Trial results are written as structured JSON files (typically named [`result.json`](https://github.com/apache/maka/blob/main/result.json)) in the trial’s working directory by the backend-specific relay agent. The test harness utilities in [`packages/eval/harbor/test_eval_framework.py`](https://github.com/apache/maka/blob/main/packages/eval/harbor/test_eval_framework.py) read these files to surface metrics to developers or CI systems.

### How does the framework ensure results come from the correct backend?

The global selector pattern guarantees that once `install(framework)` is called, all subsequent calls to `selected()` return the same framework identifier for the lifetime of the process. The relay agent uses this value to instantiate exactly one backend implementation (Harbor or Pier), ensuring result generation and storage are handled consistently by the selected platform.