Eval Subjects vs External Subjects in Maka: Benchmark Architecture Guide

Eval Subjects execute inside Maka’s Runtime Host as fully managed benchmark workloads with complete lifecycle control, while External Subjects run outside the Runtime Host via a generic adapter to compare competitor systems against Maka’s performance metrics.

In Apache Maka’s benchmark architecture, a subject represents the fundamental unit of work that an experiment executes. Understanding the distinction between Eval Subjects and External Subjects in Maka is critical for designing reproducible benchmark experiments and integrating third-party competitors. While both types appear in experiment configurations, they differ fundamentally in execution authority, runtime boundaries, and integration depth.

What Defines a Subject in Maka?

Before comparing the two types, it is important to understand that a subject is the atomic unit of execution in Maka’s evaluation framework. According to the architecture documentation, subjects determine how workloads are isolated, executed, and measured during benchmark trials, with specific semantics governing their runtime behavior.

Eval Subjects: Native Runtime Host Execution

Authority and Execution Model

Eval Subjects are executed inside the Runtime Host as regular Maka subjects. The Runtime Host maintains complete authority over the session, turn, tool runtime, and all continuation semantics. As documented in ARCHITECTURE.md (lines 37-41), a "Maka subject always crosses the public Runtime Host client/protocol boundary," ensuring the host controls lifecycle, permissions, tools, and event logging.

Implementation and Validation

These subjects represent the benchmark workload that the experiment measures, such as a model answering questions or generating code. All benchmark semantics—including tasks, repetitions, cells, attempts, result selection, budgets, and verifier configuration—are defined for Eval subjects.

In the @maka/eval package, Eval subjects run via the Agent.run() method, which flows through the Runtime Host to the task ledger. Before a trial starts, the Eval framework validates the subject’s tool chain, environment variables, and credential bindings.

To run an Eval Subject, use the standard evaluation command:

maka eval run --subject my-benchmark-task --config eval.config.js

External Subjects: The Competitor Adapter

External Execution Boundary

External Subjects are handled by a generic external-subject adapter that sits outside the Runtime Host. Unlike Eval subjects, these run in their own environment and communicate through a thin wrapper. As noted in packages/eval/README.md (lines 44-45), "External subjects declare a command and arguments" rather than using the Runtime Host client, meaning they do not participate in the Runtime Host protocol and are treated as competitors.

Configuration and Result Contracts

External Subjects are declared with a command, arguments, optional non-secret environment values, and a result contract specifying how results are returned. Maka supports two contract types:

  • exit-code: Simple process exit code interpretation
  • protocol-v1: Structured result frame protocol

The wrapper strips any provider-native web tools and enforces the result frame, ensuring isolation from Maka’s internal execution graph.

Example configuration for an External Subject:

external_subject:
  command: "python"
  args: ["competitor_script.py", "--model", "gpt-4"]
  result_contract: "exit-code"
  env:
    COMPETITOR_API_KEY: "public-value"

Key Architectural Differences

The distinction between these subject types manifests across several architectural dimensions:

Aspect Eval Subject External Subject
Execution Location Inside Runtime Host with full protocol support Outside via generic adapter with thin wrapper
Runtime Authority Runtime Host controls lifecycle, permissions, and tools No Runtime Host authority; runs independently
Use Case Reproducible benchmark workloads (maka eval run) Comparative evaluation against third-party LLMs or tools
Validation Tool chain, credentials, and environment validated by Eval framework Result contract validation only (exit-code or protocol-v1)
Source Reference ARCHITECTURE.md (lines 37-41) packages/eval/README.md (lines 44-45)

Summary

  • Eval Subjects are native Maka workloads executed under the full control of the Runtime Host, supporting complex benchmark semantics like retries, budgeting, and verification.
  • External Subjects bypass the Runtime Host to run competitor systems via a generic adapter, using command-based execution and simple result contracts.
  • The architectural boundary is strictly defined: Eval Subjects cross the public Runtime Host client/protocol boundary, while External Subjects remain outside as competitors.
  • Configuration differs fundamentally: Eval Subjects use Agent.run() and the Eval framework, whereas External Subjects declare commands and arguments with result contracts.

Frequently Asked Questions

Can External Subjects access Maka’s tool runtime?

No. External Subjects run outside the Runtime Host and cannot access Maka’s internal tool runtime or continuation semantics. They are limited to their own environment and communicate results through a thin wrapper using either the exit-code or protocol-v1 result contract, as specified in packages/eval/README.md.

How do I configure an External Subject for a competitor LLM?

Declare the External Subject in your experiment configuration with a command, arguments, and result contract. The Eval framework executes it via the generic external-subject adapter, which strips any provider-native web tools and enforces the result frame without Runtime Host integration or credential validation.

Why would I use an Eval Subject instead of an External Subject?

Use an Eval Subject when you need the Runtime Host to manage execution lifecycle, validate credentials, enforce budgets, or handle complex continuation patterns. Eval Subjects are designed for reproducible benchmark workloads where Maka controls the entire execution graph via Agent.run() and the task ledger.

Where are these subject types defined in the Apache Maka source code?

Subject execution semantics are defined in packages/eval/README.md (lines 44-45), which clarifies that "Maka subjects ask the Runtime Host client to run one owned execution… External subjects declare a command and arguments." The architectural boundary is documented in ARCHITECTURE.md (lines 37-41), which notes that Maka subjects always cross the public Runtime Host client/protocol boundary.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →