Key Features of Maka's Evaluation Framework: LLM-Driven Goal Assessment for Autonomous Coding

Maka's evaluation framework is a lightweight, LLM-driven "judge" that runs after each turn of an autonomous coding session to classify goal status as met, impossible, making progress, or waiting on an external event.

The Apache Maka project implements a deterministic evaluation loop that prevents autonomous coding agents from陷入 infinite stalls or false positives. By leveraging a structured JSON contract and tool-free model calls, the framework provides time-bounded assessments that integrate directly with the runtime's continuation logic.

Core Architecture Components

Goal-Evaluation Interface and JSON Contract

At the heart of the system lies a strict JSON schema defined in packages/runtime/src/goal-evaluator.ts (lines 37-55). The GoalEvaluation interface requires the LLM to return one of four terminal states:

  • met: The goal condition is fully satisfied
  • impossible: The goal cannot be achieved with available resources
  • progress: Forward momentum detected but goal not yet reached
  • waiting: Paused for external events (user input, API responses)

Each evaluation must include a reason string explaining the judgment. This contract ensures that downstream components in GoalContinuation can route decisions without parsing ambiguous natural language.

Deterministic System Prompt Design

The framework uses a constant EVALUATOR_SYSTEM prompt (lines 108-120) that constrains the LLM to act strictly as a judge. Unlike conversational prompts, this system prompt:

  • Explicitly forbids tool usage
  • Mandates JSON-only output
  • Prevents the model from offering solutions or code fixes
  • Eliminates self-justification bias by separating the evaluator from the code generator

This separation is critical; the evaluator assesses the conversation history without access to the agent's internal reasoning, providing objective third-party judgment.

Tool-Free Model Invocation

To prevent function-call overhead and ensure text-only responses, the evaluator invokes generateToolFreeModelCall through the generateGoalEvaluationModelCall function (lines 89-100). This method:

  1. Bypasses tool registration entirely
  2. Enforces the JSON output schema through prompt engineering rather than structured outputs
  3. Reduces token consumption by eliminating tool description overhead

Implementation Details and Robustness Mechanisms

Prompt Construction Pipeline

The buildGoalEvaluationPrompt function (lines 22-33) concatenates three critical elements:

  1. The EVALUATOR_SYSTEM instructions
  2. The specific goal condition being tested
  3. Recent conversation context (last N turns of the coding session)
import { buildGoalEvaluationPrompt } from './goal-evaluator.js';

const condition = 'The function `add(a, b)` must return the sum of its arguments.';
const recentContext = `User: Write a function that adds two numbers.
Agent: function add(a,b){return a+b;}
`;

const prompt = buildGoalEvaluationPrompt(condition, recentContext);
// Contains system prompt, condition, context, and "--- YOUR JUDGMENT (JSON only) ---" marker

Robust Parsing and Fallback Strategies

The parseGoalEvaluation function (lines 36-66) implements tolerant JSON extraction to handle malformed LLM outputs. When the model generates extra markdown fences or explanatory text, the parser extracts the first valid JSON object using regex patterns.

If parsing completely fails, the framework returns a neutral "failed" evaluation with evaluatorFailed: true. This prevents the autonomous loop from crashing or making irreversible decisions based on garbled output.

Hard Timeout and Cancellation Protection

Production deployments require strict latency bounds. The evaluateGoal function (lines 70-136) wraps the LLM call with:

  • Configurable timeout: Default 30-second ceiling via Promise.race
  • AbortController integration: Respects external cancellation signals
  • Cleanup guarantees: Timer cleanup occurs regardless of success or failure

When timeouts trigger, the cancelledEvaluation helper propagates a special flag indicating the assessment aborted rather than returned a substantive judgment.

Reason Text Truncation

To prevent context window pollution, the EVALUATOR_REASON_TEXT_LIMIT constant (lines 103-106) caps the reason field at 200 code units (approximately 600 bytes). The truncateGoalText utility applies this limit before the evaluation result propagates to GoalContinuation, ensuring concise steering signals that fit within token budgets.

Integration with Goal Continuation

The evaluation framework feeds directly into Maka's lane management system. In packages/runtime/src/goal-continuation.ts (lines 440-447), the GoalEvaluation object attaches to lane.intent, driving:

  • Continuation routing: Determines whether to proceed, retry, or terminate the session
  • Stall detection: Distinguishes between genuine waiting states and infinite loops
  • Recovery logic: Triggers alternative strategies when impossible or evaluatorFailed states occur

This tight integration ensures that evaluation results translate immediately into runtime behavior without intermediate translation layers.

Working with the Evaluation Framework

Executing Goal Assessments

The primary entry point evaluateGoal accepts dependencies, goal conditions, and session context:

import { evaluateGoal } from './goal-evaluator.js';

const deps = {
  evaluate: async (prompt, sessionId, signal) => {
    // Tool-free model invocation (OpenAI, Anthropic, etc.)
    return await myModel.call({ prompt, signal });
  },
};

const result = await evaluateGoal(
  deps,
  condition,
  recentContext,
  'session-123'
);

if (result.met) {
  console.log('✅ Goal achieved!');
} else if (result.impossible) {
  console.log('❌ Goal impossible – aborting.');
} else if (result.progress) {
  console.log('▶️ Made progress, continue.');
} else if (result.waiting) {
  console.log('⏳ Waiting on external event.');
}

Handling Evaluator Failures

When timeouts or parsing errors occur, the framework signals failure explicitly:

if (result.evaluatorFailed) {
  // Treat as neutral – do not reset stall counters or assume progress
  console.warn('Evaluator failed (timeout or error); proceeding without progress.');
}

This pattern prevents the autonomous agent from misinterpreting technical failures as substantive progress reports.

Summary

  • Structured JSON Contract: The GoalEvaluation interface enforces strict typing with four terminal states (met, impossible, progress, waiting) and mandatory reasoning.
  • Tool-Free Architecture: Uses generateGoalEvaluationModelCall to avoid function-calling overhead and maintain text-only judgment.
  • Defensive Parsing: parseGoalEvaluation extracts valid JSON from malformed responses with fallback to evaluatorFailed states.
  • Production Hardening: Implements 30-second default timeouts, AbortController support, and reason text truncation (200 code units).
  • Runtime Integration: Directly attaches to GoalContinuation lane intents via packages/runtime/src/goal-continuation.ts for immediate routing decisions.

Frequently Asked Questions

What is the purpose of Maka's evaluation framework?

The evaluation framework serves as an automated judge that runs after each turn of an autonomous coding session. According to the Apache Maka source code, it determines whether the current goal state is met, impossible, making progress, or waiting, providing the runtime with deterministic signals to continue, retry, or abort the session.

How does the evaluation framework prevent self-justification bias?

The framework prevents self-justification by using a deterministic EVALUATOR_SYSTEM prompt that forbids the LLM from generating code or offering solutions. Additionally, the generateToolFreeModelCall function operates on the same session history but without access to the agent's internal tool calls, effectively creating a third-party assessment that cannot defend its own previous outputs.

What happens when the evaluator times out or fails?

When the 30-second timeout triggers or parsing fails, the evaluateGoal function returns an object with evaluatorFailed: true. This flag propagates through GoalContinuation logic to indicate that no valid assessment occurred, preventing the system from treating the failure as progress or a stall, and allowing the runtime to implement retry or fallback strategies.

How does the evaluation result drive autonomous coding decisions?

The result attaches directly to lane.intent in packages/runtime/src/goal-continuation.ts (lines 440-447). The runtime checks result.met to terminate successfully, result.impossible to abort the lane, result.progress to reset stall counters, and result.waiting to pause execution. This tight coupling ensures that LLM judgments translate immediately into concrete runtime actions without intermediate processing layers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →