# Key Features of Maka's Evaluation Framework: LLM-Driven Goal Assessment for Autonomous Coding

> Discover Maka's LLM-driven evaluation framework for autonomous coding. It assesses goal status after each turn to optimize development.

- Repository: [The Apache Software Foundation/maka](https://github.com/apache/maka)
- Tags: deep-dive
- Published: 2026-08-22

---

**Maka's evaluation framework is a lightweight, LLM-driven "judge" that runs after each turn of an autonomous coding session to classify goal status as met, impossible, making progress, or waiting on an external event.**

The Apache Maka project implements a deterministic evaluation loop that prevents autonomous coding agents from陷入 infinite stalls or false positives. By leveraging a structured JSON contract and tool-free model calls, the framework provides time-bounded assessments that integrate directly with the runtime's continuation logic.

## Core Architecture Components

### Goal-Evaluation Interface and JSON Contract

At the heart of the system lies a strict JSON schema defined in [`packages/runtime/src/goal-evaluator.ts`](https://github.com/apache/maka/blob/main/packages/runtime/src/goal-evaluator.ts) (lines 37-55). The `GoalEvaluation` interface requires the LLM to return one of four terminal states:

- **`met`**: The goal condition is fully satisfied
- **`impossible`**: The goal cannot be achieved with available resources
- **`progress`**: Forward momentum detected but goal not yet reached
- **`waiting`**: Paused for external events (user input, API responses)

Each evaluation must include a `reason` string explaining the judgment. This contract ensures that downstream components in `GoalContinuation` can route decisions without parsing ambiguous natural language.

### Deterministic System Prompt Design

The framework uses a constant `EVALUATOR_SYSTEM` prompt (lines 108-120) that constrains the LLM to act strictly as a judge. Unlike conversational prompts, this system prompt:

- Explicitly forbids tool usage
- Mandates JSON-only output
- Prevents the model from offering solutions or code fixes
- Eliminates self-justification bias by separating the evaluator from the code generator

This separation is critical; the evaluator assesses the conversation history without access to the agent's internal reasoning, providing objective third-party judgment.

### Tool-Free Model Invocation

To prevent function-call overhead and ensure text-only responses, the evaluator invokes `generateToolFreeModelCall` through the `generateGoalEvaluationModelCall` function (lines 89-100). This method:

1. Bypasses tool registration entirely
2. Enforces the JSON output schema through prompt engineering rather than structured outputs
3. Reduces token consumption by eliminating tool description overhead

## Implementation Details and Robustness Mechanisms

### Prompt Construction Pipeline

The `buildGoalEvaluationPrompt` function (lines 22-33) concatenates three critical elements:

1. The `EVALUATOR_SYSTEM` instructions
2. The specific goal condition being tested
3. Recent conversation context (last N turns of the coding session)

```typescript
import { buildGoalEvaluationPrompt } from './goal-evaluator.js';

const condition = 'The function `add(a, b)` must return the sum of its arguments.';
const recentContext = `User: Write a function that adds two numbers.
Agent: function add(a,b){return a+b;}
`;

const prompt = buildGoalEvaluationPrompt(condition, recentContext);
// Contains system prompt, condition, context, and "--- YOUR JUDGMENT (JSON only) ---" marker

```

### Robust Parsing and Fallback Strategies

The `parseGoalEvaluation` function (lines 36-66) implements tolerant JSON extraction to handle malformed LLM outputs. When the model generates extra markdown fences or explanatory text, the parser extracts the first valid JSON object using regex patterns.

If parsing completely fails, the framework returns a neutral "failed" evaluation with `evaluatorFailed: true`. This prevents the autonomous loop from crashing or making irreversible decisions based on garbled output.

### Hard Timeout and Cancellation Protection

Production deployments require strict latency bounds. The `evaluateGoal` function (lines 70-136) wraps the LLM call with:

- **Configurable timeout**: Default 30-second ceiling via `Promise.race`
- **AbortController integration**: Respects external cancellation signals
- **Cleanup guarantees**: Timer cleanup occurs regardless of success or failure

When timeouts trigger, the `cancelledEvaluation` helper propagates a special flag indicating the assessment aborted rather than returned a substantive judgment.

### Reason Text Truncation

To prevent context window pollution, the `EVALUATOR_REASON_TEXT_LIMIT` constant (lines 103-106) caps the `reason` field at 200 code units (approximately 600 bytes). The `truncateGoalText` utility applies this limit before the evaluation result propagates to `GoalContinuation`, ensuring concise steering signals that fit within token budgets.

## Integration with Goal Continuation

The evaluation framework feeds directly into Maka's lane management system. In [`packages/runtime/src/goal-continuation.ts`](https://github.com/apache/maka/blob/main/packages/runtime/src/goal-continuation.ts) (lines 440-447), the `GoalEvaluation` object attaches to `lane.intent`, driving:

- **Continuation routing**: Determines whether to proceed, retry, or terminate the session
- **Stall detection**: Distinguishes between genuine waiting states and infinite loops
- **Recovery logic**: Triggers alternative strategies when `impossible` or `evaluatorFailed` states occur

This tight integration ensures that evaluation results translate immediately into runtime behavior without intermediate translation layers.

## Working with the Evaluation Framework

### Executing Goal Assessments

The primary entry point `evaluateGoal` accepts dependencies, goal conditions, and session context:

```typescript
import { evaluateGoal } from './goal-evaluator.js';

const deps = {
  evaluate: async (prompt, sessionId, signal) => {
    // Tool-free model invocation (OpenAI, Anthropic, etc.)
    return await myModel.call({ prompt, signal });
  },
};

const result = await evaluateGoal(
  deps,
  condition,
  recentContext,
  'session-123'
);

if (result.met) {
  console.log('✅ Goal achieved!');
} else if (result.impossible) {
  console.log('❌ Goal impossible – aborting.');
} else if (result.progress) {
  console.log('▶️ Made progress, continue.');
} else if (result.waiting) {
  console.log('⏳ Waiting on external event.');
}

```

### Handling Evaluator Failures

When timeouts or parsing errors occur, the framework signals failure explicitly:

```typescript
if (result.evaluatorFailed) {
  // Treat as neutral – do not reset stall counters or assume progress
  console.warn('Evaluator failed (timeout or error); proceeding without progress.');
}

```

This pattern prevents the autonomous agent from misinterpreting technical failures as substantive progress reports.

## Summary

- **Structured JSON Contract**: The `GoalEvaluation` interface enforces strict typing with four terminal states (met, impossible, progress, waiting) and mandatory reasoning.
- **Tool-Free Architecture**: Uses `generateGoalEvaluationModelCall` to avoid function-calling overhead and maintain text-only judgment.
- **Defensive Parsing**: `parseGoalEvaluation` extracts valid JSON from malformed responses with fallback to `evaluatorFailed` states.
- **Production Hardening**: Implements 30-second default timeouts, AbortController support, and reason text truncation (200 code units).
- **Runtime Integration**: Directly attaches to `GoalContinuation` lane intents via [`packages/runtime/src/goal-continuation.ts`](https://github.com/apache/maka/blob/main/packages/runtime/src/goal-continuation.ts) for immediate routing decisions.

## Frequently Asked Questions

### What is the purpose of Maka's evaluation framework?

The evaluation framework serves as an automated judge that runs after each turn of an autonomous coding session. According to the Apache Maka source code, it determines whether the current goal state is **met**, **impossible**, **making progress**, or **waiting**, providing the runtime with deterministic signals to continue, retry, or abort the session.

### How does the evaluation framework prevent self-justification bias?

The framework prevents self-justification by using a deterministic `EVALUATOR_SYSTEM` prompt that forbids the LLM from generating code or offering solutions. Additionally, the `generateToolFreeModelCall` function operates on the same session history but without access to the agent's internal tool calls, effectively creating a third-party assessment that cannot defend its own previous outputs.

### What happens when the evaluator times out or fails?

When the 30-second timeout triggers or parsing fails, the `evaluateGoal` function returns an object with `evaluatorFailed: true`. This flag propagates through `GoalContinuation` logic to indicate that no valid assessment occurred, preventing the system from treating the failure as progress or a stall, and allowing the runtime to implement retry or fallback strategies.

### How does the evaluation result drive autonomous coding decisions?

The result attaches directly to `lane.intent` in [`packages/runtime/src/goal-continuation.ts`](https://github.com/apache/maka/blob/main/packages/runtime/src/goal-continuation.ts) (lines 440-447). The runtime checks `result.met` to terminate successfully, `result.impossible` to abort the lane, `result.progress` to reset stall counters, and `result.waiting` to pause execution. This tight coupling ensures that LLM judgments translate immediately into concrete runtime actions without intermediate processing layers.