Key Features of Maka's Evaluation Framework: LLM-Driven Goal Assessment for Autonomous Coding
Maka's evaluation framework is a lightweight, LLM-driven "judge" that runs after each turn of an autonomous coding session to classify goal status as met, impossible, making progress, or waiting on an external event.
The Apache Maka project implements a deterministic evaluation loop that prevents autonomous coding agents from陷入 infinite stalls or false positives. By leveraging a structured JSON contract and tool-free model calls, the framework provides time-bounded assessments that integrate directly with the runtime's continuation logic.
Core Architecture Components
Goal-Evaluation Interface and JSON Contract
At the heart of the system lies a strict JSON schema defined in packages/runtime/src/goal-evaluator.ts (lines 37-55). The GoalEvaluation interface requires the LLM to return one of four terminal states:
met: The goal condition is fully satisfiedimpossible: The goal cannot be achieved with available resourcesprogress: Forward momentum detected but goal not yet reachedwaiting: Paused for external events (user input, API responses)
Each evaluation must include a reason string explaining the judgment. This contract ensures that downstream components in GoalContinuation can route decisions without parsing ambiguous natural language.
Deterministic System Prompt Design
The framework uses a constant EVALUATOR_SYSTEM prompt (lines 108-120) that constrains the LLM to act strictly as a judge. Unlike conversational prompts, this system prompt:
- Explicitly forbids tool usage
- Mandates JSON-only output
- Prevents the model from offering solutions or code fixes
- Eliminates self-justification bias by separating the evaluator from the code generator
This separation is critical; the evaluator assesses the conversation history without access to the agent's internal reasoning, providing objective third-party judgment.
Tool-Free Model Invocation
To prevent function-call overhead and ensure text-only responses, the evaluator invokes generateToolFreeModelCall through the generateGoalEvaluationModelCall function (lines 89-100). This method:
- Bypasses tool registration entirely
- Enforces the JSON output schema through prompt engineering rather than structured outputs
- Reduces token consumption by eliminating tool description overhead
Implementation Details and Robustness Mechanisms
Prompt Construction Pipeline
The buildGoalEvaluationPrompt function (lines 22-33) concatenates three critical elements:
- The
EVALUATOR_SYSTEMinstructions - The specific goal condition being tested
- Recent conversation context (last N turns of the coding session)
import { buildGoalEvaluationPrompt } from './goal-evaluator.js';
const condition = 'The function `add(a, b)` must return the sum of its arguments.';
const recentContext = `User: Write a function that adds two numbers.
Agent: function add(a,b){return a+b;}
`;
const prompt = buildGoalEvaluationPrompt(condition, recentContext);
// Contains system prompt, condition, context, and "--- YOUR JUDGMENT (JSON only) ---" marker
Robust Parsing and Fallback Strategies
The parseGoalEvaluation function (lines 36-66) implements tolerant JSON extraction to handle malformed LLM outputs. When the model generates extra markdown fences or explanatory text, the parser extracts the first valid JSON object using regex patterns.
If parsing completely fails, the framework returns a neutral "failed" evaluation with evaluatorFailed: true. This prevents the autonomous loop from crashing or making irreversible decisions based on garbled output.
Hard Timeout and Cancellation Protection
Production deployments require strict latency bounds. The evaluateGoal function (lines 70-136) wraps the LLM call with:
- Configurable timeout: Default 30-second ceiling via
Promise.race - AbortController integration: Respects external cancellation signals
- Cleanup guarantees: Timer cleanup occurs regardless of success or failure
When timeouts trigger, the cancelledEvaluation helper propagates a special flag indicating the assessment aborted rather than returned a substantive judgment.
Reason Text Truncation
To prevent context window pollution, the EVALUATOR_REASON_TEXT_LIMIT constant (lines 103-106) caps the reason field at 200 code units (approximately 600 bytes). The truncateGoalText utility applies this limit before the evaluation result propagates to GoalContinuation, ensuring concise steering signals that fit within token budgets.
Integration with Goal Continuation
The evaluation framework feeds directly into Maka's lane management system. In packages/runtime/src/goal-continuation.ts (lines 440-447), the GoalEvaluation object attaches to lane.intent, driving:
- Continuation routing: Determines whether to proceed, retry, or terminate the session
- Stall detection: Distinguishes between genuine waiting states and infinite loops
- Recovery logic: Triggers alternative strategies when
impossibleorevaluatorFailedstates occur
This tight integration ensures that evaluation results translate immediately into runtime behavior without intermediate translation layers.
Working with the Evaluation Framework
Executing Goal Assessments
The primary entry point evaluateGoal accepts dependencies, goal conditions, and session context:
import { evaluateGoal } from './goal-evaluator.js';
const deps = {
evaluate: async (prompt, sessionId, signal) => {
// Tool-free model invocation (OpenAI, Anthropic, etc.)
return await myModel.call({ prompt, signal });
},
};
const result = await evaluateGoal(
deps,
condition,
recentContext,
'session-123'
);
if (result.met) {
console.log('✅ Goal achieved!');
} else if (result.impossible) {
console.log('❌ Goal impossible – aborting.');
} else if (result.progress) {
console.log('▶️ Made progress, continue.');
} else if (result.waiting) {
console.log('⏳ Waiting on external event.');
}
Handling Evaluator Failures
When timeouts or parsing errors occur, the framework signals failure explicitly:
if (result.evaluatorFailed) {
// Treat as neutral – do not reset stall counters or assume progress
console.warn('Evaluator failed (timeout or error); proceeding without progress.');
}
This pattern prevents the autonomous agent from misinterpreting technical failures as substantive progress reports.
Summary
- Structured JSON Contract: The
GoalEvaluationinterface enforces strict typing with four terminal states (met, impossible, progress, waiting) and mandatory reasoning. - Tool-Free Architecture: Uses
generateGoalEvaluationModelCallto avoid function-calling overhead and maintain text-only judgment. - Defensive Parsing:
parseGoalEvaluationextracts valid JSON from malformed responses with fallback toevaluatorFailedstates. - Production Hardening: Implements 30-second default timeouts, AbortController support, and reason text truncation (200 code units).
- Runtime Integration: Directly attaches to
GoalContinuationlane intents viapackages/runtime/src/goal-continuation.tsfor immediate routing decisions.
Frequently Asked Questions
What is the purpose of Maka's evaluation framework?
The evaluation framework serves as an automated judge that runs after each turn of an autonomous coding session. According to the Apache Maka source code, it determines whether the current goal state is met, impossible, making progress, or waiting, providing the runtime with deterministic signals to continue, retry, or abort the session.
How does the evaluation framework prevent self-justification bias?
The framework prevents self-justification by using a deterministic EVALUATOR_SYSTEM prompt that forbids the LLM from generating code or offering solutions. Additionally, the generateToolFreeModelCall function operates on the same session history but without access to the agent's internal tool calls, effectively creating a third-party assessment that cannot defend its own previous outputs.
What happens when the evaluator times out or fails?
When the 30-second timeout triggers or parsing fails, the evaluateGoal function returns an object with evaluatorFailed: true. This flag propagates through GoalContinuation logic to indicate that no valid assessment occurred, preventing the system from treating the failure as progress or a stall, and allowing the runtime to implement retry or fallback strategies.
How does the evaluation result drive autonomous coding decisions?
The result attaches directly to lane.intent in packages/runtime/src/goal-continuation.ts (lines 440-447). The runtime checks result.met to terminate successfully, result.impossible to abort the lane, result.progress to reset stall counters, and result.waiting to pause execution. This tight coupling ensures that LLM judgments translate immediately into concrete runtime actions without intermediate processing layers.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →