How Apache Maka's RecoveryResolver Classifies Tool State After Runtime Crashes
The RecoveryResolver analyzes immutable RuntimeEvent records to classify every tool invocation into one of five definitive states—completed, not-dispatched, unknown, parked, or corrupt—ensuring a single source of truth for crash recovery decisions.
When a model invokes external tools in Apache Maka, the runtime must survive abrupt crashes without corrupting execution state. The RecoveryResolver serves as the sole authority that examines the immutable event ledger to classify tool state, enabling safe operational resumption or fresh execution triggers.
The Five Tool State Classifications
The RecoveryResolver categorizes every tool operation into exactly one of five mutually exclusive classifications based on evidence found in the event ledger. According to the implementation in packages/runtime/src/recovery-resolver.ts, these states cover every possible lifecycle outcome.
completed
A tool invocation reaches the completed state when the ledger contains definitive proof that the tool finished execution and recorded a result. This represents the only terminal state where the runtime can safely consume tool output without re-execution.
not-dispatched
The not-dispatched (or "definitely not dispatched") state applies when the event ledger shows the tool call was aborted before leaving the runtime boundary. The invocation never reached the external tool, making it safe to retry without side-effect concerns.
unknown
Tools enter the unknown state when dispatched but no result arrived—typically due to network failures or external service crashes mid-flight. The resolver cannot determine whether the side effects actually occurred, requiring manual review or idempotent retry strategies.
parked
A parked classification indicates the tool awaits external input or human decision. The runtime pauses execution at this boundary until the parking condition resolves through external intervention.
corrupt
The corrupt state triggers when the event ledger contains inconsistencies or missing crucial events that make state determination impossible. This represents a critical system error requiring operator intervention.
How RecoveryResolver Processes Runtime Events
The classification algorithm follows a strict pipeline that treats the event store as the sole source of truth, preventing conflicting interpretations across system components.
Immutable Event Ledger as Source of Truth
All runtime actions persist as immutable RuntimeEvent records. After a crash, the resolver re-reads these events from the ledger—located conceptually within the runtime event store—to reconstruct the exact sequence of tool invocations. Because events are append-only and immutable, they provide tamper-proof evidence for classification decisions.
Scanning and Analysis
The RecoveryResolver.analyze() method accepts the event stream and examines each tool-related record type, including "tool-invoked", "tool-dispatch-sent", and "tool-result-received" events. It correlates these entries by tool call ID to determine which of the five states applies to each unique invocation.
Single Authority Constraint
All higher-level components—including the Planner, CLI, UI, and future reconcilers—must consume the resolver's output directly. They are explicitly prohibited from recomputing their own state judgments. This architectural constraint, enforced by the runtime design, guarantees a consistent view of tool status across the entire system.
Safe Resume Boundaries
The runtime only initiates a new Run when the RecoveryResolver confirms that history, execution context, and workspace boundaries are provably safe. The old process is never resurrected; "try again" operations become fresh executions rather than recovery attempts. This safety check prevents partial execution corruption.
Implementation Details and Source Locations
The RecoveryResolver implementation and its architectural specifications reside in specific paths within the apache/maka repository:
- Core implementation:
packages/runtime/src/recovery-resolver.tscontains theRecoveryResolverclass and itsanalyze()method - Architecture decision record:
docs/architecture/runtime-recovery-resolver-adr.zh-CN.mddetails the state classification logic and design rationale - Runtime resume specification:
docs/architecture/runtime-resume-architecture.mddescribes how the resolver integrates into crash recovery flows - Test suite:
packages/runtime/src/__tests__/recovery-resolver.test.tsprovides comprehensive coverage of state transition edge cases
Working with RecoveryResolver in Code
To classify tool states after a crash, load the event ledger and invoke the analyzer:
// Example: Using RecoveryResolver to check a tool's state after a crash
import { RecoveryResolver } from '@apache/maka/runtime';
// Assume we have a Run ID that crashed
const runId = 'run-12345';
// Load the immutable event ledger for that run
const events = await RuntimeEventReadModel.load(runId);
// Resolve tool states
const toolDecisions = RecoveryResolver.analyze(events);
// Inspect a specific tool invocation (identified by its tool call ID)
const toolId = 'tool-call-abcdef';
const decision = toolDecisions.get(toolId);
switch (decision.state) {
case 'completed':
console.log('Tool finished – result:', decision.result);
break;
case 'not-dispatched':
console.log('Tool never left the runtime – safe to retry.');
break;
case 'unknown':
console.warn('Tool was dispatched but no result – manual review needed.');
break;
case 'parked':
console.log('Tool is waiting for external input – show UI to user.');
break;
case 'corrupt':
console.error('Ledger corrupted – abort run and alert operators.');
break;
}
When implementing resume logic, validate that all tool decisions meet safety criteria before proceeding:
// Example: Integrating the resolver into a resume flow
async function resumeRun(runId: string) {
const events = await RuntimeEventReadModel.load(runId);
const decisions = RecoveryResolver.analyze(events);
// Verify that **all** tool calls are either completed or safely retryable
const safe = [...decisions.values()].every(dec =>
dec.state === 'completed' || dec.state === 'not-dispatched'
);
if (!safe) {
throw new Error('Cannot safely resume – some tools are in unknown/parked/corrupt state.');
}
// Create a fresh Run that will continue from the last safe checkpoint
const newRun = await RunFactory.startFromCheckpoint(runId);
return newRun;
}
Summary
- The
RecoveryResolverinpackages/runtime/src/recovery-resolver.tsprovides the definitive classification of tool states after crashes. - It recognizes five distinct states:
completed,not-dispatched,unknown,parked, andcorrupt. - Classification relies exclusively on immutable
RuntimeEventrecords, making the event ledger the single source of truth. - All system components must consume resolver output rather than performing independent state calculations.
- Safe resume operations require all tools to be either completed or not-dispatched; other states prevent automatic continuation.
Frequently Asked Questions
What distinguishes the 'unknown' state from 'not-dispatched'?
The not-dispatched state indicates the tool call never left the runtime boundary, while unknown means the tool was dispatched but the result never arrived due to external failures. The former is safe to retry; the latter requires manual verification because side effects may have occurred.
Why must higher-level components use RecoveryResolver output instead of their own logic?
Architectural constraints in Apache Maka require all components—Planner, CLI, and UI—to consume RecoveryResolver classifications as the single source of truth. This prevents state inconsistencies that would occur if different subsystems computed conflicting interpretations of the same event ledger.
Where is the RecoveryResolver implementation located?
The core implementation resides in packages/runtime/src/recovery-resolver.ts within the apache/maka repository. The accompanying architecture decision record is located at docs/architecture/runtime-recovery-resolver-adr.zh-CN.md.
How does RecoveryResolver handle corrupt event ledgers?
When the event ledger contains inconsistencies or missing crucial records, the resolver classifies the affected tool invocation as corrupt. This triggers error handling that aborts the run and alerts operators, preventing execution on potentially compromised state.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →