How Maka's runtime-resume.ts Module Orchestrates Crash Recovery
Maka's crash-recovery mechanism centers on the RuntimeContinuationPlanner class in runtime-resume.ts, which reconstructs execution state by validating continuation lineage, reconciling immutable prefixes, and determining whether a crashed run can safely resume or must be parked for manual intervention.
When an Apache Maka runtime crashes, the system must reconstruct the exact execution context to determine if work can safely proceed. The runtime-resume.ts module implements this orchestration through a deterministic pipeline that validates immutable history, checks durable claims, and enforces safety boundaries. This article examines how the RuntimeContinuationPlanner analyzes crashed runs and makes continuation decisions based on lineage integrity and provider compatibility.
The RuntimeContinuationPlanner Architecture
The RuntimeContinuationPlanner serves as the central orchestrator for crash recovery in the Apache Maka runtime. When invoked, it executes a seven-step workflow that transforms a crashed run identifier into either a resumable RuntimeContinuation object or a parked plan requiring manual intervention. The planner depends on several injected interfaces—including readSourceRun, readImmutableRuntimePrefix, and readContinuationClaimStateByBoundary—to interact with storage layers without coupling to specific persistence mechanisms.
Step-by-Step Recovery Orchestration
Loading the Source Run and Validating Headers
The recovery process begins by loading the crashed run's header via readSourceRun(input.sessionId, input.sourceRunId). This retrieves the AgentRunHeader containing the run's metadata and state pointers. If the header is missing or unreadable, the planner immediately returns a parked plan with the reason 'source_run_unreadable', halting further processing to prevent operations on undefined state.
Traversing the Immutable Lineage
Next, the planner invokes readLineagePrefixes to fetch the immutable runtime prefix of the crashed run and recursively walk its continuation lineage. This process loads each ancestor's prefix to construct a complete history chain. The lineage validation enforces strict invariants: it detects cycles, validates that depth does not exceed 64 ancestors, and verifies identity consistency across the chain. Failures here raise RuntimeLineageError and result in parked plans with specific diagnostics such as 'runtime_lineage_cycle' or 'runtime_identity_mismatch'.
Generating the Provider Replay Plan
With validated prefixes collected, the planner calls buildContinuationReplayPlan from continuation-replay.ts, passing the prefixes and PROVIDER_REPLAY_PROJECTION_VERSION. This function constructs an immutable provider-side view of the execution history. If the replay plan encounters non-suffix gaps or unsupported projection features, the planner blocks continuation and parks the run with reasons 'provider_replay_non_suffix_gap' or 'provider_replay_unsupported'. Additionally, recovery-resolver.ts contributes corruption detection; if tool recovery reveals unsettled decisions or corruption, the plan is blocked with 'tool_recovery_corruption' or 'tool_recovery_unsettled'.
Reconciling Durable Continuation Claims
The planner queries readContinuationClaimStateByBoundary to check for existing durable claims associated with the run's digest. When a claim exists, the system verifies that the claim's boundary, digest, and projection version exactly match the generated replay plan. Any discrepancy indicates potential state divergence, causing the planner to return a parked plan with reason 'continuation_claim_repair_required' to signal that manual repair is necessary before resumption.
Preventing Duplicate Continuations
Before creating new state, the planner invokes findExistingContinuation to search for continuations already created for the same source run and high-water mark. If a duplicate exists, the planner prevents double-processing by returning a parked plan with reason 'continuation_already_exists'. This idempotency guard ensures that transient failures during recovery do not spawn multiple active continuations for the same crash event.
Enforcing Safe Boundary Constraints
When no blocking claims or duplicates exist, the planner delegates to buildSafeBoundaryContinuationPlan to perform comprehensive safety checks. This function validates workspace identity consistency, confirms that background operations have settled, verifies that required tools remain available in the catalog, and checks for pending permission requests. Violations accumulate in phaseOneDiagnostics and typically force a 'park' disposition, while passing all checks produces a 'continue' disposition.
Returning the SafeBoundaryContinuationPlan
The final step assembles and returns a SafeBoundaryContinuationPlan object. This structure contains the disposition ('continue' or 'park'), diagnostic messages, rejection reasons, and—if continuation is permitted—the fully populated RuntimeContinuation object. The continuation field includes the reconstructed runtimeContext required to resume execution exactly where the crash occurred.
Failure Modes and Diagnostic Parking
The runtime-resume.ts module categorizes failure modes into specific parking reasons to facilitate automated monitoring and manual remediation:
- Source run unreadable: Immediate parking when
readSourceRunfails to retrieve theAgentRunHeader. - Lineage integrity violations: Cycle detection, depth limit breaches (>64), or identity mismatches trigger
RuntimeLineageErrorand specific park reasons. - Provider replay incompatibility: Non-suffix gaps or unsupported projections in the immutable history block replay.
- Tool recovery corruption:
recovery-resolver.tsdetects unsettled tool recovery decisions or corruption states. - Durable claim mismatches: Boundary, digest, or version mismatches between the claim and replay plan require repair.
- Duplicate continuation attempts: Existing continuations for the same source run prevent redundant creation.
- Safety boundary violations: Workspace identity changes, pending permissions, or unavailable tools trigger parking with detailed diagnostics.
Practical Implementation Example
The following TypeScript example demonstrates how to initialize the planner and execute crash recovery:
// 1. Initialize the planner with required dependencies
const planner = new RuntimeContinuationPlanner({
readSourceRun: async (sessionId, runId) => {
// Fetch the AgentRunHeader from your storage layer
return await runHeaderStore.get(sessionId, runId);
},
readImmutableRuntimePrefix: async ({ sessionId, runId, upToEventSeq }) => {
// Retrieve the immutable prefix (compact ledger) for the run
return await prefixStore.get(sessionId, runId, upToEventSeq);
},
readContinuationClaimStateByBoundary: async (digest) => {
// Optional: query a durable claim service (e.g., a DB) for a claim
return await claimStore.getByDigest(digest);
},
findExistingContinuation: async (sessionId, sourceRunId, highWater) => {
// Optional: look for an already-created continuation run
return await continuationIndex.find(sessionId, sourceRunId, highWater);
},
newId: () => crypto.randomUUID(),
});
// 2. Build the input describing the current environment
const input: RuntimeContinuationPlannerInput = {
sessionId: 'abc123',
sourceRunId: crashedRunId,
currentCwd: process.cwd(),
sourceWorkspaceIdentity: 'workspace-001',
currentWorkspaceIdentity: 'workspace-001',
backgroundOperationsSettled: true,
availableToolNames: ['search', 'summarize'],
// Optional expected high-water mark if you stored a checkpoint
expectedRuntimeEventHighWater: 42,
};
// 3. Ask the planner for a recovery plan
const plan = await planner.plan(input);
if (plan.disposition === 'continue') {
// Safe to resume – you now have a RuntimeContinuation object
const continuation = plan.continuation!;
// Pass continuation.runtimeContext to your runtime executor
await resumeExecution(continuation);
} else {
// Parked – inspect diagnostics for manual remediation
console.warn('Recovery parked:', plan.diagnostics);
}
Key Source Files in the Recovery Pipeline
The crash recovery system spans multiple files in the Apache Maka repository:
runtime-resume.ts: ImplementsRuntimeContinuationPlanner, orchestrating lineage traversal, claim validation, and safe-boundary planning.continuation-replay.ts: Generates immutable provider-replay plans viabuildContinuationReplayPlanand validates segment suitability.recovery-resolver.ts: SuppliesresolveRuntimeRecoveryto detect tool-recovery status, corruption, and pending decisions used throughout the planner.model-history.ts: Contains model-level replay logic (buildRuntimeEventModelReplayPlan) that feeds intocontinuation-replay.ts.runtime-event.ts: Defines theRuntimeEventschema and helpers (isPartialRuntimeEvent,isTerminalRuntimeEvent) used in prefix validation.
Summary
- The
RuntimeContinuationPlannerinruntime-resume.tsprovides the central orchestration for Apache Maka crash recovery. - Recovery proceeds through seven validated steps: loading headers, traversing lineage, building replay plans, checking claims, preventing duplicates, enforcing safety boundaries, and returning continuation plans.
- The system uses immutable prefixes and provider-replay projections to ensure deterministic state reconstruction.
- Multiple failure modes—including lineage cycles, claim mismatches, and safety violations—result in parked plans with specific diagnostic codes for manual intervention.
- Implementation requires injecting storage adapters for run headers, prefixes, and claims, then polling the planner's disposition to determine if resumption is safe.
Frequently Asked Questions
What is the role of the RuntimeContinuationPlanner in Apache Maka?
The RuntimeContinuationPlanner serves as the core orchestrator that determines whether a crashed Maka run can safely resume. It validates the execution lineage, reconciles immutable state prefixes, checks for existing durable claims, and enforces safety boundaries before permitting a new continuation to proceed.
How does Maka prevent duplicate continuations during crash recovery?
The planner queries findExistingContinuation to check if a continuation already exists for the same source run and high-water mark. If a duplicate is detected, it returns a parked plan with reason 'continuation_already_exists', preventing redundant execution and ensuring idempotency across recovery attempts.
What causes a run to be parked instead of continued?
A run is parked when any validation step fails, including unreadable source headers, lineage cycles or depth limits exceeding 64 ancestors, tool recovery corruption, provider replay incompatibility, durable claim mismatches, or safety boundary violations such as workspace identity changes or pending permissions.
How does the lineage validation work in runtime-resume.ts?
The readLineagePrefixes function recursively walks the continuation ancestry, loading each ancestor's immutable prefix. It validates that no cycles exist, that the depth does not exceed 64, and that workspace identities match across the chain. Any violation raises RuntimeLineageError and triggers a parked plan with specific diagnostic codes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →