How Maka's Crash Recovery and Resume System Achieves Process-Crash Convergence
Maka guarantees process-crash convergence by replaying immutable events from a file-backed ledger, detecting a verifiable safe boundary, and deterministically repairing or discarding in-flight tool calls before allowing new execution.
Apache Maka implements a deterministic crash recovery mechanism that ensures no committed work is lost after a process-level failure. The system treats the runtime as an immutable event ledger, using a three-phase recovery protocol to classify events, detect safe resumption points, and repair corrupted states. According to the Apache Maka source code, this approach guarantees that only fully verified state is carried forward while partial tool executions are either idempotently repaired or safely discarded.
The Three-Phase Recovery Architecture
Maka's recovery flow operates across three conceptual phases implemented in packages/runtime/src/ and documented in the architecture specifications.
Phase 0: Establishing the Crash Contract
The foundation of Maka's recovery system rests on process-crash committed-prefix semantics. This phase defines what state remains durable after a crash and what guarantees the runtime must uphold. The system relies on an immutable RuntimeEventStore that maintains a deterministic safe boundary separating trusted history from untrusted recent events. The formal contract is documented in docs/architecture/runtime-resume-phase0-crash-contract.md, which specifies that committed entries are never rewritten and that the ledger provides a consistent prefix guarantee even after unexpected termination.
Phase 1: Safe Boundary Detection
During this phase, the RecoveryResolver class (defined in packages/runtime/src/recovery-resolver.ts) scans the ledger from newest to oldest to locate the safe_boundary_continuation. The resolver classifies each RuntimeEvent into three categories:
- Completed: Events fully persisted to the store with
resumeTrust=trusted - In-flight: Tool calls partially executed or interrupted
- Corrupted: Events with inconsistent checksums or partial writes
The resolver verifies that no tool remains in a half-executed state at the boundary. It also respects hard gates such as parked and recovery.hasCorruption that block automatic resume until user approval is obtained, as detailed in docs/architecture/runtime-resume-phase3-phase4-workspace-checkpoint-design.md.
Phase 2: Replay and Repair
Once the safe boundary is identified, the AgentRunRecovery class (in packages/runtime/src/agent-run-recovery.ts) re-opens the RuntimeEventStore and replays verified events to reconstruct the in-memory AgentRun state. For interrupted tool calls, the engine executes a repair saga:
- Repair: Re-executes the tool with identical arguments when the tool is idempotent
- Abort: Discards the incomplete tool call and continues with the next model turn
Only after successful repair does the system allow a new Run object to be created, with metadata recording both the repair actions and the original crash cause.
Core Recovery Mechanics
Immutable Event Ledger and RuntimeEventStore
All observable actions—including tool calls, model turns, and filesystem writes—are appended to a file-backed RuntimeEventStore. Each entry is immutable; later events may supersede earlier partial snapshots but never modify committed history. This design ensures that after a crash, the host can re-open the store and read a consistent prefix of history without encountering partial writes or torn pages.
Event Classification with RecoveryResolver
The RecoveryResolver implements the logic for determining whether a session can resume safely. It evaluates the resumeTrust property of each event:
trusted: Fully persisted and safe to continueuntrusted: Tool in-flight or corrupted; requires repair or discard
The resolver returns the index of the last trusted event, which defines the safe boundary for reconstruction.
import { RecoveryResolver } from '@maka/runtime';
import { RuntimeEventStore } from '@maka/runtime';
const store = new RuntimeEventStore('/path/to/store');
const resolver = new RecoveryResolver(store);
const safeBoundary = await resolver.findSafeBoundary();
// Returns the last trusted event index
Deterministic State Restoration via AgentRunRecovery
The AgentRunRecovery class manages the transition from persisted ledger to active runtime. It replays events up to the safe boundary to recreate the AgentRun state, then handles any necessary repairs. The recovery path is fully deterministic because it depends solely on immutable events; there is no reliance on in-memory caches that could diverge after a crash.
User-Controlled Resume API
Maka exposes resume functionality through both CLI and desktop interfaces, requiring explicit user opt-in to prevent accidental continuation.
The CLI provides the /resume command, while the desktop interface offers sessions:resumeLatest. However, these commands only execute full recovery when the MAKA_RUNTIME_SAFE_BOUNDARY_RESUME environment flag is set to 1. Without this flag, the system falls back to a best-effort retry that does not guarantee true continuity.
import { RuntimeHostCLI } from '@maka/cli';
// Request safe resume from CLI
await RuntimeHostCLI.command('/resume');
The safe boundary contract is formally defined in docs/architecture/runtime-resume-phase1-safe-boundary-contract.md, which mandates that users must explicitly enable the feature flag and approve repair actions before the system resumes from a verified checkpoint.
Guarantees and Safety Properties
Maka's crash recovery system provides three fundamental guarantees:
-
Process-Crash Convergence: The durable prefix of the ledger remains identical across the original and restarted host. No partial tool output is ever resurrected or double-executed without explicit repair handling.
-
Deterministic Restart: Recovery depends only on the immutable event sequence. Given the same ledger, the system always reconstructs identical runtime state, eliminating nondeterminism from cache contents or timing-dependent variables.
-
Explicit User Control: Automatic resume requires both the safe boundary feature flag and explicit approval of any repair actions. This prevents the system from misrepresenting a retry attempt as true session continuity.
Summary
- Maka implements a three-phase recovery protocol: Crash Contract definition, Safe Boundary Detection, and Replay & Repair.
- The
RecoveryResolverinpackages/runtime/src/recovery-resolver.tsclassifies events as trusted or untrusted to identify the safe resumption point. - The
AgentRunRecoveryclass replays immutable events and executes repair sagas for interrupted tool calls. - All state changes are persisted to an immutable
RuntimeEventStorethat guarantees process-crash committed-prefix semantics. - Safe resume requires the
MAKA_RUNTIME_SAFE_BOUNDARY_RESUME=1flag and user approval of any repair actions.
Frequently Asked Questions
What happens if a tool call is interrupted during a process crash?
If a tool call is in-flight when the crash occurs, Maka's AgentRunRecovery class detects the incomplete execution during the Replay & Repair phase. Depending on the tool's idempotency guarantee, the system either re-executes the tool with identical arguments (repair) or discards the partial call and continues with the next turn (abort). The outcome is recorded in the new Run object's metadata.
How does Maka distinguish between trusted and untrusted events during recovery?
The RecoveryResolver scans the event ledger from newest to oldest, assigning each RuntimeEvent a resumeTrust classification. Events marked trusted are fully persisted and safe to replay. Events marked untrusted indicate in-flight tool execution or corruption; these block automatic resume until the user explicitly approves repair or the system discards the suffix.
Is automatic crash recovery enabled by default in Maka?
No. Automatic resume is strictly opt-in. Users must enable the MAKA_RUNTIME_SAFE_BOUNDARY_RESUME environment variable and set it to 1. Additionally, hard gates such as parked or recovery.hasCorruption flags prevent automatic resumption until the user explicitly approves the safe continuation, preventing accidental retry semantics.
Where is the crash recovery logic implemented in the Maka source code?
The core recovery logic resides in packages/runtime/src/recovery-resolver.ts (event classification and safe boundary detection) and packages/runtime/src/agent-run-recovery.ts (state replay and repair execution). The formal contracts and safety guarantees are documented in docs/architecture/runtime-resume-phase0-crash-contract.md and docs/architecture/runtime-resume-phase1-safe-boundary-contract.md.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →