How Apache Maka Handles Crash Recovery and Resuming Operations: A Technical Deep Dive

Apache Maka ensures operational continuity after process crashes through a three-tier architecture of crash-boundary transactions, safe-boundary continuations, and quarantine-based recovery bundles that validate and replay interrupted work from durable commit points.

Apache Maka's approach to crash recovery centers on preventing data loss during unexpected terminations while enabling seamless resuming of operations. The system implements durable transaction boundaries and continuation replay mechanisms that allow agents to pick up exactly where they left off, even after complete process failure.

Core Crash Recovery Mechanisms

Maka's crash recovery strategy relies on three tightly coupled subsystems that coordinate to detect, isolate, and replay interrupted operations.

Crash-Boundary Transactions

Critical state-changing actions—such as persisting plan executions, committing workspace baselines, or updating agent graphs—execute inside crash-boundary sections. Each boundary writes a durable commit record before performing the actual mutation and validates it after completion.

According to the Apache Maka source code in packages/storage/src/plan-store.ts, if a process crashes between these two points, the system detects the incomplete transaction at lines 377-388 and triggers rollback or retry logic. This ensures that partial writes never leave the plan store in an inconsistent state.

Safe-Boundary Continuations

The runtime kernel exposes a resumeContinuation method that replays execution state after a crash, but only when the continuation was captured at a safe-boundary—meaning all required durable writes have already succeeded.

In packages/runtime/src/session-manager.ts (lines 2270-2290), the session-manager validates that resumes are permitted only when the runtime kernel supports the feature and the continuation passes resumeTrust checks. This prevents replay of untrusted or corrupted state segments.

Quarantine and Recovery Bundles

When crashes occur during workspace staging, Maka's quarantine logic preserves the staging inode intact without touching decoy files. The system later re-hydrates the workspace from a recovery bundle rather than attempting to salvage partial state.

As implemented in packages/storage/src/git-workspace-service.ts (lines 1472-1507), this mechanism guarantees that crashes leave no orphaned state and that subsequent runs can reclaim staging areas cleanly without contaminating the working directory.

Implementing Resume Operations

Maka exposes both high-level APIs and CLI interfaces for resuming operations explicitly or automatically after crash detection.

Resuming Plan Executions Programmatically

Applications can invoke resumePlanExecution through the session manager to restart specific operations using their original session and execution identifiers:

// Resume a previously interrupted plan execution
await runtime.sessionManager.resumePlanExecution({
  sessionId: 'my-session-id',
  executionId: 'exec-123',
  operationId: 'op-456',
});

This method locates the persisted session header in packages/storage/src/session-store.ts (lines 759-760) to identify the appropriate resume point before invoking the safe-boundary continuation logic.

CLI-Based Resume Workflows

For manual recovery scenarios, operators use the --resume flag with a legacy run identifier:


# Invoke resume command with a legacy run identifier

maka --resume <legacy-run-id>

The CLI validates the resume candidate against stored session headers before attempting reconstruction.

Low-Level Continuation Replay

Advanced use cases can manually invoke the kernel's continuation replay mechanism with proper abort signal handling:

// Manually invoking a safe-boundary continuation
if (runtimeKernel.resumeContinuation) {
  const continuation = await loadContinuationFromStore(...);
  for await (const chunk of runtimeKernel.resumeContinuation.call(
    runtimeKernel,
    continuation,
    { abortSignal: new AbortController().signal }
  )) {
    // Process resumed output chunks
  }
}

This pattern allows fine-grained control over timeout and cancellation policies during recovery.

Safety Guarantees and Error Handling

When resuming operations, Maka enforces strict validation to prevent execution of corrupted or unsupported continuations.

If a required resume candidate is missing or the model does not support the requested resume boundary, the system reports a clear error rather than attempting unsafe execution. In packages/ui/src/runtime-resume-copy.ts (lines 54-58), the source code defines specific error strings for these scenarios, ensuring users receive actionable feedback instead of opaque crash dumps.

The resumeTrust validation layer checks that all durable writes referenced by a continuation remain intact before replay begins. If the underlying storage reports mismatched checksums or missing commit records, the resume operation aborts immediately, forcing a clean restart from the last known good boundary.

Summary

Apache Maka's crash recovery architecture provides robust guarantees against data loss through:

  • Crash-boundary transactions that detect incomplete mutations via pre/post commit records in plan-store.ts
  • Safe-boundary continuations that validate durable write completion before replay via session-manager.ts
  • Quarantine logic that preserves workspace integrity using recovery bundles in git-workspace-service.ts
  • Explicit error handling that prevents unsafe resumes when candidates are missing or untrusted

Frequently Asked Questions

How does Maka detect incomplete transactions after a process crash?

Maka writes a durable commit record before executing state mutations and validates it afterward. If the validation marker is missing when the system restarts—as checked in packages/storage/src/plan-store.ts—the system identifies the transaction as incomplete and triggers rollback or retry protocols automatically.

What distinguishes a crash-boundary from a safe-boundary in Maka's architecture?

A crash-boundary wraps individual state-changing operations with pre/post validation to detect partial writes, while a safe-boundary marks points in execution where all required durable writes have succeeded and the system can safely capture continuations for later replay. Safe-boundaries depend on crash-boundaries completing successfully first.

Can Maka resume operations if the original session data is missing?

No. If the session header or continuation data cannot be located in packages/storage/src/session-store.ts, or if the resumeTrust validation fails, Maka explicitly refuses to resume. According to packages/ui/src/runtime-resume-copy.ts, the system raises a clear error indicating the missing resume candidate rather than attempting speculative reconstruction.

How does Maka prevent workspace corruption during crash recovery?

The quarantine mechanism in packages/storage/src/git-workspace-service.ts keeps staging inodes isolated from the working directory during crashes. Upon restart, Maka re-hydrates the workspace from a recovery bundle instead of attempting to repair partial staging state, ensuring no decoy files or orphaned data contaminate the production environment.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →