How Runtime Continuation (Resume) in Apache Maka Reconstructs Interrupted Turns

Apache Maka's runtime continuation rebuilds interrupted turns by creating a new, safe execution from an immutable event ledger rather than reviving crashed processes.

The resume mechanism in Apache Maka solves a critical problem: when a model crashes mid-turn, how do you safely resume without replaying dangerous side effects or losing conversational context? The answer lies in a fact-based reconstruction pipeline that treats crashes as opportunities to build provably safe continuations from durable history.

The Core Philosophy: Facts Over Processes

Maka's resume system follows a strict principle: never revive dead processes. Instead of attempting process resurrection or code replay, the runtime treats the RuntimeEvent ledger as the single source of truth. Every decision—from tool call classification to safety validation—derives from this immutable record.

This design, implemented across packages/runtime/src/runtime-resume.ts and packages/runtime/src/runtime-continuation-planner.ts, ensures that resumed turns are provably safe before any model provider receives a single token.

The Eight-Phase Reconstruction Pipeline

Phase 1: Startup Repair

When the host (Desktop, CLI, or runtime-host) initializes, it scans for non-terminal AgentRun records and repairs their terminal state by reading committed RuntimeEvent entries.

This phase prepares the ground for potential resumption by ensuring all historical runs have consistent terminal boundaries.

Phase 2: Load Immutable History

The system loads all committed RuntimeEvents for the source run—the "high-water" prefix that survived any crash.

// packages/runtime/src/runtime-resume.ts (lines 19-24)
const events = await this.runtimeStore.readRuntimeEvents(
  sourceRunId,
  { upToHighWater: true }
);

This immutable prefix becomes the foundation for all subsequent decisions. No UI state, CLI buffer, or model self-report can override these facts.

Phase 3: Classify Tool Operations with RecoveryResolver

The RecoveryResolver examines events and classifies each tool call into one of five states:

State Meaning
completed Tool executed and result committed to ledger
definitely-not-dispatched Crash occurred before dispatch
indeterminate Dispatch occurred but completion status unclear
parked Explicitly held for human review
corrupt Event sequence violates invariants

Source: packages/runtime/src/recovery-resolver.ts via projectToolOperationsFromRuntimeEvents()

This classification isolates tool side effects from the model trace, preventing dangerous replays.

Phase 4: Gather Safety Observations

The host collects current environment state:

These observations feed into the continuation safety inspector (packages/runtime/src/continuation-safety.ts) before reaching the planner.

Phase 5: Safe-Boundary Verification

RuntimeContinuationPlanner.plan() validates critical invariants (lines 94-128 in runtime-resume.ts):

// packages/runtime/src/runtime-resume.ts
if (sourcePrefix.identity.sessionId !== input.sessionId ||
    sourcePrefix.identity.runId !== input.sourceRunId) {
  return parkedPlan('runtime_identity_mismatch',
    'RuntimeEvent ledger does not belong to the requested source run');
}

// Additional checks: workspace identity, tool availability,
// permission status, provider boundary alignment...

Failure modes trigger parking with diagnostic codes:

  • runtime_lineage_cycle – would create circular run dependencies
  • workspace_identity_mismatch – workspace state changed significantly
  • tool_unavailable – required tools no longer present
  • pending_permission – unresolved permission request from crashed run

Phase 6: Continuation Construction

When all invariants pass, buildSafeBoundaryContinuationPlan() (lines 1069-1125) creates fresh identity:

// packages/runtime/src/runtime-resume.ts
return {
  disposition: 'continue',
  rejectionReasons: [],
  continuation: {
    sessionId: source.sessionId,
    runId: deps.newId(),              // Fresh UUID
    invocationId: deps.newId(),       // Fresh UUID  
    turnId: deps.newId(),             // Fresh UUID
    sourceInvocationId: source.invocationId,
    sourceRunId: source.runId,
    sourceTurnId: source.turnId,
    sourceRuntimeEventHighWater: highWaterMark,
    runtimeContext: [...modelRuntimeContext],  // Replayable history
    safetySnapshot: {
      workspaceIdentity: facts.currentWorkspaceIdentity,
      backgroundOperationsSettled: true,
      availableToolNames: [...new Set(facts.availableToolNames)].sort(),
    },
  },
};

The continuation-start event records this transition permanently.

Phase 7: Kernel Execution and Provider Handoff

RuntimeKernel performs final validation before model contact:

// packages/runtime/src/runtime-kernel.ts
if (plan.disposition === 'continue') {
  const continuation = plan.continuation!;
  await this.runtimeStore.appendContinuationStart(continuation);
  await this.provider.sendHistory(continuation.runtimeContext);
}

Critical property: The provider receives clean history without the duplicate user message that initiated the crashed run. The continuation-start event ensures ledger consistency before any external API call.

Phase 8: Parking When Safety Cannot Be Proved

Any unprovable condition triggers parkedPlan() (lines 1261-1270):

// packages/runtime/src/runtime-resume.ts
function parkedPlan(
  reason: TurnResumeParkReason,
  message: string
): SafeBoundaryContinuationPlan {
  return {
    disposition: 'park',
    rejectionReasons: [{ code: reason, message }],
    continuation: undefined,
  };
}

Parked turns persist their facts and surface machine-readable diagnostics to UI/CLI layers—no automatic retries, no unsafe assumptions.

Triggering Resume from CLI

// packages/cli/src/runtime-host-session-driver.ts
import { SessionManager } from '@maka/runtime';

// User types "/resume" in TUI
const plan = await sessionManager.planLatestAuthoritativeSafeBoundaryContinuation(sessionId);

if (plan.disposition === 'continue') {
  await sessionManager.resumeSafeBoundaryContinuation(plan);
} else {
  // Display diagnostic: plan.rejectionReasons[0].code
  console.error(`Resume blocked: ${plan.rejectionReasons[0].message}`);
}

Key Architectural Components

Component Responsibility Location
RuntimeEvent Immutable ledger contract packages/core/src/runtime-event.ts
RecoveryResolver Tool operation classification packages/runtime/src/recovery-resolver.ts
RuntimeContinuationPlanner Safety validation and planning packages/runtime/src/runtime-continuation-planner.ts
ContinuationSafety Host-side environment inspection packages/runtime/src/continuation-safety.ts
RuntimeKernel Final validation and provider execution packages/runtime/src/runtime-kernel.ts

Summary

  • Resume creates new executions, never revives crashed processes
  • RuntimeEvent ledger is the sole source of reconstruction truth
  • RecoveryResolver isolates tool side effects through state classification
  • Continuation Planner enforces provable safety via invariant checking
  • Fresh identity allocation (runId, invocationId, turnId) prevents lineage corruption
  • Parking mechanism provides graceful degradation with actionable diagnostics

Frequently Asked Questions

What happens if workspace files changed after a crash?

The continuation planner detects this via workspace_identity_mismatch. The ContinuationSafety inspector (packages/runtime/src/continuation-safety.ts) captures a fingerprint of .maka-workspace.json and the planner compares it against the source run's recorded identity. Any deviation parks the turn with a diagnostic code, requiring human intervention.

Can a resumed turn create infinite loops through repeated crashes?

No. The planner explicitly checks for runtime_lineage_cycle—if the new run would create a circular dependency in the run ancestry graph, it returns a parked plan. Additionally, depth limits on continuation chains prevent unbounded nesting.

Why generate new IDs instead of reusing the crashed run's identity?

Fresh runId, invocationId, and turnId allocation (via deps.newId()) ensures immutable lineage tracking. The source identities are preserved as sourceRunId, sourceTurnId, etc., creating an audit trail while guaranteeing that each execution has a unique, non-conflicting identity in the ledger.

How does the provider avoid seeing duplicate user messages?

The runtimeContext constructed by buildSafeBoundaryContinuationPlan() contains only the replayable model events—assistant messages, tool results, and continuation metadata. The original user message that initiated the crashed run is excluded because the provider receives a clean continuation start. The continuation-start event in the ledger bridges this gap without exposing duplicates to the model API.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →