Resume vs Retry in Apache Maka: Understanding the Recovery Process Difference
In Apache Maka, resume continues an interrupted session from a verified safe checkpoint while preserving model context and workspace snapshots, whereas retry launches a fresh attempt that only retains the task definition without restoring the previous execution state.
The Apache Maka project implements sophisticated recovery mechanisms for handling interrupted autonomous AI workflows. Understanding the difference between resume and retry in Maka's recovery process is critical for maintaining state consistency across long-running tasks. While both mechanisms address execution failures, they operate at fundamentally different architectural levels—resume performs continuation from exact snapshots while retry initiates new execution contexts.
Core Architectural Distinctions
Resume as State-Preserving Continuation
Resume operations in Maka function as true continuations from exact interruption points. According to the safe boundary contract defined in docs/architecture/runtime-resume-phase1-safe-boundary-contract.md, the runtime performs a repair → resolve/reconcile → resume sequence (lines 28-31) that ensures all model context remains intact. This includes tokens already consumed by the language model and workspace snapshots up to the interruption point.
The implementation requires a hard gate verification checking for parked status or recovery.hasCorruption before permitting continuation. This ensures that underlying corruption cannot propagate through the resume operation, maintaining the integrity of tool side-effects and completed operations.
Retry as Fresh Attempt Initialization
Retry operations represent distinct re-execution attempts rather than continuations. As documented in docs/archive/heavy-task-mainline-system-design.md (lines 111-119), timeout retry is not checkpoint resume—it launches a new attempt often with a fresh prompt projection without guaranteeing that the model sees the same hidden state. This mechanism targets transient failures, timeouts, or tool-level errors where re-execution provides better reliability than state restoration.
State Preservation Guarantees
Resume Retains Full Execution Context
When Maka executes a resume operation, it preserves:
- Model context including all tokens previously consumed
- Workspace snapshots up to the exact point of interruption
- Completed tool calls which are replayed or reconciled against the current state
The runtime-resume-architecture.md documentation (lines 69-73) specifies that this reconciliation occurs through explicit repair and resolve phases before execution continues.
Retry Initializes Clean State
Retry operations intentionally discard execution context:
- Only the task definition (prompt and parameters) persists from the failed attempt
- No guarantee of identical hidden state or workspace consistency with previous runs
- Previous attempts are marked as failed or retryable while the new attempt proceeds independently
This distinction appears explicitly in the UI implementation within packages/ui/src/tool-activity/copy.ts (lines 337-340), where retry actions are isolated from resume terminology through the "Switch and retry" affordance.
User Interface and CLI Implementation
Resume Commands and Safe Boundaries
Users trigger resume through specific CLI commands that enforce the safe boundary contract:
# Resume the most recent paused session
maka --resume
# Resume a specific legacy session ID
maka --resume <legacy-id>
The UI presents resume options only when a goal indicates paused status, displaying buttons with aria-label attributes defined by copy.resumeGoalAriaLabel as implemented in packages/ui/src/session-context-layer.tsx.
Retry Actions and Failure Handling
Retry mechanisms appear as reactive solutions to specific failure modes:
// In the tool-activity panel, the retry action is defined as:
<Button
label={copy.switchAndRetry}
onClick={handleSwitchAndRetry}
/>
This implementation in packages/ui/src/tool-activity/copy.ts demonstrates the deliberate UI separation between resuming paused workflows and retrying failed tool executions.
Practical Usage Examples
Recovering from System Interruptions
For sandbox crashes or system interruptions where maintaining exact state matters:
# Check session status and resume safely
maka status --session-id <id>
maka --resume <id>
Handling Transient Tool Failures
For timeout errors where fresh execution is preferable:
# After a task timeout, retry with fresh state
maka task retry <task-id>
Summary
- Resume continues execution from exact checkpoints via the repair → resolve/reconcile → resume sequence defined in
runtime-resume-architecture.md, preserving model tokens and workspace snapshots. - Retry launches new attempts that only retain task definitions, explicitly documented in
heavy-task-mainline-system-design.mdas distinct from checkpoint resume operations. - State guarantees differ fundamentally: resume reconciles previous tool calls while retry initializes clean contexts.
- UI separation exists intentionally, with safe resume buttons appearing only for paused goals while "Switch and retry" handles tool failures in
packages/ui/src/tool-activity/copy.ts. - Safety gates protect resume operations through corruption checks in
runtime-resume-phase1-safe-boundary-contract.md, whereas retry lacks such verification by design.
Frequently Asked Questions
When should I use resume instead of retry in Maka?
Use resume when recovering from crashes, sandbox errors, or intentionally paused sessions where preserving the exact conversation context and workspace state is essential. Resume operates only when the runtime has recorded a safe checkpoint and passed the hard gate verification for corruption. Use retry for transient failures, timeouts, or tool-level errors where starting fresh with a new attempt provides better reliability than restoring potentially problematic state.
Does retry preserve the conversation history in Maka?
No. Retry operations only retain the task definition including the original prompt and parameters. They do not guarantee that the model sees the same hidden state or that previously consumed tokens remain in context. According to heavy-task-mainline-system-design.md (lines 111-119), retry explicitly projects prior progress without restoring the exact in-flight model or workspace snapshot.
What happens to workspace files during a resume operation?
During resume, Maka reconciles the workspace snapshot up to the interruption point, replaying or reconciling completed tool calls against the current filesystem state. The runtime-resume-phase1-safe-boundary-contract.md specifies that this reconciliation occurs during the resolve phase before execution continues, ensuring file system consistency with the pre-interruption checkpoint.
How does Maka prevent corruption during the resume process?
The safe boundary contract implemented in runtime-resume-phase1-safe-boundary-contract.md (lines 28-31) enforces a hard gate that checks for parked status and recovery.hasCorruption flags before permitting any resume operation. This ensures that simple "click-to-continue" actions cannot mask underlying corruption, requiring explicit repair and reconciliation phases before state restoration occurs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →