How to Handle Agent Recovery After an Unexpected Restart in SwarmForge
Agent recovery in SwarmForge is handled automatically by running ready_for_next.sh, which detects partially processed work, resumes it, or pulls the next item from the durable inbox queue without risking duplicate handoffs.
SwarmForge is a distributed multi-agent orchestration framework where each agent runs as a terminal pane with a designated role. When an agent process crashes or the host restarts, the system must guarantee exactly-once processing of handoffs. This article breaks down the architectural mechanisms that make this recovery seamless, with direct references to the source implementation in unclebob/swarm-forge.
The Recovery Entry Point: ready_for_next.sh
On restart, an agent executes the ready_for_next.sh script as its first action. According to the handoff protocol specification:
"On restart, an agent should run
ready_for_next.shand follow its output." — [handoff-protocol.md](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/handoff-protocol.md#L509)
This script serves as the dispatcher that determines which recovery logic to invoke based on the role's configuration. The script itself delegates to a Babashka implementation in ready_for_next.bb, which handles the low-level orchestration.
Role-Based Dispatch: Task Mode vs. Batch Mode
SwarmForge roles operate in one of two receive modes:
- Task mode — processes one handoff at a time sequentially
- Batch mode — groups multiple handoffs with identical priority for concurrent processing
The dispatcher reads the current role from SWARMFORGE_ROLE, queries the role's receive mode from .swarmforge/roles.tsv, and branches accordingly:
| Receive Mode | Helper Script | Purpose |
|---|---|---|
task |
ready_for_next_task.sh |
Resume or dequeue single handoffs |
batch |
ready_for_next_batch.sh |
Resume or dequeue grouped handoffs |
From the protocol documentation:
"Read the current role from the
SWARMFORGE_ROLEenvironment variable. Read that role's receive mode. Dispatch toready_for_next_task.shif the mode istask. Dispatch toready_for_next_batch.shif the mode isbatch." — [handoff-protocol.md](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/handoff-protocol.md#L67-L75)
Task Mode Recovery: The in_process Check
The ready_for_next_task.sh helper implements crash-safe recovery through a simple but powerful filesystem state machine. It operates on three directories:
inbox/new/— handoffs waiting to be processedinbox/in_process/— handoff currently being worked oninbox/completed/— finished handoffs (archived)
The recovery logic follows this exact sequence:
# The agent (or its supervisor) runs:
$ ready_for_next.sh
# Possible outputs:
TASK: .swarmforge/handoffs/inbox/in_process/00_20260615T140531Z_000042_from_architect_to_coder.handoff
# Or if nothing is waiting:
NO_TASK
Behind the scenes, the Babashka implementation in ready_for_next_task.bb performs these steps:
- Check for existing in_process work — if a file exists, assume the previous run crashed mid-processing and resume it
- Otherwise, atomically move the oldest new handoff to
in_process/and add adequeued_attimestamp - Print the task details for the agent to consume
The protocol specifies this precisely:
"Check
inbox/in_process/first. If a file exists, report that it must be resumed. If no in-process file exists, select the first file ininbox/new/and atomically move that file toinbox/in_process/." — [handoff-protocol.md](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/handoff-protocol.md#L90-L100)
This atomic move operation ensures that even if the agent crashes immediately after, the handoff remains in in_process/ and will be resumed on the next restart.
Batch Mode Recovery: Grouped Handoff Processing
For roles configured with batch receive mode, recovery works similarly but operates on grouped handoffs. The ready_for_next_batch.sh helper:
- Checks
inbox/in_process/for an existing batch directory - If found, resumes that batch
- If not, creates a new batch directory with timestamp and suffix, moves all
new/handoffs sharing the same priority as the first file, and assigns the batch
From the protocol:
"If no in-process work exists, select every queued handoff with the same priority as the first new file. Move those files into one
inbox/in_process/batch_<timestamp>_<suffix>/directory." — [handoff-protocol.md](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/handoff-protocol.md#L44-L50)
Idempotency Guarantees
The SwarmForge recovery system is naturally idempotent because:
- State is encoded entirely in immutable filesystem operations (atomic moves, not rewrites)
- Timestamps and headers (
dequeued_at,completed_at) are only appended, never modified - Duplicate detection happens automatically — an
in_processfile prevents re-dequeuing until explicitly completed
Re-running ready_for_next.sh multiple times on the same agent state produces the same output without side effects until the current work is finished.
Completing Work and Looping Back
After processing resumes or fresh work, the agent signals completion through done_with_current.sh:
# Agent finishes processing the resumed task:
$ done_with_current.sh
# Outputs indicate next action:
MAIL_WAITING # More work available — loop back to ready_for_next.sh
NO_TASK # Inbox empty — agent can idle or terminate
This helper dispatches to done_with_current_task.sh or done_with_current_batch.sh, which:
- Add
completed_attimestamp header - Move the handoff to the receiving role's
outbox/(or archive location) - Clean up the
in_processstate
The agent then loops back to ready_for_next.sh if MAIL_WAITING was returned, creating a continuous processing cycle that survives arbitrary crashes.
Code Example: Full Recovery Workflow
#!/bin/bash
# Example agent supervisor script that handles restart scenarios
while true; do
# Step 1: Discover and resume or fetch work
NEXT_TASK=$(ready_for_next.sh)
if [[ "$NEXT_TASK" == "NO_TASK" ]]; then
echo "No work available, sleeping..."
sleep 30
continue
fi
# Step 2: Parse the task or batch identifier
if [[ "$NEXT_TASK" == TASK:* ]]; then
HANDOFF_PATH="${NEXT_TASK#TASK: }"
echo "Processing task: $HANDOFF_PATH"
# Role-specific processing (example: coder merges and processes)
merge_and_process.sh "$HANDOFF_PATH"
elif [[ "$NEXT_TASK" == BATCH:* ]]; then
BATCH_PATH="${NEXT_TASK#BATCH: }"
echo "Processing batch: $BATCH_PATH"
# Process all handoffs in batch directory
for handoff in "$BATCH_PATH"/*.handoff; do
merge_and_process.sh "$handoff"
done
fi
# Step 3: Signal completion and check for more work
STATUS=$(done_with_current.sh)
echo "Completion status: $STATUS"
if [[ "$STATUS" == "NO_TASK" ]]; then
break # Exit cleanly when no more work
fi
# Otherwise loop immediately to ready_for_next.sh
done
Key Implementation Files
Summary
- Recovery starts with
ready_for_next.sh— this is the mandatory entry point after any restart - Role receive mode determines the recovery path — task mode for sequential processing, batch mode for grouped processing
- The
in_process/directory acts as a crash journal — existing files indicate interrupted work that must be resumed - Atomic filesystem operations guarantee exactly-once semantics — no handoff can be duplicated or lost
- The agent loops through
ready_for_next.sh→ process →done_with_current.shuntil no work remains
Frequently Asked Questions
What happens if ready_for_next.sh is run while work is already in progress?
The helper detects the existing in_process/ file and reports it for resumption without modifying state. The protocol explicitly states: "Check inbox/in_process/ first. If a file exists, report that it must be resumed." This prevents duplicate processing and allows safe re-execution.
How does SwarmForge ensure handoffs aren't lost during a crash?
All state changes use atomic move operations and append-only headers. A handoff only leaves new/ when successfully moved to in_process/, and only leaves in_process/ when explicitly completed. Even if the agent terminates abruptly, the handoff remains in in_process/ and will be resumed on restart.
Can batch and task modes be mixed within the same role?
No — receive mode is a per-role configuration stored in .swarmforge/roles.tsv. Each role must declare task or batch mode. However, different roles in the same swarm can use different modes, allowing flexible pipeline design where some stages process sequentially and others process in parallel groups.
What triggers the automatic restart behavior in production deployments?
The README notes that "Recipient agents run ready_for_next.sh when notified or after restart." In practice, agent panes are typically supervised by process managers (systemd, supervisord, or container orchestrators) that restart the process on failure. The first command in the agent's startup sequence should always be ready_for_next.sh to ensure immediate recovery.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →