How to Handle Agent Recovery After an Unexpected Restart in SwarmForge

Agent recovery in SwarmForge is handled automatically by running ready_for_next.sh, which detects partially processed work, resumes it, or pulls the next item from the durable inbox queue without risking duplicate handoffs.

SwarmForge is a distributed multi-agent orchestration framework where each agent runs as a terminal pane with a designated role. When an agent process crashes or the host restarts, the system must guarantee exactly-once processing of handoffs. This article breaks down the architectural mechanisms that make this recovery seamless, with direct references to the source implementation in unclebob/swarm-forge.


The Recovery Entry Point: ready_for_next.sh

On restart, an agent executes the ready_for_next.sh script as its first action. According to the handoff protocol specification:

"On restart, an agent should run ready_for_next.sh and follow its output." — [handoff-protocol.md](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/handoff-protocol.md#L509)

This script serves as the dispatcher that determines which recovery logic to invoke based on the role's configuration. The script itself delegates to a Babashka implementation in ready_for_next.bb, which handles the low-level orchestration.


Role-Based Dispatch: Task Mode vs. Batch Mode

SwarmForge roles operate in one of two receive modes:

  • Task mode — processes one handoff at a time sequentially
  • Batch mode — groups multiple handoffs with identical priority for concurrent processing

The dispatcher reads the current role from SWARMFORGE_ROLE, queries the role's receive mode from .swarmforge/roles.tsv, and branches accordingly:

Receive Mode Helper Script Purpose
task ready_for_next_task.sh Resume or dequeue single handoffs
batch ready_for_next_batch.sh Resume or dequeue grouped handoffs

From the protocol documentation:

"Read the current role from the SWARMFORGE_ROLE environment variable. Read that role's receive mode. Dispatch to ready_for_next_task.sh if the mode is task. Dispatch to ready_for_next_batch.sh if the mode is batch." — [handoff-protocol.md](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/handoff-protocol.md#L67-L75)


Task Mode Recovery: The in_process Check

The ready_for_next_task.sh helper implements crash-safe recovery through a simple but powerful filesystem state machine. It operates on three directories:

  • inbox/new/ — handoffs waiting to be processed
  • inbox/in_process/ — handoff currently being worked on
  • inbox/completed/ — finished handoffs (archived)

The recovery logic follows this exact sequence:


# The agent (or its supervisor) runs:

$ ready_for_next.sh

# Possible outputs:

TASK: .swarmforge/handoffs/inbox/in_process/00_20260615T140531Z_000042_from_architect_to_coder.handoff

# Or if nothing is waiting:

NO_TASK

Behind the scenes, the Babashka implementation in ready_for_next_task.bb performs these steps:

  1. Check for existing in_process work — if a file exists, assume the previous run crashed mid-processing and resume it
  2. Otherwise, atomically move the oldest new handoff to in_process/ and add a dequeued_at timestamp
  3. Print the task details for the agent to consume

The protocol specifies this precisely:

"Check inbox/in_process/ first. If a file exists, report that it must be resumed. If no in-process file exists, select the first file in inbox/new/ and atomically move that file to inbox/in_process/." — [handoff-protocol.md](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/handoff-protocol.md#L90-L100)

This atomic move operation ensures that even if the agent crashes immediately after, the handoff remains in in_process/ and will be resumed on the next restart.


Batch Mode Recovery: Grouped Handoff Processing

For roles configured with batch receive mode, recovery works similarly but operates on grouped handoffs. The ready_for_next_batch.sh helper:

  1. Checks inbox/in_process/ for an existing batch directory
  2. If found, resumes that batch
  3. If not, creates a new batch directory with timestamp and suffix, moves all new/ handoffs sharing the same priority as the first file, and assigns the batch

From the protocol:

"If no in-process work exists, select every queued handoff with the same priority as the first new file. Move those files into one inbox/in_process/batch_<timestamp>_<suffix>/ directory." — [handoff-protocol.md](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/handoff-protocol.md#L44-L50)


Idempotency Guarantees

The SwarmForge recovery system is naturally idempotent because:

  • State is encoded entirely in immutable filesystem operations (atomic moves, not rewrites)
  • Timestamps and headers (dequeued_at, completed_at) are only appended, never modified
  • Duplicate detection happens automatically — an in_process file prevents re-dequeuing until explicitly completed

Re-running ready_for_next.sh multiple times on the same agent state produces the same output without side effects until the current work is finished.


Completing Work and Looping Back

After processing resumes or fresh work, the agent signals completion through done_with_current.sh:


# Agent finishes processing the resumed task:

$ done_with_current.sh

# Outputs indicate next action:

MAIL_WAITING   # More work available — loop back to ready_for_next.sh

NO_TASK        # Inbox empty — agent can idle or terminate

This helper dispatches to done_with_current_task.sh or done_with_current_batch.sh, which:

  • Add completed_at timestamp header
  • Move the handoff to the receiving role's outbox/ (or archive location)
  • Clean up the in_process state

The agent then loops back to ready_for_next.sh if MAIL_WAITING was returned, creating a continuous processing cycle that survives arbitrary crashes.


Code Example: Full Recovery Workflow

#!/bin/bash

# Example agent supervisor script that handles restart scenarios

while true; do
    # Step 1: Discover and resume or fetch work

    NEXT_TASK=$(ready_for_next.sh)
    
    if [[ "$NEXT_TASK" == "NO_TASK" ]]; then
        echo "No work available, sleeping..."
        sleep 30
        continue
    fi
    
    # Step 2: Parse the task or batch identifier

    if [[ "$NEXT_TASK" == TASK:* ]]; then
        HANDOFF_PATH="${NEXT_TASK#TASK: }"
        echo "Processing task: $HANDOFF_PATH"
        
        # Role-specific processing (example: coder merges and processes)

        merge_and_process.sh "$HANDOFF_PATH"
        
    elif [[ "$NEXT_TASK" == BATCH:* ]]; then
        BATCH_PATH="${NEXT_TASK#BATCH: }"
        echo "Processing batch: $BATCH_PATH"
        
        # Process all handoffs in batch directory

        for handoff in "$BATCH_PATH"/*.handoff; do
            merge_and_process.sh "$handoff"
        done
    fi
    
    # Step 3: Signal completion and check for more work

    STATUS=$(done_with_current.sh)
    echo "Completion status: $STATUS"
    
    if [[ "$STATUS" == "NO_TASK" ]]; then
        break  # Exit cleanly when no more work

    fi
    # Otherwise loop immediately to ready_for_next.sh

done

Key Implementation Files

File Purpose Location
ready_for_next.sh Main restart entry point and dispatcher [swarmforge/scripts/ready_for_next.sh](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/scripts/ready_for_next.sh)
ready_for_next.bb Babashka dispatcher implementation swarmforge/scripts/ready_for_next.bb
ready_for_next_task.sh / ready_for_next_task.bb Task-mode recovery and dequeuing [swarmforge/scripts/ready_for_next_task.sh](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/scripts/ready_for_next_task.sh)
ready_for_next_batch.sh / ready_for_next_batch.bb Batch-mode recovery and grouping [swarmforge/scripts/ready_for_next_batch.sh](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/scripts/ready_for_next_batch.sh)
done_with_current.sh Completion dispatcher [swarmforge/scripts/done_with_current.sh](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/scripts/done_with_current.sh)
handoff_lib.bb Core library for role lookup, file operations, and header management swarmforge/scripts/handoff_lib.bb
handoff-protocol.md Complete protocol specification with recovery semantics [swarmforge/handoff-protocol.md](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/handoff-protocol.md)

Summary

  • Recovery starts with ready_for_next.sh — this is the mandatory entry point after any restart
  • Role receive mode determines the recovery path — task mode for sequential processing, batch mode for grouped processing
  • The in_process/ directory acts as a crash journal — existing files indicate interrupted work that must be resumed
  • Atomic filesystem operations guarantee exactly-once semantics — no handoff can be duplicated or lost
  • The agent loops through ready_for_next.sh → process → done_with_current.sh until no work remains

Frequently Asked Questions

What happens if ready_for_next.sh is run while work is already in progress?

The helper detects the existing in_process/ file and reports it for resumption without modifying state. The protocol explicitly states: "Check inbox/in_process/ first. If a file exists, report that it must be resumed." This prevents duplicate processing and allows safe re-execution.

How does SwarmForge ensure handoffs aren't lost during a crash?

All state changes use atomic move operations and append-only headers. A handoff only leaves new/ when successfully moved to in_process/, and only leaves in_process/ when explicitly completed. Even if the agent terminates abruptly, the handoff remains in in_process/ and will be resumed on restart.

Can batch and task modes be mixed within the same role?

No — receive mode is a per-role configuration stored in .swarmforge/roles.tsv. Each role must declare task or batch mode. However, different roles in the same swarm can use different modes, allowing flexible pipeline design where some stages process sequentially and others process in parallel groups.

What triggers the automatic restart behavior in production deployments?

The README notes that "Recipient agents run ready_for_next.sh when notified or after restart." In practice, agent panes are typically supervised by process managers (systemd, supervisord, or container orchestrators) that restart the process on failure. The first command in the agent's startup sequence should always be ready_for_next.sh to ensure immediate recovery.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →