# How to Handle Agent Recovery After an Unexpected Restart in SwarmForge

> Learn how SwarmForge handles agent recovery after unexpected restarts. Discover the automated ready_for_next.sh script that resumes work or fetches new tasks securely.

- Repository: [Robert C. Martin/swarm-forge](https://github.com/unclebob/swarm-forge)
- Tags: how-to-guide
- Published: 2026-09-01

---

**Agent recovery in SwarmForge is handled automatically by running [`ready_for_next.sh`](https://github.com/unclebob/swarm-forge/blob/main/ready_for_next.sh), which detects partially processed work, resumes it, or pulls the next item from the durable inbox queue without risking duplicate handoffs.**

SwarmForge is a distributed multi-agent orchestration framework where each agent runs as a terminal pane with a designated **role**. When an agent process crashes or the host restarts, the system must guarantee exactly-once processing of handoffs. This article breaks down the architectural mechanisms that make this recovery seamless, with direct references to the source implementation in `unclebob/swarm-forge`.

---

## The Recovery Entry Point: ready_for_next.sh

On restart, an agent executes the [`ready_for_next.sh`](https://github.com/unclebob/swarm-forge/blob/main/ready_for_next.sh) script as its first action. According to the handoff protocol specification:

> "On restart, an agent should run [`ready_for_next.sh`](https://github.com/unclebob/swarm-forge/blob/main/ready_for_next.sh) and follow its output." — [[`handoff-protocol.md`](https://github.com/unclebob/swarm-forge/blob/main/handoff-protocol.md)](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/handoff-protocol.md#L509)

This script serves as the **dispatcher** that determines which recovery logic to invoke based on the role's configuration. The script itself delegates to a Babashka implementation in `ready_for_next.bb`, which handles the low-level orchestration.

---

## Role-Based Dispatch: Task Mode vs. Batch Mode

SwarmForge roles operate in one of two **receive modes**:

- **Task mode** — processes one handoff at a time sequentially
- **Batch mode** — groups multiple handoffs with identical priority for concurrent processing

The dispatcher reads the current role from `SWARMFORGE_ROLE`, queries the role's receive mode from `.swarmforge/roles.tsv`, and branches accordingly:

| Receive Mode | Helper Script | Purpose |
|-------------|-------------|---------|
| `task` | [`ready_for_next_task.sh`](https://github.com/unclebob/swarm-forge/blob/main/ready_for_next_task.sh) | Resume or dequeue single handoffs |
| `batch` | [`ready_for_next_batch.sh`](https://github.com/unclebob/swarm-forge/blob/main/ready_for_next_batch.sh) | Resume or dequeue grouped handoffs |

From the protocol documentation:

> "Read the current role from the `SWARMFORGE_ROLE` environment variable. Read that role's receive mode. Dispatch to [`ready_for_next_task.sh`](https://github.com/unclebob/swarm-forge/blob/main/ready_for_next_task.sh) if the mode is `task`. Dispatch to [`ready_for_next_batch.sh`](https://github.com/unclebob/swarm-forge/blob/main/ready_for_next_batch.sh) if the mode is `batch`." — [[`handoff-protocol.md`](https://github.com/unclebob/swarm-forge/blob/main/handoff-protocol.md)](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/handoff-protocol.md#L67-L75)

---

## Task Mode Recovery: The in_process Check

The [`ready_for_next_task.sh`](https://github.com/unclebob/swarm-forge/blob/main/ready_for_next_task.sh) helper implements **crash-safe recovery** through a simple but powerful filesystem state machine. It operates on three directories:

- `inbox/new/` — handoffs waiting to be processed
- `inbox/in_process/` — handoff currently being worked on
- `inbox/completed/` — finished handoffs (archived)

The recovery logic follows this exact sequence:

```bash

# The agent (or its supervisor) runs:

$ ready_for_next.sh

# Possible outputs:

TASK: .swarmforge/handoffs/inbox/in_process/00_20260615T140531Z_000042_from_architect_to_coder.handoff

# Or if nothing is waiting:

NO_TASK

```

Behind the scenes, the Babashka implementation in `ready_for_next_task.bb` performs these steps:

1. **Check for existing in_process work** — if a file exists, assume the previous run crashed mid-processing and **resume it**
2. **Otherwise, atomically move the oldest new handoff** to `in_process/` and add a `dequeued_at` timestamp
3. **Print the task details** for the agent to consume

The protocol specifies this precisely:

> "Check `inbox/in_process/` first. If a file exists, report that it must be resumed. If no in-process file exists, select the first file in `inbox/new/` and atomically move that file to `inbox/in_process/`." — [[`handoff-protocol.md`](https://github.com/unclebob/swarm-forge/blob/main/handoff-protocol.md)](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/handoff-protocol.md#L90-L100)

This atomic move operation ensures that even if the agent crashes immediately after, the handoff remains in `in_process/` and will be resumed on the next restart.

---

## Batch Mode Recovery: Grouped Handoff Processing

For roles configured with **batch receive mode**, recovery works similarly but operates on grouped handoffs. The [`ready_for_next_batch.sh`](https://github.com/unclebob/swarm-forge/blob/main/ready_for_next_batch.sh) helper:

1. Checks `inbox/in_process/` for an existing batch directory
2. If found, resumes that batch
3. If not, creates a new batch directory with timestamp and suffix, moves all `new/` handoffs sharing the same priority as the first file, and assigns the batch

From the protocol:

> "If no in-process work exists, select every queued handoff with the same priority as the first new file. Move those files into one `inbox/in_process/batch_<timestamp>_<suffix>/` directory." — [[`handoff-protocol.md`](https://github.com/unclebob/swarm-forge/blob/main/handoff-protocol.md)](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/handoff-protocol.md#L44-L50)

---

## Idempotency Guarantees

The SwarmForge recovery system is **naturally idempotent** because:

- State is encoded entirely in **immutable filesystem operations** (atomic moves, not rewrites)
- **Timestamps and headers** (`dequeued_at`, `completed_at`) are only appended, never modified
- **Duplicate detection** happens automatically — an `in_process` file prevents re-dequeuing until explicitly completed

Re-running [`ready_for_next.sh`](https://github.com/unclebob/swarm-forge/blob/main/ready_for_next.sh) multiple times on the same agent state produces the same output without side effects until the current work is finished.

---

## Completing Work and Looping Back

After processing resumes or fresh work, the agent signals completion through [`done_with_current.sh`](https://github.com/unclebob/swarm-forge/blob/main/done_with_current.sh):

```bash

# Agent finishes processing the resumed task:

$ done_with_current.sh

# Outputs indicate next action:

MAIL_WAITING   # More work available — loop back to ready_for_next.sh

NO_TASK        # Inbox empty — agent can idle or terminate

```

This helper dispatches to [`done_with_current_task.sh`](https://github.com/unclebob/swarm-forge/blob/main/done_with_current_task.sh) or [`done_with_current_batch.sh`](https://github.com/unclebob/swarm-forge/blob/main/done_with_current_batch.sh), which:

- Add `completed_at` timestamp header
- Move the handoff to the receiving role's `outbox/` (or archive location)
- Clean up the `in_process` state

The agent then **loops back** to [`ready_for_next.sh`](https://github.com/unclebob/swarm-forge/blob/main/ready_for_next.sh) if `MAIL_WAITING` was returned, creating a continuous processing cycle that survives arbitrary crashes.

---

## Code Example: Full Recovery Workflow

```bash
#!/bin/bash

# Example agent supervisor script that handles restart scenarios

while true; do
    # Step 1: Discover and resume or fetch work

    NEXT_TASK=$(ready_for_next.sh)
    
    if [[ "$NEXT_TASK" == "NO_TASK" ]]; then
        echo "No work available, sleeping..."
        sleep 30
        continue
    fi
    
    # Step 2: Parse the task or batch identifier

    if [[ "$NEXT_TASK" == TASK:* ]]; then
        HANDOFF_PATH="${NEXT_TASK#TASK: }"
        echo "Processing task: $HANDOFF_PATH"
        
        # Role-specific processing (example: coder merges and processes)

        merge_and_process.sh "$HANDOFF_PATH"
        
    elif [[ "$NEXT_TASK" == BATCH:* ]]; then
        BATCH_PATH="${NEXT_TASK#BATCH: }"
        echo "Processing batch: $BATCH_PATH"
        
        # Process all handoffs in batch directory

        for handoff in "$BATCH_PATH"/*.handoff; do
            merge_and_process.sh "$handoff"
        done
    fi
    
    # Step 3: Signal completion and check for more work

    STATUS=$(done_with_current.sh)
    echo "Completion status: $STATUS"
    
    if [[ "$STATUS" == "NO_TASK" ]]; then
        break  # Exit cleanly when no more work

    fi
    # Otherwise loop immediately to ready_for_next.sh

done

```

---

## Key Implementation Files

| File | Purpose | Location |
|------|---------|----------|
| [`ready_for_next.sh`](https://github.com/unclebob/swarm-forge/blob/main/ready_for_next.sh) | Main restart entry point and dispatcher | [[`swarmforge/scripts/ready_for_next.sh`](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/scripts/ready_for_next.sh)](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/scripts/ready_for_next.sh) |
| `ready_for_next.bb` | Babashka dispatcher implementation | [`swarmforge/scripts/ready_for_next.bb`](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/scripts/ready_for_next.bb) |
| [`ready_for_next_task.sh`](https://github.com/unclebob/swarm-forge/blob/main/ready_for_next_task.sh) / `ready_for_next_task.bb` | Task-mode recovery and dequeuing | [[`swarmforge/scripts/ready_for_next_task.sh`](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/scripts/ready_for_next_task.sh)](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/scripts/ready_for_next_task.sh) |
| [`ready_for_next_batch.sh`](https://github.com/unclebob/swarm-forge/blob/main/ready_for_next_batch.sh) / `ready_for_next_batch.bb` | Batch-mode recovery and grouping | [[`swarmforge/scripts/ready_for_next_batch.sh`](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/scripts/ready_for_next_batch.sh)](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/scripts/ready_for_next_batch.sh) |
| [`done_with_current.sh`](https://github.com/unclebob/swarm-forge/blob/main/done_with_current.sh) | Completion dispatcher | [[`swarmforge/scripts/done_with_current.sh`](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/scripts/done_with_current.sh)](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/scripts/done_with_current.sh) |
| `handoff_lib.bb` | Core library for role lookup, file operations, and header management | [`swarmforge/scripts/handoff_lib.bb`](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/scripts/handoff_lib.bb) |
| [`handoff-protocol.md`](https://github.com/unclebob/swarm-forge/blob/main/handoff-protocol.md) | Complete protocol specification with recovery semantics | [[`swarmforge/handoff-protocol.md`](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/handoff-protocol.md)](https://github.com/unclebob/swarm-forge/blob/main/swarmforge/handoff-protocol.md) |

---

## Summary

- **Recovery starts with [`ready_for_next.sh`](https://github.com/unclebob/swarm-forge/blob/main/ready_for_next.sh)** — this is the mandatory entry point after any restart
- **Role receive mode determines the recovery path** — task mode for sequential processing, batch mode for grouped processing
- **The `in_process/` directory acts as a crash journal** — existing files indicate interrupted work that must be resumed
- **Atomic filesystem operations guarantee exactly-once semantics** — no handoff can be duplicated or lost
- **The agent loops through [`ready_for_next.sh`](https://github.com/unclebob/swarm-forge/blob/main/ready_for_next.sh) → process → [`done_with_current.sh`](https://github.com/unclebob/swarm-forge/blob/main/done_with_current.sh)** until no work remains

---

## Frequently Asked Questions

### What happens if ready_for_next.sh is run while work is already in progress?

The helper detects the existing `in_process/` file and **reports it for resumption** without modifying state. The protocol explicitly states: "Check `inbox/in_process/` first. If a file exists, report that it must be resumed." This prevents duplicate processing and allows safe re-execution.

### How does SwarmForge ensure handoffs aren't lost during a crash?

All state changes use **atomic move operations** and **append-only headers**. A handoff only leaves `new/` when successfully moved to `in_process/`, and only leaves `in_process/` when explicitly completed. Even if the agent terminates abruptly, the handoff remains in `in_process/` and will be resumed on restart.

### Can batch and task modes be mixed within the same role?

No — receive mode is a **per-role configuration** stored in `.swarmforge/roles.tsv`. Each role must declare `task` or `batch` mode. However, different roles in the same swarm can use different modes, allowing flexible pipeline design where some stages process sequentially and others process in parallel groups.

### What triggers the automatic restart behavior in production deployments?

The README notes that "Recipient agents run [`ready_for_next.sh`](https://github.com/unclebob/swarm-forge/blob/main/ready_for_next.sh) when notified or after **restart**." In practice, agent panes are typically supervised by process managers (systemd, supervisord, or container orchestrators) that restart the process on failure. The first command in the agent's startup sequence should always be [`ready_for_next.sh`](https://github.com/unclebob/swarm-forge/blob/main/ready_for_next.sh) to ensure immediate recovery.