# How Firstmate Handles Stuck or Unresponsive Crewmates: Recovery Playbook Explained

> Firstmate's stuck-crewmate-recovery playbook ensures crewmate state is reconciled and unlanded work preserved before interrupting or relaunching actions. Learn how Firstmate keeps your operations running smoothly.

- Repository: [Kun Chen/firstmate](https://github.com/kunchenguid/firstmate)
- Tags: how-to-guide
- Published: 2026-08-13

---

**Firstmate treats crewmates as alive until proven otherwise, invoking the `stuck-crewmate-recovery` playbook to reconcile state and preserve unlanded work before escalating to interrupt or relaunch actions.**

In the `kunchenguid/firstmate` repository, a "crewmate" represents the worker process that implements a task, and its reliability is critical to mission success. When session-start digests report dead endpoints, missing windows, or stale wakes, Firstmate executes a deterministic recovery workflow defined in [`.agents/skills/stuck-crewmate-recovery/SKILL.md`](https://github.com/kunchenguid/firstmate/blob/main/.agents/skills/stuck-crewmate-recovery/SKILL.md) rather than immediately terminating the worker.

## The Stuck Crewmate Recovery Philosophy

Firstmate operates on a **"alive until proven otherwise"** premise, avoiding premature termination of potentially viable workers. The recovery mechanism is encapsulated in the **`stuck-crewmate-recovery`** skill, which is restricted to agent-only invocation to prevent accidental captain intervention. Throughout recovery, Firstmate **never discards unlanded work**, consistently reusing existing work-trees and only spawning new workers when safe state inheritance is guaranteed. The harness recorded in `state/<id>.meta` (specifically the `harness=` field) guides runtime-specific adjustments during reconciliation.

## Three-Stage Recovery Workflow

The recovery process progresses through three distinct phases, each designed to minimize disruption while maximizing the chance of autonomous resolution.

### Stage 1: Session-Start Reconciliation

When digest reports indicate potential issues with `kind=ship` or `kind=scout` workers, Firstmate initiates state inspection using **[`bin/fm-crew-state.sh`](https://github.com/kunchenguid/firstmate/blob/main/bin/fm-crew-state.sh)** to load the worker's current runtime status. If the worker shows an active no-mistakes run, Firstmate allows the existing lifecycle to continue uninterrupted. Otherwise, it inspects the recorded backend and work-tree inventory, checking outputs like `treehouse status` for tmux/herdr/zellij sessions or `orca_worktree_id=` for Orca harnesses.

### Stage 2: Live-Endpoint Escalation

If reconciliation confirms the worker is unresponsive, Firstmate follows a deterministic escalation ladder:

1. **Peek the pane** to assess current activity without interrupting flow.
2. **Answer pending questions** directly via **[`fm-send.sh`](https://github.com/kunchenguid/firstmate/blob/main/fm-send.sh)** if the brief provides sufficient context.
3. **Interrupt confused or looping workers** using **`fm-control.sh <id> interrupt`**.
4. **Relaunch genuinely wedged workers** with **`fm-control.sh <id> relaunch`**, preserving the work-tree, uncommitted changes, original brief, and appending a concise progress note via `--note`.

The relaunch command accepts explicit runtime flags including `--harness`, `--model`, and `--effort` to adjust execution parameters. If a second relaunch fails, Firstmate records a `failed` status in the backlog and notifies the captain, strictly avoiding exposure of internal metadata in captain-facing messages.

### Stage 3: Second-Mate Delegation

Workers with `kind=secondmate` follow a separate recovery path. Instead of invoking `stuck-crewmate-recovery`, Firstmate delegates to the **`secondmate-provisioning`** skill, as referenced in [`AGENTS.md`](https://github.com/kunchenguid/firstmate/blob/main/AGENTS.md) (line 215), recognizing that second-mate workers require specialized provisioning logic distinct from standard ship or scout tasks.

## Key Recovery Commands and Usage

The following commands implement the recovery workflow directly from the shell:

```bash

# Check current state of a potentially stuck task

FM_HOME=/path/to/firstmate-home bin/fm-crew-state.sh 1234abcd

# Interrupt a looping or confused worker

FM_HOME=/path/to/firstmate-home bin/fm-control.sh 1234abcd interrupt

# Answer a pending question directly

FM_HOME=/path/to/firstmate-home bin/fm-send.sh 1234abcd "Yes, the value is 42"

# Relaunch a wedged worker with preserved state and custom runtime

FM_HOME=/path/to/firstmate-home \
  bin/fm-control.sh 1234abcd relaunch \
  --note "Recovered after loop; continue from step 3" \
  --harness="tmux" \
  --model="gpt-4o" \
  --effort="low"

```

## Summary

- Firstmate assumes crewmates are alive until session-start digests or stale wakes prove otherwise.
- Recovery executes via the agent-only `stuck-crewmate-recovery` skill defined in [`.agents/skills/stuck-crewmate-recovery/SKILL.md`](https://github.com/kunchenguid/firstmate/blob/main/.agents/skills/stuck-crewmate-recovery/SKILL.md).
- **State inspection** uses [`bin/fm-crew-state.sh`](https://github.com/kunchenguid/firstmate/blob/main/bin/fm-crew-state.sh) to check backend inventory without disrupting active no-mistakes runs.
- **Escalation** follows a four-rung ladder: peek, answer via [`fm-send.sh`](https://github.com/kunchenguid/firstmate/blob/main/fm-send.sh), interrupt via [`fm-control.sh`](https://github.com/kunchenguid/firstmate/blob/main/fm-control.sh), then relaunch with preserved work-trees and briefs.
- **Second-mate workers** delegate to `secondmate-provisioning` rather than the standard recovery playbook.
- Failed recoveries after two relaunch attempts record `failed` status and notify the captain without exposing internal metadata.

## Frequently Asked Questions

### What triggers the stuck crewmate recovery playbook in Firstmate?

Firstmate triggers recovery when session-start digests report that a direct-report worker's endpoint is dead, when worker metadata lacks a window, or when the worker generates a stale or blocked wake. These conditions signal potential unresponsiveness that requires agent intervention before captain notification.

### How does Firstmate preserve work when relaunching a stuck worker?

During relaunch, Firstmate explicitly preserves the existing work-tree, uncommitted changes, and the original brief while appending a progress note via the `--note` flag. According to [`.agents/skills/harness-adapters/SKILL.md`](https://github.com/kunchenguid/firstmate/blob/main/.agents/skills/harness-adapters/SKILL.md) (line 57), the harness recorded in `state/<id>.meta` ensures runtime-specific state inheritance, preventing work loss even when changing harnesses or models via `--harness`, `--model`, or `--effort` flags.

### What is the difference between interrupting and relaunching a crewmate?

**Interrupting** via `fm-control.sh <id> interrupt` sends a signal to break confused or looping execution without terminating the worker process, allowing recovery to continue within the same session. **Relaunching** via `fm-control.sh <id> relaunch` terminates the wedged worker and spawns a new process that inherits the preserved state, work-tree, and brief—necessary when the worker is genuinely frozen or unrecoverable.

### Why are second-mate workers handled differently from ship or scout workers?

Second-mate workers (`kind=secondmate`) delegate recovery to the `secondmate-provisioning` skill rather than `stuck-crewmate-recovery` because they operate under distinct lifecycle constraints and provisioning requirements documented in [`AGENTS.md`](https://github.com/kunchenguid/firstmate/blob/main/AGENTS.md) (line 215). This separation ensures that specialized second-mate initialization logic—potentially involving different authentication or resource allocation—remains isolated from standard crewmate recovery procedures.