How Firstmate Handles Stuck or Unresponsive Crewmates: Recovery Playbook Explained

Firstmate treats crewmates as alive until proven otherwise, invoking the stuck-crewmate-recovery playbook to reconcile state and preserve unlanded work before escalating to interrupt or relaunch actions.

In the kunchenguid/firstmate repository, a "crewmate" represents the worker process that implements a task, and its reliability is critical to mission success. When session-start digests report dead endpoints, missing windows, or stale wakes, Firstmate executes a deterministic recovery workflow defined in .agents/skills/stuck-crewmate-recovery/SKILL.md rather than immediately terminating the worker.

The Stuck Crewmate Recovery Philosophy

Firstmate operates on a "alive until proven otherwise" premise, avoiding premature termination of potentially viable workers. The recovery mechanism is encapsulated in the stuck-crewmate-recovery skill, which is restricted to agent-only invocation to prevent accidental captain intervention. Throughout recovery, Firstmate never discards unlanded work, consistently reusing existing work-trees and only spawning new workers when safe state inheritance is guaranteed. The harness recorded in state/<id>.meta (specifically the harness= field) guides runtime-specific adjustments during reconciliation.

Three-Stage Recovery Workflow

The recovery process progresses through three distinct phases, each designed to minimize disruption while maximizing the chance of autonomous resolution.

Stage 1: Session-Start Reconciliation

When digest reports indicate potential issues with kind=ship or kind=scout workers, Firstmate initiates state inspection using bin/fm-crew-state.sh to load the worker's current runtime status. If the worker shows an active no-mistakes run, Firstmate allows the existing lifecycle to continue uninterrupted. Otherwise, it inspects the recorded backend and work-tree inventory, checking outputs like treehouse status for tmux/herdr/zellij sessions or orca_worktree_id= for Orca harnesses.

Stage 2: Live-Endpoint Escalation

If reconciliation confirms the worker is unresponsive, Firstmate follows a deterministic escalation ladder:

  1. Peek the pane to assess current activity without interrupting flow.
  2. Answer pending questions directly via fm-send.sh if the brief provides sufficient context.
  3. Interrupt confused or looping workers using fm-control.sh <id> interrupt.
  4. Relaunch genuinely wedged workers with fm-control.sh <id> relaunch, preserving the work-tree, uncommitted changes, original brief, and appending a concise progress note via --note.

The relaunch command accepts explicit runtime flags including --harness, --model, and --effort to adjust execution parameters. If a second relaunch fails, Firstmate records a failed status in the backlog and notifies the captain, strictly avoiding exposure of internal metadata in captain-facing messages.

Stage 3: Second-Mate Delegation

Workers with kind=secondmate follow a separate recovery path. Instead of invoking stuck-crewmate-recovery, Firstmate delegates to the secondmate-provisioning skill, as referenced in AGENTS.md (line 215), recognizing that second-mate workers require specialized provisioning logic distinct from standard ship or scout tasks.

Key Recovery Commands and Usage

The following commands implement the recovery workflow directly from the shell:


# Check current state of a potentially stuck task

FM_HOME=/path/to/firstmate-home bin/fm-crew-state.sh 1234abcd

# Interrupt a looping or confused worker

FM_HOME=/path/to/firstmate-home bin/fm-control.sh 1234abcd interrupt

# Answer a pending question directly

FM_HOME=/path/to/firstmate-home bin/fm-send.sh 1234abcd "Yes, the value is 42"

# Relaunch a wedged worker with preserved state and custom runtime

FM_HOME=/path/to/firstmate-home \
  bin/fm-control.sh 1234abcd relaunch \
  --note "Recovered after loop; continue from step 3" \
  --harness="tmux" \
  --model="gpt-4o" \
  --effort="low"

Summary

  • Firstmate assumes crewmates are alive until session-start digests or stale wakes prove otherwise.
  • Recovery executes via the agent-only stuck-crewmate-recovery skill defined in .agents/skills/stuck-crewmate-recovery/SKILL.md.
  • State inspection uses bin/fm-crew-state.sh to check backend inventory without disrupting active no-mistakes runs.
  • Escalation follows a four-rung ladder: peek, answer via fm-send.sh, interrupt via fm-control.sh, then relaunch with preserved work-trees and briefs.
  • Second-mate workers delegate to secondmate-provisioning rather than the standard recovery playbook.
  • Failed recoveries after two relaunch attempts record failed status and notify the captain without exposing internal metadata.

Frequently Asked Questions

What triggers the stuck crewmate recovery playbook in Firstmate?

Firstmate triggers recovery when session-start digests report that a direct-report worker's endpoint is dead, when worker metadata lacks a window, or when the worker generates a stale or blocked wake. These conditions signal potential unresponsiveness that requires agent intervention before captain notification.

How does Firstmate preserve work when relaunching a stuck worker?

During relaunch, Firstmate explicitly preserves the existing work-tree, uncommitted changes, and the original brief while appending a progress note via the --note flag. According to .agents/skills/harness-adapters/SKILL.md (line 57), the harness recorded in state/<id>.meta ensures runtime-specific state inheritance, preventing work loss even when changing harnesses or models via --harness, --model, or --effort flags.

What is the difference between interrupting and relaunching a crewmate?

Interrupting via fm-control.sh <id> interrupt sends a signal to break confused or looping execution without terminating the worker process, allowing recovery to continue within the same session. Relaunching via fm-control.sh <id> relaunch terminates the wedged worker and spawns a new process that inherits the preserved state, work-tree, and brief—necessary when the worker is genuinely frozen or unrecoverable.

Why are second-mate workers handled differently from ship or scout workers?

Second-mate workers (kind=secondmate) delegate recovery to the secondmate-provisioning skill rather than stuck-crewmate-recovery because they operate under distinct lifecycle constraints and provisioning requirements documented in AGENTS.md (line 215). This separation ensures that specialized second-mate initialization logic—potentially involving different authentication or resource allocation—remains isolated from standard crewmate recovery procedures.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →