How Firstmate Handles Stuck or Unresponsive Crewmates: Recovery Playbook Explained
Firstmate treats crewmates as alive until proven otherwise, invoking the stuck-crewmate-recovery playbook to reconcile state and preserve unlanded work before escalating to interrupt or relaunch actions.
In the kunchenguid/firstmate repository, a "crewmate" represents the worker process that implements a task, and its reliability is critical to mission success. When session-start digests report dead endpoints, missing windows, or stale wakes, Firstmate executes a deterministic recovery workflow defined in .agents/skills/stuck-crewmate-recovery/SKILL.md rather than immediately terminating the worker.
The Stuck Crewmate Recovery Philosophy
Firstmate operates on a "alive until proven otherwise" premise, avoiding premature termination of potentially viable workers. The recovery mechanism is encapsulated in the stuck-crewmate-recovery skill, which is restricted to agent-only invocation to prevent accidental captain intervention. Throughout recovery, Firstmate never discards unlanded work, consistently reusing existing work-trees and only spawning new workers when safe state inheritance is guaranteed. The harness recorded in state/<id>.meta (specifically the harness= field) guides runtime-specific adjustments during reconciliation.
Three-Stage Recovery Workflow
The recovery process progresses through three distinct phases, each designed to minimize disruption while maximizing the chance of autonomous resolution.
Stage 1: Session-Start Reconciliation
When digest reports indicate potential issues with kind=ship or kind=scout workers, Firstmate initiates state inspection using bin/fm-crew-state.sh to load the worker's current runtime status. If the worker shows an active no-mistakes run, Firstmate allows the existing lifecycle to continue uninterrupted. Otherwise, it inspects the recorded backend and work-tree inventory, checking outputs like treehouse status for tmux/herdr/zellij sessions or orca_worktree_id= for Orca harnesses.
Stage 2: Live-Endpoint Escalation
If reconciliation confirms the worker is unresponsive, Firstmate follows a deterministic escalation ladder:
- Peek the pane to assess current activity without interrupting flow.
- Answer pending questions directly via
fm-send.shif the brief provides sufficient context. - Interrupt confused or looping workers using
fm-control.sh <id> interrupt. - Relaunch genuinely wedged workers with
fm-control.sh <id> relaunch, preserving the work-tree, uncommitted changes, original brief, and appending a concise progress note via--note.
The relaunch command accepts explicit runtime flags including --harness, --model, and --effort to adjust execution parameters. If a second relaunch fails, Firstmate records a failed status in the backlog and notifies the captain, strictly avoiding exposure of internal metadata in captain-facing messages.
Stage 3: Second-Mate Delegation
Workers with kind=secondmate follow a separate recovery path. Instead of invoking stuck-crewmate-recovery, Firstmate delegates to the secondmate-provisioning skill, as referenced in AGENTS.md (line 215), recognizing that second-mate workers require specialized provisioning logic distinct from standard ship or scout tasks.
Key Recovery Commands and Usage
The following commands implement the recovery workflow directly from the shell:
# Check current state of a potentially stuck task
FM_HOME=/path/to/firstmate-home bin/fm-crew-state.sh 1234abcd
# Interrupt a looping or confused worker
FM_HOME=/path/to/firstmate-home bin/fm-control.sh 1234abcd interrupt
# Answer a pending question directly
FM_HOME=/path/to/firstmate-home bin/fm-send.sh 1234abcd "Yes, the value is 42"
# Relaunch a wedged worker with preserved state and custom runtime
FM_HOME=/path/to/firstmate-home \
bin/fm-control.sh 1234abcd relaunch \
--note "Recovered after loop; continue from step 3" \
--harness="tmux" \
--model="gpt-4o" \
--effort="low"
Summary
- Firstmate assumes crewmates are alive until session-start digests or stale wakes prove otherwise.
- Recovery executes via the agent-only
stuck-crewmate-recoveryskill defined in.agents/skills/stuck-crewmate-recovery/SKILL.md. - State inspection uses
bin/fm-crew-state.shto check backend inventory without disrupting active no-mistakes runs. - Escalation follows a four-rung ladder: peek, answer via
fm-send.sh, interrupt viafm-control.sh, then relaunch with preserved work-trees and briefs. - Second-mate workers delegate to
secondmate-provisioningrather than the standard recovery playbook. - Failed recoveries after two relaunch attempts record
failedstatus and notify the captain without exposing internal metadata.
Frequently Asked Questions
What triggers the stuck crewmate recovery playbook in Firstmate?
Firstmate triggers recovery when session-start digests report that a direct-report worker's endpoint is dead, when worker metadata lacks a window, or when the worker generates a stale or blocked wake. These conditions signal potential unresponsiveness that requires agent intervention before captain notification.
How does Firstmate preserve work when relaunching a stuck worker?
During relaunch, Firstmate explicitly preserves the existing work-tree, uncommitted changes, and the original brief while appending a progress note via the --note flag. According to .agents/skills/harness-adapters/SKILL.md (line 57), the harness recorded in state/<id>.meta ensures runtime-specific state inheritance, preventing work loss even when changing harnesses or models via --harness, --model, or --effort flags.
What is the difference between interrupting and relaunching a crewmate?
Interrupting via fm-control.sh <id> interrupt sends a signal to break confused or looping execution without terminating the worker process, allowing recovery to continue within the same session. Relaunching via fm-control.sh <id> relaunch terminates the wedged worker and spawns a new process that inherits the preserved state, work-tree, and brief—necessary when the worker is genuinely frozen or unrecoverable.
Why are second-mate workers handled differently from ship or scout workers?
Second-mate workers (kind=secondmate) delegate recovery to the secondmate-provisioning skill rather than stuck-crewmate-recovery because they operate under distinct lifecycle constraints and provisioning requirements documented in AGENTS.md (line 215). This separation ensures that specialized second-mate initialization logic—potentially involving different authentication or resource allocation—remains isolated from standard crewmate recovery procedures.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →