How Firstmate's Watcher Continuity and Stale Detection Works: A Deep Dive into the Shell Supervision Logic
Firstmate guarantees reliable worker supervision through a singleton watcher process that maintains continuity via filesystem locks and heartbeats, while detecting stale panes through double-hash verification and busy-state contracts.
Firstmate's architecture relies on a single-process watcher (bin/fm-watch.sh) to supervise every running worker, whether crewmate, secondmate, or external tool. This component implements watcher continuity through durable locks and heartbeat beacons, while its stale detection system identifies hung panes by comparing output hashes and inspecting busy-state contracts. Understanding these mechanisms is essential for debugging supervision failures or extending Firstmate's support for new backends.
Watcher Continuity Mechanisms
The watcher guarantees that only one healthy instance runs at any time, surviving crashes and preventing duplicate processes through explicit state management.
Singleton Lock Acquisition and Health Verification
When bin/fm-watch.sh starts, it attempts to acquire a filesystem lock by creating the directory state/.watch.lock and writing its PID, home directory, and path inside. The function fm_lock_try_acquire "$WATCH_LOCK" (line 27) handles this atomic acquisition. If the lock already exists, the new instance does not immediately exit; instead, it calls fm_watcher_healthy (defined in fm-wake-lib.sh lines 12‑29) to verify whether the existing lock holder is still alive.
This health check inspects three conditions: the PID recorded in the lock must exist in /proc or ps, the identity metadata must match the current filesystem context, and the heartbeat file must be recent. If any check fails, the new watcher considers the old lock stale and proceeds with takeover.
Heartbeat Beacon and Self-Eviction
Every loop iteration, the watcher updates state/.last-watcher-beat via touch "$STATE/.last-watcher-beat" (lines 40‑42). Other components, including bin/fm-guard.sh and the turn-end guard, monitor this file to determine if the watcher has become unresponsive.
The watcher also performs self-eviction checks. At lines 36‑38, it verifies that $(cat "$WATCH_LOCK/pid") still equals its own $WATCHER_PID. If a duplicate process has started and claimed the lock, the current instance detects the mismatch and exits immediately with exit 0, ensuring only the newest healthy process continues.
Recovery Handling After Downtime
When the lock contains a "downtime" marker indicating a previous crash, the watcher executes resurface_after_downtime (lines 63‑71). This routine re-queues any pending wakes that were lost during the outage, specifically processing check: rearm-resurface entries to ensure no signal is dropped during watcher replacement.
Stale Detection Logic
The watcher classifies wakes into categories—signal, stale, check, or heartbeat—and determines whether to absorb them or surface them as actionable items requiring LLM triage.
Signal Classification and Absorption
Signal files (matching *.status or *.turn-ended) are processed by scan_signals. A signal becomes an actionable wake only if:
- AFK mode is active (
state/.afkexists), or - The status contains a captain-relevant verb (
signal_reason_is_actionable), or - The crew is not provably working (
! signal_crew_provably_working).
If none of these conditions hold, the signal is absorbed: the watcher updates its .seen-* marker but does not append to the wake queue (lines 62‑78).
Pane-Hash Staleness Detection
For each window in recorded_windows, the watcher captures the last 40 bytes of pane output (tail40) and generates a hash using hash_pane. It compares this hash against the previous value stored in state/.hash-<key>.
If the hashes match twice consecutively (count ≥ 2), the pane becomes a stale candidate (lines 36‑41). This double-hash requirement prevents false positives from transient output pauses.
Provably-Working Override via Busy-State Contract
Even with matching hashes, the watcher checks the busy-state contract implemented in bin/fm-busy-lib.sh (lines 35‑36):
if window_is_busy "$w" "$tail40"; then busy_now=0; else busy_now=1; fi
- If
busy_now=0(busy), the stale detection is absorbed—the worker is deemed active despite static output. - If
busy_now=1(idle), the watcher proceeds to stale-handling paths.
Stale Handling Paths and Escalation
The watcher implements differentiated handling based on context:
Secondmate Pause State: If pause_state_class returns paused, the watcher calls handle_paused_stale (lines 42‑47), which records a long-cadence resurfacing timer rather than immediate escalation.
AFK Mode: When state/.afk exists, stale panes are enqueued directly via fm_wake_append stale (lines 48‑53).
Terminal Status Check: The function stale_is_terminal inspects whether the crew is provably working. If working, the stale is absorbed and a wedge timer starts; if not working, a stale wake emits immediately (lines 54‑71).
Wedge Escalation: After STALE_ESCALATE_SECS of repeated idle hashes, wedge_timer_check escalates the issue. Upon reaching FM_WEDGE_DEMAND_INSPECT_COUNT repetitions, the system adds a demand-deep-inspection marker to force LLM attention (lines 67‑87).
All paths ultimately surface via wake "stale: $w" to alert Firstmate's supervision loop.
Practical Examples
Starting the Watcher for Debugging
Launch the watcher manually to observe lock acquisition and heartbeat behavior:
# From the repository root
FM_HOME="$(pwd)" FM_STATE_OVERRIDE="${PWD}/state" bin/fm-watch.sh &
The script will acquire state/.watch.lock, begin writing to state/.last-watcher-beat, and start scanning recorded windows.
Simulating Stale Pane Detection
To manually trigger stale detection for testing:
# Assume window identifier stored in $WIN (e.g., tmux:0:mytask)
# Create a stable hash by capturing identical output twice:
tail -c 40 <(tmux capture-pane -t "$WIN") | md5sum > state/.hash-$(echo "$WIN" | tr ':/.' '___')
# Force the watcher to re-run its loop:
kill -USR1 $(cat state/.watch.lock/pid)
The watcher will detect the repeated hash, check the busy state via window_is_busy, and either absorb the signal or append stale: tmux:0:mytask to the queue.
Reading the Wake Queue
Inspect durable wake entries to verify stale detection results:
while read -r line; do
echo "Queued wake: $line"
done < state/.wake-queue
After detection, expect entries formatted as stale: <window_identifier>.
Summary
- Continuity relies on a singleton filesystem lock (
state/.watch.lock) combined with a heartbeat file (state/.last-watcher-beat) that other processes monitor to detect watcher death or duplication. - Stale detection uses double-hash verification of pane output (comparing
tail40hashes instate/.hash-*files) to identify static content that persists across two scanning cycles. - The busy-state contract (
window_is_busyinbin/fm/busy-lib.sh) provides a provably-working override that absorbs false positives from workers that appear idle but are actually processing. - Escalation paths include wedge timers (
STALE_ESCALATE_SECS) that eventually trigger demand-deep-inspection wakes, and special handling for paused secondmate windows viahandle_paused_stale.
Frequently Asked Questions
How does Firstmate prevent multiple watcher processes from running simultaneously?
Firstmate uses an atomic lock directory at state/.watch.lock combined with PID verification. When bin/fm-watch.sh starts, it calls fm_lock_try_acquire to claim the lock. If the lock exists, the new instance verifies the old PID through fm_watcher_healthy (checking /proc or ps, plus heartbeat recency). Only if the existing watcher is definitively dead does the new process proceed; otherwise, it exits. Additionally, running watchers perform self-eviction checks at lines 36‑38 to exit if another process steals the lock.
What determines whether a stale pane is absorbed versus surfaced as a wake?
A stale pane is absorbed when the busy-state contract reports the window as busy (window_is_busy returns true in bin/fm-busy-lib.sh), or when the crew is provably working during terminal status checks. The pane is surfaced as a stale wake when hashes match twice consecutively, the busy-state indicates idle, and either AFK mode is active or the crew is not provably working. In the latter case, fm_wake_append stale queues the notification.
Where does the watcher store its continuity and state markers?
The watcher maintains all state in the state/ directory relative to FM_HOME. Critical files include:
state/.watch.lock/(directory containingpid,home, andpathfiles for singleton enforcement)state/.last-watcher-beat(heartbeat timestamp)state/.hash-<window_key>(previous pane hashes for stale detection)state/.wake-queue(durable queue of actionable wakes)state/.afk(flag file indicating away-mode)
How does the wedge escalation mechanism work?
When a pane remains stale (idle hash repeated) beyond STALE_ESCALATE_SECS, the watcher increments an internal counter via wedge_timer_check. After FM_WEDGE_DEMAND_INSPECT_COUNT consecutive escalations, the system appends a demand-deep-inspection marker to the wake queue. This forces the LLM supervisor to perform a detailed analysis of the hung window, preventing indefinite silent stalling.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →