# How Firstmate's Watcher Continuity and Stale Detection Works: A Deep Dive into the Shell Supervision Logic

> Discover how Firstmate's watcher ensures continuity and detects stale panes using filesystem locks, heartbeats, and busy-state contracts for robust worker supervision.

- Repository: [Kun Chen/firstmate](https://github.com/kunchenguid/firstmate)
- Tags: deep-dive
- Published: 2026-08-13

---

**Firstmate guarantees reliable worker supervision through a singleton watcher process that maintains continuity via filesystem locks and heartbeats, while detecting stale panes through double-hash verification and busy-state contracts.**

Firstmate's architecture relies on a **single-process watcher** ([`bin/fm-watch.sh`](https://github.com/kunchenguid/firstmate/blob/main/bin/fm-watch.sh)) to supervise every running worker, whether crewmate, secondmate, or external tool. This component implements **watcher continuity** through durable locks and heartbeat beacons, while its **stale detection** system identifies hung panes by comparing output hashes and inspecting busy-state contracts. Understanding these mechanisms is essential for debugging supervision failures or extending Firstmate's support for new backends.

## Watcher Continuity Mechanisms

The watcher guarantees that only one healthy instance runs at any time, surviving crashes and preventing duplicate processes through explicit state management.

### Singleton Lock Acquisition and Health Verification

When [`bin/fm-watch.sh`](https://github.com/kunchenguid/firstmate/blob/main/bin/fm-watch.sh) starts, it attempts to acquire a filesystem lock by creating the directory `state/.watch.lock` and writing its PID, home directory, and path inside. The function `fm_lock_try_acquire "$WATCH_LOCK"` (line 27) handles this atomic acquisition. If the lock already exists, the new instance does not immediately exit; instead, it calls `fm_watcher_healthy` (defined in [`fm-wake-lib.sh`](https://github.com/kunchenguid/firstmate/blob/main/fm-wake-lib.sh) lines 12‑29) to verify whether the existing lock holder is still alive.

This health check inspects three conditions: the PID recorded in the lock must exist in `/proc` or `ps`, the identity metadata must match the current filesystem context, and the heartbeat file must be recent. If any check fails, the new watcher considers the old lock stale and proceeds with takeover.

### Heartbeat Beacon and Self-Eviction

Every loop iteration, the watcher updates `state/.last-watcher-beat` via `touch "$STATE/.last-watcher-beat"` (lines 40‑42). Other components, including [`bin/fm-guard.sh`](https://github.com/kunchenguid/firstmate/blob/main/bin/fm-guard.sh) and the turn-end guard, monitor this file to determine if the watcher has become unresponsive.

The watcher also performs **self-eviction** checks. At lines 36‑38, it verifies that `$(cat "$WATCH_LOCK/pid")` still equals its own `$WATCHER_PID`. If a duplicate process has started and claimed the lock, the current instance detects the mismatch and exits immediately with `exit 0`, ensuring only the newest healthy process continues.

### Recovery Handling After Downtime

When the lock contains a "downtime" marker indicating a previous crash, the watcher executes `resurface_after_downtime` (lines 63‑71). This routine re-queues any pending wakes that were lost during the outage, specifically processing `check: rearm-resurface` entries to ensure no signal is dropped during watcher replacement.

## Stale Detection Logic

The watcher classifies wakes into categories—signal, stale, check, or heartbeat—and determines whether to absorb them or surface them as actionable items requiring LLM triage.

### Signal Classification and Absorption

Signal files (matching `*.status` or `*.turn-ended`) are processed by `scan_signals`. A signal becomes an actionable wake only if:

- **AFK mode** is active (`state/.afk` exists), **or**
- The status contains a **captain-relevant verb** (`signal_reason_is_actionable`), **or**
- The crew is **not provably working** (`! signal_crew_provably_working`).

If none of these conditions hold, the signal is **absorbed**: the watcher updates its `.seen-*` marker but does not append to the wake queue (lines 62‑78).

### Pane-Hash Staleness Detection

For each window in `recorded_windows`, the watcher captures the last 40 bytes of pane output (`tail40`) and generates a hash using `hash_pane`. It compares this hash against the previous value stored in `state/.hash-<key>`.

If the hashes match **twice consecutively** (`count ≥ 2`), the pane becomes a stale candidate (lines 36‑41). This double-hash requirement prevents false positives from transient output pauses.

### Provably-Working Override via Busy-State Contract

Even with matching hashes, the watcher checks the **busy-state contract** implemented in [`bin/fm-busy-lib.sh`](https://github.com/kunchenguid/firstmate/blob/main/bin/fm-busy-lib.sh) (lines 35‑36):

```bash
if window_is_busy "$w" "$tail40"; then busy_now=0; else busy_now=1; fi

```

- If `busy_now=0` (busy), the stale detection is **absorbed**—the worker is deemed active despite static output.
- If `busy_now=1` (idle), the watcher proceeds to stale-handling paths.

### Stale Handling Paths and Escalation

The watcher implements differentiated handling based on context:

**Secondmate Pause State:** If `pause_state_class` returns `paused`, the watcher calls `handle_paused_stale` (lines 42‑47), which records a long-cadence resurfacing timer rather than immediate escalation.

**AFK Mode:** When `state/.afk` exists, stale panes are enqueued directly via `fm_wake_append stale` (lines 48‑53).

**Terminal Status Check:** The function `stale_is_terminal` inspects whether the crew is provably working. If working, the stale is absorbed and a **wedge timer** starts; if not working, a stale wake emits immediately (lines 54‑71).

**Wedge Escalation:** After `STALE_ESCALATE_SECS` of repeated idle hashes, `wedge_timer_check` escalates the issue. Upon reaching `FM_WEDGE_DEMAND_INSPECT_COUNT` repetitions, the system adds a *demand-deep-inspection* marker to force LLM attention (lines 67‑87).

All paths ultimately surface via `wake "stale: $w"` to alert Firstmate's supervision loop.

## Practical Examples

### Starting the Watcher for Debugging

Launch the watcher manually to observe lock acquisition and heartbeat behavior:

```bash

# From the repository root

FM_HOME="$(pwd)" FM_STATE_OVERRIDE="${PWD}/state" bin/fm-watch.sh &

```

The script will acquire `state/.watch.lock`, begin writing to `state/.last-watcher-beat`, and start scanning recorded windows.

### Simulating Stale Pane Detection

To manually trigger stale detection for testing:

```bash

# Assume window identifier stored in $WIN (e.g., tmux:0:mytask)

# Create a stable hash by capturing identical output twice:

tail -c 40 <(tmux capture-pane -t "$WIN") | md5sum > state/.hash-$(echo "$WIN" | tr ':/.' '___')

# Force the watcher to re-run its loop:

kill -USR1 $(cat state/.watch.lock/pid)

```

The watcher will detect the repeated hash, check the busy state via `window_is_busy`, and either absorb the signal or append `stale: tmux:0:mytask` to the queue.

### Reading the Wake Queue

Inspect durable wake entries to verify stale detection results:

```bash
while read -r line; do
  echo "Queued wake: $line"
done < state/.wake-queue

```

After detection, expect entries formatted as `stale: <window_identifier>`.

## Summary

- **Continuity** relies on a **singleton filesystem lock** (`state/.watch.lock`) combined with a **heartbeat file** (`state/.last-watcher-beat`) that other processes monitor to detect watcher death or duplication.
- **Stale detection** uses **double-hash verification** of pane output (comparing `tail40` hashes in `state/.hash-*` files) to identify static content that persists across two scanning cycles.
- The **busy-state contract** (`window_is_busy` in [`bin/fm/busy-lib.sh`](https://github.com/kunchenguid/firstmate/blob/main/bin/fm/busy-lib.sh)) provides a provably-working override that absorbs false positives from workers that appear idle but are actually processing.
- **Escalation paths** include wedge timers (`STALE_ESCALATE_SECS`) that eventually trigger *demand-deep-inspection* wakes, and special handling for paused secondmate windows via `handle_paused_stale`.

## Frequently Asked Questions

### How does Firstmate prevent multiple watcher processes from running simultaneously?

Firstmate uses an atomic **lock directory** at `state/.watch.lock` combined with PID verification. When [`bin/fm-watch.sh`](https://github.com/kunchenguid/firstmate/blob/main/bin/fm-watch.sh) starts, it calls `fm_lock_try_acquire` to claim the lock. If the lock exists, the new instance verifies the old PID through `fm_watcher_healthy` (checking `/proc` or `ps`, plus heartbeat recency). Only if the existing watcher is definitively dead does the new process proceed; otherwise, it exits. Additionally, running watchers perform self-eviction checks at lines 36‑38 to exit if another process steals the lock.

### What determines whether a stale pane is absorbed versus surfaced as a wake?

A stale pane is **absorbed** when the **busy-state contract** reports the window as busy (`window_is_busy` returns true in [`bin/fm-busy-lib.sh`](https://github.com/kunchenguid/firstmate/blob/main/bin/fm-busy-lib.sh)), or when the crew is provably working during terminal status checks. The pane is **surfaced** as a stale wake when hashes match twice consecutively, the busy-state indicates idle, and either AFK mode is active or the crew is not provably working. In the latter case, `fm_wake_append stale` queues the notification.

### Where does the watcher store its continuity and state markers?

The watcher maintains all state in the `state/` directory relative to `FM_HOME`. Critical files include:
- `state/.watch.lock/` (directory containing `pid`, `home`, and `path` files for singleton enforcement)
- `state/.last-watcher-beat` (heartbeat timestamp)
- `state/.hash-<window_key>` (previous pane hashes for stale detection)
- `state/.wake-queue` (durable queue of actionable wakes)
- `state/.afk` (flag file indicating away-mode)

### How does the wedge escalation mechanism work?

When a pane remains stale (idle hash repeated) beyond `STALE_ESCALATE_SECS`, the watcher increments an internal counter via `wedge_timer_check`. After `FM_WEDGE_DEMAND_INSPECT_COUNT` consecutive escalations, the system appends a *demand-deep-inspection* marker to the wake queue. This forces the LLM supervisor to perform a detailed analysis of the hung window, preventing indefinite silent stalling.