How the Witness Agent Monitors Polecats Without Becoming a Bottleneck in Gastown
The Witness agent avoids bottlenecks by treating tmux as the single source of truth, reading push-based heartbeats only when needed, and performing selective asynchronous actions exclusively on problem sessions rather than polling all Polecats continuously.
The gastownhall/gastown project deploys the Witness agent as a lightweight watchdog to ensure Polecat sessions (Claude code agents) remain healthy without introducing coordination overhead. Unlike traditional monitoring systems that impose heavy polling loops or maintain complex state mirrors, the Witness relies on tmux session health checks and sparse file-based heartbeats to detect stuck or dead Polecats reactively.
Core Design Principles That Prevent Bottlenecks
Treating tmux as the Single Source of Truth
The Witness never maintains its own heavy state database. Instead, it delegates all liveness checks to tmux directly. In internal/witness/manager.go, the Manager.IsRunning() and Manager.IsHealthy() methods call tmux.CheckSessionHealth() to verify both the existence of the tmux window and the liveness of the underlying Claude process.
This approach is O(1) and lock-free because the system only queries tmux rather than writing to extra state files or maintaining internal mirrors. The CheckSessionHealth function in internal/tmux/tmux.go implements a three-state health model (healthy, agent-dead, session-dead) that the Witness consumes without modification.
Push-Based Heartbeat Architecture
Rather than continuously polling every Polecat, the Witness implements a push-based notification system. Polecats write a small JSON heartbeat file (read via polecat.ReadSessionHeartbeat) each time they change state. The Witness reads this file only when it scans a specific Polecat, avoiding tight loops.
According to the implementation in internal/witness/handlers.go (lines 1770-1790), the heartbeat carries a version field (IsV2) and a timestamp that lets the Witness determine if the session is fresh, stuck, or exiting. This file-based heartbeat is computationally cheap—a single read operation examined only when needed—ensuring the Witness never spins CPU cycles on idle agents.
Selective Asynchronous Actions
The Witness patrol logic in internal/witness/handlers.go employs aggressive filtering to minimize work:
- Idle Polecats (those with no work and a clean sandbox) are skipped entirely via a
continuestatement (lines 31-36, corresponding to lines 1712-1736 in the patrol loop). - Stuck or dead Polecats trigger a
RestartPolecatSessionrather than a destructive worktree teardown. - TOCTOU Protection prevents accidental double restarts. Before restarting, the code checks
if alive, _ := t.HasSession(sessionName); !alive { ... }(lines 17-22, corresponding to lines 1817-1822), ensuring a Polecat that exited microseconds ago is not needlessly restarted. - Escalations are sent asynchronously to the Refinery via
nudgeWitness, as seen ininternal/cmd/done.go(lines 1729-1734), where completed Polecats notify the Witness without blocking.
By acting only on problem cases and preferring safe restarts over destruction, the Witness eliminates unnecessary load on the system.
Config-Driven Thresholds
All timeouts—including DoneIntentStuckTimeoutD and HeartbeatStartupGraceD—are loaded from operational configuration via config.LoadOperationalConfig. In internal/witness/manager.go (lines 1650-1660), the witCfg object populates these thresholds at startup.
Dynamic configuration allows the system to adapt to slower machines or higher network latency without code changes, preventing false-positive restarts that would otherwise generate extra load and create cascading bottlenecks.
Periodic Polling Without Busy-Wait
The Witness patrol runs as a dedicated tmux pane (gt-witness-patrol formula) that wakes on a timer—defaulting to approximately 30 seconds—and processes the Polecat list once per cycle. The rest of the time, the pane remains idle.
This periodic, bounded scan guarantees a predictable CPU footprint regardless of how many Polecats are active, avoiding the resource exhaustion typical of busy-wait monitoring loops.
Key Implementation Files
| File | Role |
|---|---|
internal/witness/manager.go |
Life-cycle handling and tmux session health checks via IsRunning() and IsHealthy(). |
internal/witness/handlers.go |
Core patrol logic for idle detection, zombie detection, and restart/escalate actions. |
internal/tmux/tmux.go |
Low-level tmux wrapper with CheckSessionHealth() implementing the three-state health model. |
internal/cmd/done.go |
Polecat completion notifications via nudgeWitness() to alert the Witness asynchronously. |
Practical Code Examples
// Checking Witness session health with a 30-second inactivity threshold
m := witness.NewManager(rig) // rig is the current gas-town rig
healthy := m.IsHealthy(30 * time.Second)
fmt.Println("Witness healthy?", healthy)
// Restarting a stuck Polecat after TOCTOU verification
if zombie, found := detectZombieLiveSession(...); found {
// TOCTOU check already performed inside detectZombieLiveSession
log.Printf("Polecat %s restarted: %s", zombie.PolecatName, zombie.Action)
}
// Notifying Witness of completion without blocking
nudgeWitness(rigName, fmt.Sprintf("POLECAT_DONE %s exit=%s", polecatName, exitType))
Summary
- The Witness uses tmux as the source of truth, eliminating state synchronization overhead through O(1) health checks in
manager.go. - Push-based heartbeats replace continuous polling, with the Witness reading JSON heartbeat files only when scanning specific Polecats.
- Selective action ensures idle Polecats are skipped entirely, while stuck sessions trigger guarded restarts with TOCTOU protection in
handlers.go. - Configuration-driven thresholds allow runtime tuning of timeouts without recompilation, preventing false-positive bottlenecks.
- Periodic 30-second patrol cycles in a dedicated tmux pane ensure predictable resource usage without busy-waiting.
Frequently Asked Questions
What is a Polecat in the Gastown project?
A Polecat is a Claude code agent session managed within a tmux window. Each Polecat operates inside its own sandboxed environment (worktree) and performs development tasks such as coding, refactoring, or testing. The Witness monitors these sessions to ensure they remain responsive and healthy.
How does the Witness detect stuck Polecats without constant polling?
The Witness detects stuck Polecats by examining push-based heartbeat files that each Polecat updates when changing state. During its periodic patrol (approximately every 30 seconds), the Witness reads these files in handlers.go to check timestamps and version fields (IsV2). If a heartbeat exceeds the configured DoneIntentStuckTimeoutD threshold, the Polecat is flagged for restart.
What prevents the Witness from accidentally restarting active Polecats?
A Time-of-Check-Time-of-Use (TOCTOU) guard prevents accidental restarts. Before executing RestartPolecatSession, the code verifies the session is still missing by calling t.HasSession(sessionName) (lines 1817-1822 in handlers.go). If the session has already recovered or exited, the restart is skipped, eliminating race conditions that could interrupt active work.
How can operators tune Witness monitoring sensitivity?
Operators adjust sensitivity through the operational configuration loaded by config.LoadOperationalConfig in manager.go (lines 1650-1660). Key parameters include DoneIntentStuckTimeoutD for stuck session detection and HeartbeatStartupGraceD for startup timing allowances. These values can be modified at runtime without restarting the Witness binary, allowing immediate adaptation to different hardware or network conditions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →