Troubleshooting Agent Wake-Up Notifications Not Reaching tmux in SwarmForge
If agent wake-up notifications aren't appearing in tmux sessions in SwarmForge, the issue typically stems from socket path mismatches, incorrect session names, or the handoff daemon failing to execute its notify! function.
SwarmForge's agent coordination relies on a hand-off daemon (handoffd.bb) to bridge agents and tmux. Agents never communicate with tmux directly—they drop hand-off files into their outbox, and the daemon delivers wake-up messages when those files are processed. When these notifications fail to reach tmux panes, the root cause usually lies in the daemon's connection to the tmux server or its configuration.
How Wake-Up Notifications Work in SwarmForge
Understanding the notification flow is essential for effective troubleshooting. The daemon follows a strict three-step process defined in swarmforge/scripts/handoffd.bb.
The Wake-Up Message Definition
The daemon hardcodes the notification text at line 10:
(def wake-message "You have new handoff mail. If idle, run ready_for_next.sh.")
This generic wake-up message is intentionally identical for all recipients—the protocol does not include message content from the hand-off itself.
The notify! Function Implementation
The notify! function (lines 13-19 in handoffd.bb) constructs three sequential tmux commands:
send-keys -lto transmit the wake-message textsend-keys C-mto send a carriage-returnsend-keys C-jto send a line-feed
Each command's exit code is checked, and any failure throws an exception that gets logged:
(defn notify! [socket session]
(let [send-text (sh "tmux" "-S" socket "send-keys" "-t" session "-l" wake-message)
...] ...)
When notify! Is Called
The daemon triggers wake-ups in two scenarios. After copying a hand-off to each recipient's inbox, it loops through recipients and calls notify! for each role's session (lines 60-68):
(doseq [recipient recipients]
... (notify! socket (:session role-info)))
Additionally, maybe-notify-unblocked-sender! (lines 43-48) can wake the original sender once its work clears:
(when (and (approved-git-handoff? headers)
(sender-ready-work? roles sender-role)
(not (contains? (set (recipient-list headers)) sender-role)))
(notify! socket (get-in roles [sender-role :session])))
Common Causes of Missing Wake-Up Notifications
Socket Path Problems
The daemon reads the tmux socket path from .swarmforge/tmux-socket during configure! (line 44):
(alter-var-root #'socket-file (constantly (fs/path state "tmux-socket")))
If this file is missing, unreadable, or contains a stale path, notify! fails when attempting tmux commands.
Symptoms: No message appears in any tmux pane; handoffd.log shows "tmux send text failed."
Verification:
cat .swarmforge/tmux-socket # displays the socket path
ls -l $(cat .swarmforge/tmux-socket) # confirms socket exists with correct permissions
Session Name Mismatches
Each role's session name is stored in .swarmforge/roles.tsv and loaded by load-roles around line 64. If this session doesn't exist in the tmux server, notify! executes but tmux silently fails.
Symptoms: Only specific agents miss wake-ups while others receive them normally.
Verification:
tmux -S $(cat .swarmforge/tmux-socket) list-sessions
# Compare output against the session column in .swarmforge/roles.tsv
Incomplete notify! Execution
The notify! function sends three separate tmux commands. If the text delivery succeeds but the carriage-return or line-feed fails, the notification aborts partially.
Symptoms: Message text may flash briefly then disappear, or appear without being submitted.
Verification: Check handoffd.log for "tmux send carriage return failed" or "tmux send line feed failed."
Daemon-Tmux Startup Ordering
If tmux restarts after the daemon initializes, the socket file is recreated and the daemon retains the old path.
Symptoms: Wake-ups work immediately after restarting tmux, then fail permanently until daemon restart.
Fix: Always restart the daemon after tmux:
pkill -f handoffd.bb
handoffd.bb <project-root>
Daemon Not Processing Outbox
The daemon may be running but not actively processing hand-offs due to the once? flag, a stop file, or an event loop failure.
Symptoms: Files accumulate in outbox/ without triggering notifications.
Verification:
tail -f .swarmforge/daemon/handoffd.log # confirm "started" and looping behavior
ls .swarmforge/daemon/stop # should not exist
ls .swarmforge/handoffs/outbox/tmp/ # stuck files here are ignored
Step-by-Step Diagnostic Procedure
Follow this sequence to isolate wake-up failures:
-
Check daemon health
tail -f .swarmforge/daemon/handoffd.logLook for "delivered" entries and any "error" or "tmux send … failed" messages.
-
Validate the tmux socket
SOCKET=$(cat .swarmforge/tmux-socket) tmux -S "$SOCKET" list-sessions -
Verify role session configuration
awk -F'\t' '$1=="coder"{print $4}' .swarmforge/roles.tsv -
Manually test tmux commands
SOCKET=$(cat .swarmforge/tmux-socket) SESSION=worker-coder tmux -S "$SOCKET" send-keys -t "$SESSION" -l "Test wake-up" tmux -S "$SOCKET" send-keys -t "$SESSION" C-m tmux -S "$SOCKET" send-keys -t "$SESSION" C-jSuccess here indicates the daemon never reaches
notify!; failure indicates tmux connectivity issues. -
Inspect outbox state
ls .swarmforge/handoffs/outbox/*.handoff -
Restart the daemon after any configuration correction
pkill -f handoffd.bb handoffd.bb <project-root>
The Intentionally Lossy Design
According to the hand-off protocol documented in swarmforge/handoff-protocol.md, tmux wake-ups are intentionally lossy. The wake-up is merely a prompt for idle agents to run ready_for_next.sh and check their durable inbox. Busy agents safely ignore missed notifications without data loss, and the daemon does not track delivery acknowledgments. This design decouples agents from tmux availability and maintains the file-based queue as the authoritative source of truth.
Key Files for Wake-Up Troubleshooting
| File | Purpose |
|---|---|
swarmforge/scripts/handoffd.bb |
Core daemon with notify! implementation and wake-message definition |
swarmforge/handoff-protocol.md |
Protocol specification explaining lossy wake-up semantics |
swarmforge/scripts/ready_for_next.sh |
Agent entry point triggered by wake-ups |
.swarmforge/roles.tsv |
Role-to-session mapping loaded by daemon |
.swarmforge/daemon/handoffd.log |
Runtime error log including tmux command failures |
.swarmforge/tmux-socket |
Dynamic path to the active tmux socket |
Summary
- Agents never call tmux directly—only
handoffd.bbsends wake-ups vianotify! - Three tmux commands execute per wake-up: text, carriage-return, and line-feed; any failure aborts the notification
- Socket path and session name are the critical configuration points—verify both when wake-ups fail
- Wake-ups are intentionally unreliable by design—agents must poll their inbox if they miss a notification
- Always restart the daemon after tmux restarts to refresh the socket path
Frequently Asked Questions
Why do my agents sometimes miss wake-up notifications even when everything is configured correctly?
The SwarmForge hand-off protocol intentionally treats wake-ups as lossy hints rather than guaranteed delivery mechanisms. According to swarmforge/handoff-protocol.md, this design prevents coupling agents to tmux availability and allows the file-based inbox to remain the source of truth. Busy agents may safely ignore wake-ups, and the daemon does not retry failed notifications. Agents should run ready_for_next.sh periodically when idle to poll for work.
How can I manually send a wake-up notification to test my tmux setup?
Execute the three commands that notify! uses internally, substituting your socket path and session name:
SOCKET=$(cat .swarmforge/tmux-socket)
SESSION=your-role-session
tmux -S "$SOCKET" send-keys -t "$SESSION" -l "You have new handoff mail. If idle, run ready_for_next.sh."
tmux -S "$SOCKET" send-keys -t "$SESSION" C-m
tmux -S "$SOCKET" send-keys -t "$SESSION" C-j
If the message appears and submits in the target pane, your tmux configuration is correct and the issue lies upstream in the daemon.
Where does the daemon log tmux command failures?
All notify! errors are written to .swarmforge/daemon/handoffd.log via the log! function. Search for strings containing "tmux send" to find specific failure points—each of the three command types (text, carriage-return, line-feed) generates distinct error messages when they fail.
What happens if the tmux socket path changes while the daemon is running?
The daemon reads the socket path once during configure! (line 44 in handoffd.bb) and stores it in the socket-file var. If tmux restarts and recreates its socket file, the daemon retains the old path and all subsequent notify! calls fail. Always restart handoffd.bb after restarting tmux to refresh the socket reference.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →