How the Swarm Forge Agent Queue Loop Works: File-Driven Orchestration Explained
Swarm Forge processes AI agent work through a hand-off queue loop implemented in Babashka scripts that dequeue tasks from .swarmforge/handoffs/inbox, execute them atomically, and update the dashboard via timestamp-driven state transitions.
The agent queue loop in Swarm Forge is a deterministic, file-driven state machine that coordinates work between specialized AI roles. Unlike traditional message brokers, Swarm Forge uses simple filesystem operations and shell scripts to guarantee exactly-once processing and full visibility into agent state.
The Core Loop: ready_for_next as Entry Point
When an agent's tmux pane becomes idle, ready_for_next.sh triggers the queue loop. This thin shell wrapper delegates to ready_for_next.bb, the central dispatcher written in Babashka.
In swarmforge/scripts/ready_for_next.bb, the script determines the current role and queries its receive mode:
(case mode
"batch" (run-helper! "ready_for_next_batch.sh")
"task" (run-helper! "ready_for_next_task.sh")
(exit! 2 (str "INVALID_RECEIVE_MODE: " mode " for role " role-name)))
The receive mode controls throughput: "task" processes one hand-off per invocation, while "batch" exhausts the entire inbox before returning control to the dashboard.
Task Mode: Single Hand-Off Processing
Task mode (ready_for_next_task.sh) is the default for interactive roles that need frequent dashboard updates. The script relies on reusable functions in handoff_lib.bb to manipulate hand-off files.
The dequeue sequence in swarmforge/scripts/handoff_lib.bb:
- List pending files in
.swarmforge/handoffs/inbox - Sort by filesystem modification time (oldest first)
- Move the selected file to
outboxwith adequeued_attimestamp - Execute the payload via
merge_and_process.sh - Complete by writing
completed_atand moving todone
;; handoff_lib.bb – dequeue one hand-off
(defn dequeue-next! [inbox-dir]
(let [sorted (sort-by fs/mod-time (fs/glob inbox-dir "*.handoff"))
next (first sorted)]
(when next
(fs/move next ".swarmforge/handoffs/outbox")
(add-header! next "dequeued_at" (now))
next)))
# Manually trigger single-task processing
$ ./swarmforge/scripts/ready_for_next.sh
Batch Mode: Throughput-Optimized Processing
Batch mode (ready_for_next_batch.sh) wraps the task logic in a loop that continues until the inbox empties. This mode suits background roles like cleaner or tester that benefit from minimizing context switches.
The batch loop structure:
# Pseudo-code representation of ready_for_next_batch.sh
while inbox_not_empty; do
ready_for_next_task.sh
done
Batch mode reduces dashboard noise but delays visibility into individual task completion—trade-offs the role configuration controls.
Finalization: merge_and_process.sh
Both modes converge on swarmforge/scripts/merge_and_process.sh for work execution. This script performs three critical operations:
- Git synchronization: Rebases the agent's worktree onto latest
mainto prevent merge conflicts - Role invocation: Executes the agent-specific script (e.g.,
coder.sh,cleaner.sh) - Dashboard refresh: Updates UI state so users see "waiting" → "in-process" → "completed" transitions
File-Based State Machine Guarantees
The agent queue loop achieves reliability through timestamp fields embedded in hand-off files:
| Field | Purpose |
|---|---|
enqueued_at |
Records when the dashboard created the hand-off |
dequeued_at |
Set when a role claims the task (prevents double-processing) |
completed_at |
Written after merge_and_process.sh finishes successfully |
This design ensures atomicity (timestamp presence indicates state), serialization (only one agent accesses its inbox at a time via tmux pane isolation), and visibility (dashboard reads the same files without polling a database).
Trigger Sources and Watchdog Behavior
The queue loop activates from multiple sources:
- Tmux watchdog: Automatic invocation when an agent pane detects idleness
- Manual trigger: Direct execution of
ready_for_next.shfor debugging - Test harness: The
queue-handoff!helper in tests simulates dashboard injection
# Example: Inspect current queue state
$ ls -la .swarmforge/handoffs/inbox/*.handoff 2>/dev/null | wc -l
Key Files in the Agent Queue Loop
| File | Responsibility |
|---|---|
swarmforge/scripts/ready_for_next.bb |
Role detection and mode dispatch |
swarmforge/scripts/ready_for_next_task.sh |
Single hand-off dequeue and execution |
swarmforge/scripts/ready_for_next_batch.sh |
Exhaustive inbox processing |
swarmforge/scripts/handoff_lib.bb |
File operations and timestamp management |
swarmforge/scripts/merge_and_process.sh |
Git merge, role script invocation, UI update |
swarmforge/scripts/pack_web.sh |
Dashboard entry point for hand-off creation |
Summary
- The agent queue loop centers on
ready_for_next.bbdispatching to task or batch processors based on role configuration - Task mode processes one hand-off per invocation; batch mode empties the inbox before returning
handoff_lib.bbprovides atomic dequeue via filesystem moves and timestamp headersmerge_and_process.shfinalizes work with Git operations and role-specific execution- Filesystem-based state (
enqueued_at,dequeued_at,completed_at) eliminates external dependencies and enables full observability
Frequently Asked Questions
How does Swarm Forge prevent the same hand-off from being processed twice?
The dequeued_at timestamp acts as a claim marker. When handoff_lib.bb moves a file from inbox to outbox, it atomically writes this timestamp. Any concurrent attempt to process the same file fails because it no longer exists in the source directory, and the timestamp presence in outbox indicates active processing.
Can I change a role from task mode to batch mode without restarting Swarm Forge?
Yes. The receive mode is resolved dynamically in ready_for_next.bb via handoff-lib/role-receive-mode. Modify the role configuration and the next invocation of ready_for_next.sh will respect the new setting. No daemon restart is required because state lives in files, not memory.
What happens if merge_and_process.sh fails mid-execution?
The hand-off remains in outbox without a completed_at timestamp. The dashboard displays this as "in-process" until manual intervention or timeout. Because the file left inbox with dequeued_at set, it won't be re-processed automatically—preventing duplicate work but requiring operator attention for recovery.
Why use files instead of a message queue like RabbitMQ or Redis?
Filesystem operations provide built-in persistence, simple inspection (ls, cat, grep), and zero operational dependencies. As implemented in unclebob/swarm-forge, this choice prioritizes debuggability and portability over throughput—appropriate for a development-focused multi-agent system where visibility trumps latency.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →