SwarmForge Performance Considerations for Large Numbers of Concurrent Projects
SwarmForge manages concurrent projects through isolated git worktrees and tmux sessions per role, with scalability limits determined by file descriptors, pseudo-terminals, and memory consumption per backend process.
SwarmForge, an open-source multi-agent development framework by Uncle Bob, treats each project as an isolated environment where role-specific agents collaborate through git-based handoffs. When scaling to dozens of concurrent projects, understanding the architectural bottlenecks becomes essential for maintaining responsive performance on developer workstations.
Git Worktrees: Linear Growth with Project Count
Every project in SwarmForge creates git worktrees under .worktrees/<role> to provide clean checkouts for each agent. Handoffs materialize as ordinary git commits, enabling transparent collaboration between roles.
The worktree count grows linearly: projects × roles. Each worktree consumes disk space and a file-descriptor handle for the underlying git process.
Worktrees are created lazily when a project opens (see the Open Project flow in the README). Roles marked window-invisible run background agents without UI overhead, deferring worktree creation until actually needed.
To monitor worktree saturation:
# Count active worktrees under the current forge
find .worktrees -type d -mindepth 2 -maxdepth 2 | wc -l
Tmux Sessions: Socket Limits and UI Isolation
SwarmForge isolates project UI state through tmux sessions, with each project receiving its own socket at .swarmforge/tmux-socket and a session per role.
Tmux imposes hard limits on windows/panes per session. System-wide pseudo-terminal exhaustion can occur when spawning many concurrent sessions.
The dashboard mitigates this by creating a single tmux server per project, not per role—scaling session count with projects rather than roles. Packs default to window-invisible to minimize UI footprint.
When scaling beyond typical limits, override the socket path to avoid path-length issues:
# Use shorter socket path for many projects
export SWARMFORGE_SOCKET=/tmp/swarm-socket
./swarm-forge start
Handoff Files: Atomic Lock Contention
Agent communication flows through handoff files stored under .swarmforge/handoffs/ across four subdirectories: outbox, inbox, sent, and failed. This filesystem-based messaging eliminates network traffic but introduces shared-resource contention.
The handoff protocol implements an atomic lock when updating the global handoff sequence. As documented in swarmforge/handoff-protocol.md, this guarantees that only one handoff writer proceeds at a time, ensuring consistency under heavy parallel traffic.
For high-throughput scenarios, configure batch receive mode to reduce lock frequency:
# In projects/<name>/swarmforge/swarmforge.conf
window-invisible coder grok wt-coder batch back-one
Batch mode reduces wake-up notifications and handoff file writes, easing contention on the atomic lock.
Dashboard API: In-Memory State Management
The pack_web HTTP server exposes a minimal API at /api/projects/* for project lifecycle management. Every open project adds an entry to the server's in-memory state, requiring iteration across all open projects for requests like the Attention list.
The API surface is deliberately constrained to three endpoints:
GET /api/projects/openPOST /api/projects/openPOST /api/projects/close
Server-side caching via .swarmforge/open-projects avoids costly filesystem scans, as implemented in swarmforge/scripts/pack_web.bb (lines 1725–1726).
Batch project creation via the HTTP API:
for i in {1..50}; do
curl -X POST http://localhost:8080/api/projects/open \
-H "Content-Type: application/json" \
-d "{\"name\":\"proj_$i\",\"mission\":\"Demo project $i\",\"pack\":\"four-pack\"}"
done
Backend Processes: Memory as the Primary Constraint
Each role executes a backend process (Claude, Codex, Grok, etc.) that drives agent behavior. These processes represent the most significant memory bottleneck: each may consume several hundred MB of RAM, making memory the typical limiting factor for workstation deployments.
SwarmForge mitigates through role sharing—the same backend binary serves multiple agents—and headless execution for window-invisible roles, eliminating terminal emulator overhead.
Select lighter backends for high-volume roles in swarmforge.conf:
# Resource-aware backend selection
window-invisible specifier codex ...
Practical Scaling Recommendations
| Optimization | Implementation | Impact |
|---|---|---|
| Prefer invisible windows | Set window-invisible for all non-UI roles |
Eliminates tmux panes, reduces CPU/memory |
| Limit concurrent projects | Keep active count under OS pty limit (typically 1024 on Linux) | Prevents resource exhaustion |
| Batch handoffs | Configure batch receive mode in swarmforge.conf |
Reduces lock contention and wake-ups |
| Prune handoff directories | Clean failed/ and sent/ subdirectories periodically |
Controls disk growth |
| Override tmux socket | Set SWARMFORGE_SOCKET=/tmp/swarm-socket |
Avoids path-length limits |
| Choose lighter backends | Use codex vs. heavier alternatives for high-volume roles |
Reduces per-process memory |
Periodic cleanup automation:
#!/usr/bin/env bash
# Prune old handoff files every hour
find .swarmforge/handoffs/* -type f -mtime +7 -delete
Key Source Files for Performance Analysis
| File | Relevance |
|---|---|
README.md |
Architecture overview, pack defaults, role definitions |
swarmforge/scripts/pack_web.bb |
HTTP endpoints for project lifecycle (L1725–1726) |
swarmforge/handoff-protocol.md |
Atomic lock implementation for handoff sequencing (L105–108) |
swarmforge/scripts/swarm_handoff.sh |
Handoff validation and queuing logic |
swarmforge/swarmforge.conf |
Role/window configuration syntax |
Summary
- Git worktrees scale linearly with
projects × rolesbut are created lazily to defer resource consumption - Tmux sessions scale with project count, not role count, with
window-invisibledefaults minimizing overhead - Handoff atomic locks serialize concurrent writes to prevent race conditions under heavy load
- Dashboard API uses minimal endpoints and file-based caching to limit memory growth
- Backend processes typically constrain scaling before other resources due to per-role memory requirements
With invisible windows, batch handoff mode, and resource-aware backend selection, SwarmForge manages dozens of concurrent projects on standard developer workstations. Beyond OS limits, consolidate roles or distribute across multiple forge instances.
Frequently Asked Questions
How many concurrent projects can SwarmForge handle on a typical workstation?
SwarmForge comfortably manages 20–50 concurrent projects on a standard developer workstation with 32GB RAM, assuming window-invisible roles and moderate handoff traffic. The practical limit is typically pseudo-terminals (1024 on default Linux) or aggregate backend memory consumption rather than CPU or disk. Projects with many visible windows or heavyweight LLM backends reduce this capacity proportionally.
What causes handoff contention and how can I detect it?
Handoff contention occurs when multiple agents attempt simultaneous writes to .swarmforge/handoffs/, competing for the atomic lock that sequences global handoff state. Symptoms include delayed agent responses and elevated CPU in swarm_handoff.sh processes. Enable batch receive mode for non-urgent roles and monitor lock wait times through system tools like strace -e flock on the handoff script.
Does closing a project free all associated resources?
Closing a project via POST /api/projects/close releases the tmux session and removes the entry from the dashboard's in-memory state. However, git worktrees persist on disk and backend process cleanup depends on graceful termination. For complete resource reclamation, periodically prune orphaned worktrees and verify no zombie processes remain via ps aux | grep swarm.
Can I run multiple SwarmForge instances on one machine?
Multiple forge instances are supported by isolating their working directories and tmux socket paths. Set independent SWARMFORGE_SOCKET values and launch from separate directories. This effectively partitions resource pools and bypasses per-instance limits, though aggregate system resources (RAM, file descriptors) remain shared. This pattern suits CI/CD environments or team-shared servers better than single-workstation scaling.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →