How Multi-Agent Researchers Analyze the GSD-Build Codebase in Parallel
The GSD system spawns four specialized gsd-codebase-mapper agents simultaneously using background tasks to map different architectural aspects of a repository, with each agent writing directly to .planning/codebase/ to avoid context window overload.
The GSD-Build (Get-Shit-Done) repository demonstrates exactly how multi-agent researchers analyze the gsd-build codebase in parallel through a sophisticated orchestration pattern. By distributing analysis across focused mapper agents, the system produces seven structured markdown documents without overwhelming the primary context window or sacrificing execution speed.
The Map-Codebase Orchestration Workflow
The parallel analysis initiates through the /gsd:map-codebase command, which invokes the map-codebase workflow defined in get-shit-done/workflows/map-codebase.md. This orchestrator establishes the target directory .planning/codebase/ to store generated artifacts, then spawns four gsd-codebase-mapper sub-agents using the Task tool with run_in_background=true.
Each agent receives a distinct focus area and writes directly to the filesystem, enabling true parallel execution without blocking the orchestrator's main thread.
The Four Specialized Research Agents
The workflow partitions analysis across four domain-specific researchers:
- tech → Generates
STACK.mdandINTEGRATIONS.md - arch → Generates
ARCHITECTURE.mdandSTRUCTURE.md - quality → Generates
CONVENTIONS.mdandTESTING.md - concerns → Generates
CONCERNS.md
This division ensures each agent explores relevant file patterns—such as package.json for tech focus or test directories for quality focus—without duplicating effort or competing for context space.
Agent Implementation and File Exploration
Each mapper follows the gsd-codebase-mapper process defined in agents/gsd-codebase-mapper.md. After parsing its assigned focus from the prompt, the agent explores the repository using Read, Grep, and Glob commands to identify pertinent files.
For example, a tech-focused agent executes:
ls package.json 2>/dev/null
grep -r "import" src/ --include="*.ts"
The agent then composes documentation using embedded templates and writes the resulting markdown directly to .planning/codebase/, returning only a lightweight confirmation to the orchestrator rather than the full document contents.
Parallel Execution Architecture Benefits
Running multi-agent researchers in parallel provides three critical advantages:
- Fresh context isolation — Each mapper receives its own 200k-token Claude window, eliminating token "rot" and maintaining high-quality analysis throughout the session.
- Concurrent throughput — Background execution allows all four agents to run simultaneously, reducing total analysis time by approximately 75% compared to sequential processing.
- Payload minimization — Since agents write directly to files, the orchestrator maintains a minimal context footprint, carrying only confirmation messages rather than large documentation payloads.
Result Aggregation and Verification
After all agents complete their writes, the orchestrator verifies the existence of all seven expected markdown files in .planning/codebase/. It performs a security scan for secrets, then commits the generated codebase map using:
node ~/.claude/get-shit-done/bin/gsd-tools.cjs commit
These artifacts become consumable by downstream GSD phases such as plan-phase and execute-phase, providing structured context for subsequent development tasks.
Summary
- The
/gsd:map-codebasecommand triggers parallel analysis viaget-shit-done/workflows/map-codebase.md - Four
gsd-codebase-mapperagents run simultaneously withrun_in_background=true, each focusing on tech, arch, quality, or concerns - Each agent maintains an isolated 200k-token context window and writes directly to
.planning/codebase/ - The orchestrator aggregates results by verifying file existence and committing via
gsd-tools.cjs - This architecture prevents context window overflow while generating comprehensive documentation
Frequently Asked Questions
How does the orchestrator ensure agents execute truly in parallel?
The workflow utilizes the Task tool with the parameter run_in_background=true when spawning each gsd-codebase-mapper agent. This configuration allows all four research agents to execute simultaneously rather than sequentially, with each process running independently until completion.
What specific documents does each research agent generate?
The tech agent produces STACK.md and INTEGRATIONS.md; the arch agent creates ARCHITECTURE.md and STRUCTURE.md; the quality agent writes CONVENTIONS.md and TESTING.md; and the concerns agent generates a single CONCERNS.md file. These seven documents collectively map the entire codebase architecture.
Why does parallel analysis prevent context window overflow?
Each mapper agent receives its own dedicated 200k-token Claude context window, isolating the heavy lifting of codebase exploration from the orchestrator. Because agents write findings directly to the filesystem and return only brief confirmations, the orchestrator's context remains small and manageable throughout the analysis process.
Can the number of parallel research agents be customized?
The current implementation in get-shit-done/workflows/map-codebase.md hardcodes four specific focus areas (tech, arch, quality, concerns) with corresponding document templates. While the architecture supports parallel execution via background tasks, modifying the agent count requires editing the workflow definition and creating additional focus-specific templates in the mapper agent definition.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →