How the Understand-Anything Pipeline Implements Batching Strategy for Subagents with Concurrent Workers and Files Per Batch
The Understand-Anything pipeline processes up to 5 concurrent file-analyzer subagents using semantic batches capped at 35 files for code and 20 files for non-code assets, merging tiny groups below 3 files into optimized "misc" batches of 25.
The Egonex-AI/Understand-Anything repository orchestrates complex codebase analysis through a sophisticated batching strategy for subagents with concurrent workers and files per batch. This system, implemented primarily in compute-batches.mjs and executed during Phase 2 of the understand skill, balances semantic coherence with computational constraints by grouping files via graph algorithms before dispatching parallel workers.
Semantic Community Detection via Import Graph Analysis
The pipeline begins with Louvain community detection on the import dependency graph to create semantically meaningful batches.
In compute-batches.mjs, the runLouvain function analyzes file relationships to identify natural boundaries between functional modules. Files within the same import community are grouped together, ensuring that related components are analyzed by the same subagent for context preservation.
If the Louvain algorithm fails or returns insufficient communities, the system falls back to countBasedAssignment, which creates deterministic chunks of exactly 12 files each. This fallback guarantees that the pipeline continues processing even when graph analysis encounters edge cases.
Size Enforcement and Hard Limits for Code Batches
To prevent token overflow and ensure subagent efficiency, the strategy enforces strict upper bounds on batch sizes.
Communities exceeding 35 files (MAX_COMMUNITY_SIZE) are automatically split alphabetically until all batches meet this limit. This hard cap keeps the file-analyzer subagent prompts within manageable token limits while maintaining the semantic relationships identified by the Louvain algorithm.
The threshold of 35 represents a calculated balance between:
- Parallel efficiency — Larger batches reduce total worker count
- Context window constraints — Preventing prompt truncation during analysis
- Cognitive load — Ensuring subagents can effectively reason about interdependencies
Non-Code File Grouping Strategy
Non-code assets follow a separate classification path that prioritizes file type coherence over import graphs.
The system categorizes non-code files into semantic groups A-D (Dockerfiles, CI workflows, SQL migrations, and configuration files) that are explicitly marked as non-mergeable. These groups maintain isolation because their analysis requires specialized context.
Remaining non-code files are batched by their immediate parent directory with a limit of 20 files per batch (MAX_E = 20). This directory-based grouping ensures that files with shared path contexts are analyzed together, improving the subagent's ability to infer architectural patterns.
Merging Tiny Batches for Optimization
After initial grouping, the pipeline optimizes worker efficiency by consolidating undersized batches.
Any batch marked mergeable: true containing fewer than 3 files (MIN_BATCH_SIZE = 3) is pooled into a collection. The system then redistributes these files into new "misc" batches with a maximum of 25 files (`MAX_MERGE_TARGET = 25).
This merging strategy reduces the total number of subagent launches while preserving semantic boundaries for non-mergeable groups. The implementation in compute-batches.mjs sorts pooled files alphabetically before slicing to ensure deterministic output:
const MIN_BATCH_SIZE = 3;
const MAX_MERGE_TARGET = 25;
// After initial community grouping
if (smallMergeable.length) {
const pooledFiles = smallMergeable.flatMap(b => b.files).sort((a, b) =>
a.path.localeCompare(b.path)
);
const miscBatches = [];
for (let i = 0; i < pooledFiles.length; i += MAX_MERGE_TARGET) {
miscBatches.push({ files: pooledFiles.slice(i, i + MAX_MERGE_TARGET) });
}
}
Concurrent Worker Dispatch in Phase 2
Phase 2 of the understand skill dispatches the file-analyzer subagents using up to 5 concurrent workers, as documented in SKILL.md.
Each worker receives one complete batch containing the file paths, batchImportData, and neighborMap for its assigned slice. This injection allows subagents to work independently without cross-worker communication:
# Pseudo-code for the orchestrator
for batch in $(jq -c '.batches[]' batches.json); do
# Launch at most 5 parallel file-analyzer sub-agents
dispatch file-analyzer "$batch" &
[[ $(jobs -r | wc -l) -ge 5 ]] && wait -n
done
wait # Wait for remaining workers
The concurrency limit of 5 prevents resource exhaustion while maximizing throughput. When a worker completes, the orchestrator immediately dispatches the next pending batch, maintaining the worker pool until all batches are processed.
Summary
- Semantic grouping: Code files are batched via Louvain community detection on import graphs, with a deterministic fallback of 12 files per chunk.
- Size limits: Code batches are capped at 35 files (
MAX_COMMUNITY_SIZE), while non-code directory batches are limited to 20 files (MAX_E). - Optimization: Mergeable batches below 3 files are pooled and redistributed into "misc" batches of at most 25 files (
MAX_MERGE_TARGET). - Concurrency: The pipeline dispatches up to 5 concurrent file-analyzer subagents during Phase 2, with each worker processing one batch independently.
- Source locations: Core logic resides in
compute-batches.mjswith dispatch orchestration defined inSKILL.md.
Frequently Asked Questions
What is the maximum number of files per batch in the Understand-Anything pipeline?
The maximum varies by file type. Code batches are limited to 35 files after Louvain community detection splits, while non-code files grouped by directory are capped at 20 files. When merging tiny batches, the resulting "misc" batches cannot exceed 25 files.
How does the pipeline handle files that don't belong to a clear import community?
If Louvain community detection fails, the system falls back to countBasedAssignment in compute-batches.mjs, which creates deterministic chunks of exactly 12 files each. Additionally, non-code files that don't fit into semantic groups A-D are batched by their immediate parent directory rather than import relationships.
Why does the pipeline merge small batches before dispatching subagents?
Batches containing fewer than 3 files (MIN_BATCH_SIZE) are merged to reduce orchestrational overhead and minimize the total number of subagent launches. This optimization prevents the inefficiency of spawning workers for trivially small tasks while keeping the merged "misc" batches under 25 files to preserve token limits.
How many subagents run simultaneously during the analysis phase?
The understand skill runs up to 5 concurrent file-analyzer subagents during Phase 2. This concurrency limit is enforced by the orchestrator described in SKILL.md, which maintains the worker pool and dispatches new batches as existing workers complete their analysis.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →