How the Egonex AI Multi-Agent Pipeline Coordinates Between Project-Scanner, File-Analyzer, and Architecture-Analyzer
The Egonex AI multi-agent pipeline coordinates through a read-only importMap generated by the project-scanner, which enables parallel file-analyzers to resolve cross-batch dependencies without re-parsing, before the architecture-analyzer merges all outputs into a unified knowledge graph.
The Egonex-AI/Understand-Anything repository implements a sophisticated code understanding system that transforms large repositories into navigable knowledge graphs. Understanding how the Egonex AI multi-agent pipeline coordinates between its initial scanning, parallel analysis, and architectural aggregation stages is essential for optimizing performance in enterprise-scale codebases.
The Three-Stage Coordination Flow
The pipeline operates as a directed acyclic graph where each stage consumes the artifacts of the previous agent. This design ensures immutable data contracts between the project-scanner, file-analyzer, and architecture-analyzer components.
- Project-Scanner ingests the repository root and produces
scan-result.jsoncontaining the file inventory and resolved import mappings. - File-Analyzer processes batches of files in parallel, using the
importMapto emit cross-batch edges without requiring access to the actual target file contents. - Architecture-Analyzer consumes all batch results to construct the final knowledge graph, performing layer detection and staleness checking against previous runs.
Stage 1: Repository Scanning and Import Mapping
Scanning and Categorization
Defined in agents/project-scanner.md, the project-scanner performs the initial repository traversal. It applies built-in ignore patterns alongside optional .understandignore files to filter the file tree. For each discovered file, it detects language type, calculates size in lines, and assigns a category classification such as source, test, or configuration.
The ImportMap Contract
The critical output artifact is scan-result.json, which contains two top-level keys:
files: An array of file metadata objects containingpath,language,sizeLines, andcategoryimportMap: A fully resolved mapping where each source file key points to an array of internal import paths (external npm or pip dependencies are explicitly filtered out)
{
"files": [
{ "path": "src/index.ts", "language": "typescript", "sizeLines": 42, "category": "source" }
],
"importMap": {
"src/index.ts": ["src/util.ts"],
"src/util.ts": []
}
}
This importMap serves as the coordination backbone for subsequent stages, allowing analyzers to understand dependency topology without filesystem access.
Stage 2: Parallel File Analysis with Cross-Batch Context
Batching Strategy
Before invoking the file-analyzer, the pipeline computes size-limited batches from the scan-result.json file list. Up to five file-analyzer sub-agents run concurrently, each receiving a dedicated batch along with the complete read-only importMap.
According to agents/file-analyzer.md, each agent executes the bundled extract-structure.mjs script, which utilizes tree-sitter parsers to extract symbols, definitions, and edges for every file in its assigned batch.
Cross-Batch Edge Resolution
When the analyzer encounters an import referencing a file outside its current batch, it consults the neighborMap—the specific entry from the global importMap for that target file. If found, the analyzer emits a precise cross-batch edge; otherwise, it falls back to recording the raw import text.
This mechanism allows parallel processing without race conditions, as agents never write to shared state and rely solely on the pre-computed importMap for dependency resolution.
// Conceptual dispatch from the skill implementation
await dispatchSubagent('project-scanner', {
projectRoot: cwd,
ignorePatterns: ['node_modules/**', '.git/**']
});
const scanResult = await readJSON('scan-result.json');
const batches = computeBatches(scanResult.files, scanResult.importMap);
await Promise.all(
batches.map(batch => dispatchSubagent('file-analyzer', {
batchFiles: batch,
importMap: scanResult.importMap
}))
);
Stage 3: Graph Aggregation and Architecture Analysis
Merging Sub-Graphs
The architecture-analyzer, defined in agents/architecture-analyzer.md, collects all batch-result-*.json files from the previous stage. It merges these individual sub-graphs into a single coherent graph.json representation using logic implemented in packages/core/src/analyzer/graph-builder.ts.
High-Level Analysis
Beyond simple aggregation, this agent performs sophisticated architectural detection:
- Layer detection: Identifies architectural tiers and dependency constraints
- Language-lesson generation: Creates educational content about code patterns
- Staleness checking: Compares the new graph against persisted versions to enable incremental updates
The final output serves as the authoritative knowledge graph for the repository, consumed by downstream tools like the tour-builder and graph-reviewer agents.
Summary
- The project-scanner generates a centralized
importMapthat acts as the coordination contract for the entire pipeline - File-analyzer agents process batches in parallel (up to five concurrent instances) using the read-only
importMapto resolve cross-batch dependencies without filesystem conflicts - Architecture-analyzer merges distributed batch results into a unified knowledge graph while performing high-level architectural analysis
- The pipeline achieves scalability through immutable data contracts where each stage transforms artifacts rather than sharing mutable state
Frequently Asked Questions
How does the Egonex AI pipeline resolve dependencies between files in different batches?
The file-analyzer resolves cross-batch dependencies by consulting the importMap generated by the project-scanner. When encountering an import target not present in its current batch, the analyzer looks up the target in the neighborMap entry. If found, it emits a structured edge reference; otherwise, it records the raw import string. This eliminates the need for analyzers to access files outside their assigned batches.
What is the maximum number of file-analyzer agents that run simultaneously?
The pipeline dispatches up to five file-analyzer sub-agents concurrently. Each operates on a distinct batch of files while reading from the same immutable importMap. This parallelism significantly accelerates processing for large repositories while preventing race conditions since no agent writes to shared state during analysis.
What format does the architecture-analyzer produce?
The architecture-analyzer outputs a consolidated graph.json file representing the complete knowledge graph of the repository. This JSON structure includes nodes for all symbols and definitions, edges representing imports/calls/exports (including cross-batch connections), and metadata for architectural layers. The analyzer also generates supplementary files for language lessons and tour navigation.
How does the project-scanner filter irrelevant files?
The project-scanner applies a built-in ignore list covering common directories like node_modules and .git, while also respecting repository-specific patterns defined in .understandignore files. It categorizes remaining files as source, test, or configuration based on heuristics before recording them in scan-result.json.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →