How TeamAI’s Deep-Enrich Process Generates Knowledge Docs from Extracted Code
TeamAI’s deep-enrich command transforms raw code extraction evidence into structured knowledge documents—including component designs, architecture overviews, and graph-based insights—through a six-phase pipeline orchestrated in src/deep-enrich.ts.
The Tencent/teamai-cli repository provides a sophisticated documentation pipeline that bridges the gap between raw code extraction and consumable knowledge artifacts. The teamai deep-enrich process takes the output from teamai codebase --extract and automatically generates comprehensive markdown documentation describing your software’s components, architecture, and relationships.
Overview of the Deep-Enrich Architecture
The deep-enrich workflow is implemented as a deterministic, multi-phase pipeline that progressively enriches extracted code into a navigable knowledge base. According to the source code in src/deep-enrich.ts, the process begins by loading extraction manifests and proceeds through component analysis, architectural synthesis, graph generation, and index optimization.
The pipeline uses defensive programming techniques throughout: it skips unsafe component slugs, falls back from batch to sequential AI processing when rate limits hit, and tolerates non-blocking failures such as missing architecture overviews. This ensures partial extractions still yield usable documentation.
Phase-by-Phase Execution Flow
The deep-enrich process executes six distinct phases, each with specific inputs, AI interactions, and artifacts.
Phase 0: Preparation and Manifest Loading
The pipeline starts by validating the extraction evidence directory and loading the extraction manifest from _manifest.json. If the manifest is missing, the parser fails silently; if no components are found, the command aborts immediately with the log message:
deep-enrich[<project>]: No components in _manifest.json, aborting
During this phase, the script scans the modules directory to build an internal registry of components, capping the processing volume according to the opts.maxModules parameter.
Phase 1: Component Design Documentation
For each discovered component, the pipeline generates individual design documents through AI-assisted analysis. The system attempts batch processing for efficiency, but automatically falls back to sequential processing if the batch fails. Key implementation details include:
- Iteration over the component list (respecting
opts.maxModuleslimits) - Skip logic for components with invalid slugs
- Output logging:
deep-enrich[<project>]: Component doc written: <outPath>
Phase 1 produces markdown files describing each component’s API, responsibilities, and implementation details.
Phase 2: Architecture Overview Generation
After component docs are complete, the pipeline aggregates them to generate a high-level architecture overview. This phase feeds the consolidated component descriptions to an LLM to synthesize a project-wide architectural summary. The implementation in src/deep-enrich.ts includes graceful abort handling: if the LLM returns an empty result, the process logs Architecture overview generation failed, skipping and continues rather than crashing.
Phase 3: Deterministic Graph Documentation
This phase constructs static graph documentation (designated G1–G4) from the component metadata without AI involvement. The deterministic graph represents concrete relationships—imports, inheritances, and call graphs—derived directly from the extraction evidence. These files are written to <docsDir>/graph as machine-readable relationship maps.
Phase 4: AI-Enhanced Graph Insights
Building upon the deterministic graph, Phase 4 generates richer, multi-hop insights (G5/G6) using AI analysis. This includes scenario diagrams and complex relationship explanations that require inferential reasoning beyond static code structure.
The pipeline includes conditional logic to skip G5 generation when the project contains fewer than two modules or lacks an architecture overview from Phase 2. If G5 generation fails, the system logs G5 scenario generation failed, skipping and continues with remaining phases.
Phase 5: Index Enhancement and Routing Table Updates
The final phase updates the knowledge-base index for fast retrieval by TeamAI agents. The system appends new component, architecture, and graph entries to the routing table at graph/README.md and performs an incremental index update. On completion, the log shows:
deep-enrich[<project>]: Index incremental update complete
CLI Integration and Usage Patterns
The deep-enrich functionality is exposed as a hidden command in the TeamAI CLI, registered in src/index.ts at lines 919–927. You can invoke it manually or integrate it into automated import workflows.
Running Deep-Enrich Manually
teamai codebase --deep-enrich \
--project my-awesome-app \
--output ./knowledge
--projectspecifies the slug used during the initial extraction--outputdefines the target directory for generated markdown files
Integrating with Import Workflows
Higher-level commands such as teamai import optionally invoke deep-enrich after repository extraction. As implemented in src/import.ts (lines 72–119), the integration uses dynamic imports to maintain CLI performance:
import { extractRepo } from './extract-repo.js';
async function importAndEnrich(repoUrl: string) {
const { slug } = await extractRepo(repoUrl);
try {
// Dynamically imported to keep CLI lightweight when not needed
const { deepEnrich } = await import('./deep-enrich.js');
await deepEnrich({ project: slug, output: './knowledge' });
} catch (e) {
console.warn(`deep-enrich failed for ${slug}:`, e);
}
}
This pattern allows the CLI to load the deep-enrich module only when documentation generation is required, reducing startup time for other commands.
Key Implementation Files and Testing
The deep-enrich process relies on several critical files within the Tencent/teamai-cli repository:
src/deep-enrich.ts— Core implementation containing the six-phase pipeline logic and error handlingsrc/index.ts— CLI registration point for theteamai codebase --deep-enrichhidden command (lines 919–927)src/import.ts— Demonstrates integration patterns and dynamic module loading (lines 113–119)src/__tests__/deep-enrich-graph-preserve.test.ts— Unit tests ensuring graph documentation remains stable across multiple enrichment runsskills/team-wiki-codebase/references/agents/kb-doc-generator.md— Specification for the knowledge-base document generator used by the pipeline
Error Handling and Resilience
The deep-enrich process implements comprehensive defensive mechanisms to handle incomplete or corrupted extraction data. If the _manifest.json file is malformed, the parser fails silently rather than throwing. When processing individual components, unsafe slugs are automatically skipped with logging: Skipping component …: ….
The pipeline distinguishes between blocking and non-blocking failures. Missing architecture overviews or failed G5 scenario generation trigger warning messages but do not terminate the process. However, if no components exist in the manifest, the system aborts early to prevent wasted compute cycles. This resilience ensures that even partial code extractions produce valuable documentation artifacts.
Summary
- TeamAI’s deep-enrich process converts raw extraction evidence into structured knowledge through six deterministic phases implemented in
src/deep-enrich.ts. - The pipeline generates component design docs (Phase 1), architecture overviews (Phase 2), deterministic graph docs (Phase 3), and AI-enhanced graph insights (Phase 4) before updating the search index (Phase 5).
- Command registration occurs in
src/index.tswith integration patterns demonstrated insrc/import.ts, including dynamic loading for performance optimization. - Defensive error handling allows partial success, skipping unsafe components and tolerating missing architecture data while aborting only when no components are found.
Frequently Asked Questions
What triggers the deep-enrich process to skip a component during Phase 1?
The pipeline skips components with unsafe slugs or those that fail individual AI generation, logging Skipping component …: … and continuing with the next item. If batch processing fails for a group of components, the system automatically falls back to sequential processing to maximize throughput while minimizing data loss.
How does the deep-enrich process handle missing or corrupted manifest files?
If _manifest.json is missing, the parser fails silently; if it contains no components, the process logs deep-enrich[<project>]: No components in _manifest.json, aborting and exits immediately. For malformed JSON, the extraction directory scan continues based on filesystem discovery, though this may limit certain metadata-dependent features.
What is the difference between deterministic graph docs (G1–G4) and AI-enhanced graph docs (G5/G6)?
Deterministic graph docs (Phase 3) derive directly from static code analysis, mapping concrete relationships like imports and inheritance without AI inference. AI-enhanced graph docs (Phase 4) use LLM reasoning to generate multi-hop insights, scenario diagrams, and complex relationship explanations that require contextual understanding beyond raw syntax.
Can deep-enrich run incrementally on previously processed projects?
Yes. Phase 5 performs incremental index updates by appending new entries to graph/README.md rather than rebuilding the entire index. The test file src/__tests__/deep-enrich-graph-preserve.test.ts validates that graph documentation remains consistent across multiple runs, ensuring idempotent behavior for CI/CD integrations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →