# How TeamAI’s Deep-Enrich Process Generates Knowledge Docs from Extracted Code

> Discover TeamAI's deep-enrich process. This six-phase pipeline transforms extracted code into structured knowledge docs like component designs and architecture overviews. Learn more from the Tencent teamai-cli repository.

- Repository: [Tencent/teamai-cli](https://github.com/tencent/teamai-cli)
- Tags: deep-dive
- Published: 2026-09-11

---

**TeamAI’s deep-enrich command transforms raw code extraction evidence into structured knowledge documents—including component designs, architecture overviews, and graph-based insights—through a six-phase pipeline orchestrated in [`src/deep-enrich.ts`](https://github.com/Tencent/teamai-cli/blob/main/src/deep-enrich.ts).**

The **Tencent/teamai-cli** repository provides a sophisticated documentation pipeline that bridges the gap between raw code extraction and consumable knowledge artifacts. The **teamai deep-enrich process** takes the output from `teamai codebase --extract` and automatically generates comprehensive markdown documentation describing your software’s components, architecture, and relationships.

## Overview of the Deep-Enrich Architecture

The deep-enrich workflow is implemented as a deterministic, multi-phase pipeline that progressively enriches extracted code into a navigable knowledge base. According to the source code in [`src/deep-enrich.ts`](https://github.com/Tencent/teamai-cli/blob/main/src/deep-enrich.ts), the process begins by loading extraction manifests and proceeds through component analysis, architectural synthesis, graph generation, and index optimization.

The pipeline uses defensive programming techniques throughout: it skips unsafe component slugs, falls back from batch to sequential AI processing when rate limits hit, and tolerates non-blocking failures such as missing architecture overviews. This ensures partial extractions still yield usable documentation.

## Phase-by-Phase Execution Flow

The deep-enrich process executes six distinct phases, each with specific inputs, AI interactions, and artifacts.

### Phase 0: Preparation and Manifest Loading

The pipeline starts by validating the extraction evidence directory and loading the extraction manifest from [`_manifest.json`](https://github.com/Tencent/teamai-cli/blob/main/_manifest.json). If the manifest is missing, the parser fails silently; if no components are found, the command aborts immediately with the log message:

```text
deep-enrich[<project>]: No components in _manifest.json, aborting

```

During this phase, the script scans the modules directory to build an internal registry of components, capping the processing volume according to the `opts.maxModules` parameter.

### Phase 1: Component Design Documentation

For each discovered component, the pipeline generates individual design documents through AI-assisted analysis. The system attempts batch processing for efficiency, but automatically falls back to sequential processing if the batch fails. Key implementation details include:

- Iteration over the component list (respecting `opts.maxModules` limits)
- Skip logic for components with invalid slugs
- Output logging: `deep-enrich[<project>]: Component doc written: <outPath>`

Phase 1 produces markdown files describing each component’s API, responsibilities, and implementation details.

### Phase 2: Architecture Overview Generation

After component docs are complete, the pipeline aggregates them to generate a high-level architecture overview. This phase feeds the consolidated component descriptions to an LLM to synthesize a project-wide architectural summary. The implementation in [`src/deep-enrich.ts`](https://github.com/Tencent/teamai-cli/blob/main/src/deep-enrich.ts) includes graceful abort handling: if the LLM returns an empty result, the process logs `Architecture overview generation failed, skipping` and continues rather than crashing.

### Phase 3: Deterministic Graph Documentation

This phase constructs static graph documentation (designated G1–G4) from the component metadata without AI involvement. The deterministic graph represents concrete relationships—imports, inheritances, and call graphs—derived directly from the extraction evidence. These files are written to `<docsDir>/graph` as machine-readable relationship maps.

### Phase 4: AI-Enhanced Graph Insights

Building upon the deterministic graph, Phase 4 generates richer, multi-hop insights (G5/G6) using AI analysis. This includes scenario diagrams and complex relationship explanations that require inferential reasoning beyond static code structure.

The pipeline includes conditional logic to skip G5 generation when the project contains fewer than two modules or lacks an architecture overview from Phase 2. If G5 generation fails, the system logs `G5 scenario generation failed, skipping` and continues with remaining phases.

### Phase 5: Index Enhancement and Routing Table Updates

The final phase updates the knowledge-base index for fast retrieval by TeamAI agents. The system appends new component, architecture, and graph entries to the routing table at [`graph/README.md`](https://github.com/Tencent/teamai-cli/blob/main/graph/README.md) and performs an incremental index update. On completion, the log shows:

```text
deep-enrich[<project>]: Index incremental update complete

```

## CLI Integration and Usage Patterns

The deep-enrich functionality is exposed as a hidden command in the TeamAI CLI, registered in [`src/index.ts`](https://github.com/Tencent/teamai-cli/blob/main/src/index.ts) at lines 919–927. You can invoke it manually or integrate it into automated import workflows.

### Running Deep-Enrich Manually

```bash
teamai codebase --deep-enrich \
  --project my-awesome-app \
  --output ./knowledge

```

- `--project` specifies the slug used during the initial extraction
- `--output` defines the target directory for generated markdown files

### Integrating with Import Workflows

Higher-level commands such as `teamai import` optionally invoke deep-enrich after repository extraction. As implemented in [`src/import.ts`](https://github.com/Tencent/teamai-cli/blob/main/src/import.ts) (lines 72–119), the integration uses dynamic imports to maintain CLI performance:

```typescript
import { extractRepo } from './extract-repo.js';

async function importAndEnrich(repoUrl: string) {
  const { slug } = await extractRepo(repoUrl);
  try {
    // Dynamically imported to keep CLI lightweight when not needed
    const { deepEnrich } = await import('./deep-enrich.js');
    await deepEnrich({ project: slug, output: './knowledge' });
  } catch (e) {
    console.warn(`deep-enrich failed for ${slug}:`, e);
  }
}

```

This pattern allows the CLI to load the deep-enrich module only when documentation generation is required, reducing startup time for other commands.

## Key Implementation Files and Testing

The deep-enrich process relies on several critical files within the Tencent/teamai-cli repository:

- **[`src/deep-enrich.ts`](https://github.com/Tencent/teamai-cli/blob/main/src/deep-enrich.ts)** — Core implementation containing the six-phase pipeline logic and error handling
- **[`src/index.ts`](https://github.com/Tencent/teamai-cli/blob/main/src/index.ts)** — CLI registration point for the `teamai codebase --deep-enrich` hidden command (lines 919–927)
- **[`src/import.ts`](https://github.com/Tencent/teamai-cli/blob/main/src/import.ts)** — Demonstrates integration patterns and dynamic module loading (lines 113–119)
- **[`src/__tests__/deep-enrich-graph-preserve.test.ts`](https://github.com/Tencent/teamai-cli/blob/main/src/__tests__/deep-enrich-graph-preserve.test.ts)** — Unit tests ensuring graph documentation remains stable across multiple enrichment runs
- **[`skills/team-wiki-codebase/references/agents/kb-doc-generator.md`](https://github.com/Tencent/teamai-cli/blob/main/skills/team-wiki-codebase/references/agents/kb-doc-generator.md)** — Specification for the knowledge-base document generator used by the pipeline

## Error Handling and Resilience

The deep-enrich process implements comprehensive defensive mechanisms to handle incomplete or corrupted extraction data. If the [`_manifest.json`](https://github.com/Tencent/teamai-cli/blob/main/_manifest.json) file is malformed, the parser fails silently rather than throwing. When processing individual components, unsafe slugs are automatically skipped with logging: `Skipping component …: …`.

The pipeline distinguishes between blocking and non-blocking failures. Missing architecture overviews or failed G5 scenario generation trigger warning messages but do not terminate the process. However, if no components exist in the manifest, the system aborts early to prevent wasted compute cycles. This resilience ensures that even partial code extractions produce valuable documentation artifacts.

## Summary

- **TeamAI’s deep-enrich process** converts raw extraction evidence into structured knowledge through six deterministic phases implemented in [`src/deep-enrich.ts`](https://github.com/Tencent/teamai-cli/blob/main/src/deep-enrich.ts).
- The pipeline generates **component design docs** (Phase 1), **architecture overviews** (Phase 2), **deterministic graph docs** (Phase 3), and **AI-enhanced graph insights** (Phase 4) before updating the search **index** (Phase 5).
- Command registration occurs in [`src/index.ts`](https://github.com/Tencent/teamai-cli/blob/main/src/index.ts) with integration patterns demonstrated in [`src/import.ts`](https://github.com/Tencent/teamai-cli/blob/main/src/import.ts), including dynamic loading for performance optimization.
- Defensive error handling allows partial success, skipping unsafe components and tolerating missing architecture data while aborting only when no components are found.

## Frequently Asked Questions

### What triggers the deep-enrich process to skip a component during Phase 1?

The pipeline skips components with unsafe slugs or those that fail individual AI generation, logging `Skipping component …: …` and continuing with the next item. If batch processing fails for a group of components, the system automatically falls back to sequential processing to maximize throughput while minimizing data loss.

### How does the deep-enrich process handle missing or corrupted manifest files?

If [`_manifest.json`](https://github.com/Tencent/teamai-cli/blob/main/_manifest.json) is missing, the parser fails silently; if it contains no components, the process logs `deep-enrich[<project>]: No components in _manifest.json, aborting` and exits immediately. For malformed JSON, the extraction directory scan continues based on filesystem discovery, though this may limit certain metadata-dependent features.

### What is the difference between deterministic graph docs (G1–G4) and AI-enhanced graph docs (G5/G6)?

**Deterministic graph docs** (Phase 3) derive directly from static code analysis, mapping concrete relationships like imports and inheritance without AI inference. **AI-enhanced graph docs** (Phase 4) use LLM reasoning to generate multi-hop insights, scenario diagrams, and complex relationship explanations that require contextual understanding beyond raw syntax.

### Can deep-enrich run incrementally on previously processed projects?

Yes. Phase 5 performs incremental index updates by appending new entries to [`graph/README.md`](https://github.com/Tencent/teamai-cli/blob/main/graph/README.md) rather than rebuilding the entire index. The test file [`src/__tests__/deep-enrich-graph-preserve.test.ts`](https://github.com/Tencent/teamai-cli/blob/main/src/__tests__/deep-enrich-graph-preserve.test.ts) validates that graph documentation remains consistent across multiple runs, ensuring idempotent behavior for CI/CD integrations.