# How /understand-knowledge Analyzes LLM Wikis into Knowledge Graphs: A 5-Phase Technical Breakdown

> Discover how /understand-knowledge transforms LLM wikis into knowledge graphs. Explore the 5-phase technical pipeline combining Python parsing and LLM semantic analysis for interactive visualization.

- Repository: [Yuxiang Lin/Understand-Anything](https://github.com/Lum1104/Understand-Anything)
- Tags: deep-dive
- Published: 2026-06-06

---

**The `/understand-knowledge` skill converts Karpathy-pattern LLM wikis into interactive force-directed knowledge graphs through a deterministic five-phase pipeline that combines Python-based structural parsing with LLM-driven semantic analysis.**

The `Lum1104/Understand-Anything` repository provides a Claude Code plugin that transforms markdown-based LLM wikis into structured knowledge graphs. The `/understand-knowledge` skill orchestrates this conversion through a hybrid architecture that merges deterministic file parsing with parallel LLM agent processing to extract both explicit wikilinks and implicit conceptual relationships.

## The Five-Phase Pipeline

The skill processes wiki directories through five deterministic phases, each implemented by specialized Python scripts and sub-agents defined in [`understand-anything-plugin/skills/understand-knowledge/SKILL.md`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand-knowledge/SKILL.md).

### Phase 1: Detect and Parse

The pipeline begins by determining the target directory and executing the deterministic parser [`parse-knowledge-base.py`](https://github.com/Lum1104/Understand-Anything/blob/main/parse-knowledge-base.py). This script scans the wiki structure and writes a [`scan-manifest.json`](https://github.com/Lum1104/Understand-Anything/blob/main/scan-manifest.json) file that describes raw sources, wiki pages, topics, and wikilinks. According to **SKILL.md** lines 30-34, this phase establishes the foundational structure without LLM involvement, ensuring a baseline graph exists even if subsequent AI processing fails.

### Phase 2: Scan and Structure

Phase 2 operates entirely within the manifest generation completed by Phase 1. The [`scan-manifest.json`](https://github.com/Lum1104/Understand-Anything/blob/main/scan-manifest.json) contains:

- **Article nodes**: One per markdown file with extracted wikilinks, headings, and front-matter
- **Source nodes**: Files located under the `raw/` directory
- **Topic nodes**: Extracted from [`index.md`](https://github.com/Lum1104/Understand-Anything/blob/main/index.md) headings
- **Edges**: `related` links (wikilinks) and `categorized_under` relationships (index sections)

As documented in **SKILL.md** lines 43-48, this manifest captures the explicit structural connections present in the markdown files.

### Phase 3: Analyze with LLM Agents

The system dispatches the `article-analyzer` sub-agent (defined in [`understand-anything-plugin/agents/article-analyzer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/article-analyzer.md)) to process batches of 10-15 articles concurrently. Each batch receives the article metadata plus the complete node-ID list, enabling the LLM to infer implicit relationships, entities, and claims that exist conceptually but not explicitly as wikilinks. The orchestration steps in **SKILL.md** lines 54-68 describe how up to three batches run in parallel to minimize latency for large wikis.

Each batch writes its results to `analysis-batch-{N}.json` files containing enriched semantic connections.

### Phase 4: Merge and Normalize

The [`merge-knowledge-graph.py`](https://github.com/Lum1104/Understand-Anything/blob/main/merge-knowledge-graph.py) script (referenced in **SKILL.md** lines 74-85) consolidates the deterministic manifest with the LLM-generated analysis files. This phase performs:

1. **Entity deduplication**: Case-insensitive matching to eliminate redundant nodes
2. **Type normalization**: Standardizing `article`, `source`, `topic`, `entity`, and `claim` node types
3. **Layer construction**: Building hierarchical layers from [`index.md`](https://github.com/Lum1104/Understand-Anything/blob/main/index.md) categories
4. **Tour ordering**: Creating navigational sequences through the graph

The output is [`assembled-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/assembled-graph.json), a unified representation combining explicit structural links with AI-inferred semantic connections.

### Phase 5: Validate and Persist

The final phase validates the assembled graph using logic from [`understand-anything-plugin/packages/core/src/analyzer/graph-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/analyzer/graph-builder.ts). Validation checks that every edge's source and target nodes exist and that all required node fields are present. As detailed in **SKILL.md** lines 91-125, the system then:

- Writes the final [`knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/knowledge-graph.json) to `.<understand-anything>/`
- Persists a small [`meta.json`](https://github.com/Lum1104/Understand-Anything/blob/main/meta.json) with processing metadata
- Cleans up the intermediate folder
- Auto-triggers `/understand-dashboard` to render the visualization

## Architecture Highlights

**Deterministic Pre-processing**

The [`parse-knowledge-base.py`](https://github.com/Lum1104/Understand-Anything/blob/main/parse-knowledge-base.py) script extracts only structural information—wikilinks, headings, front-matter, and categories—providing a guaranteed baseline graph independent of LLM availability or reliability.

**LLM-Driven Enrichment**

The `article-analyzer` agents receive the deterministic graph context, allowing them to add implicit edges such as entity-to-entity relationships or claim-based links (`implies` edges) that markdown syntax alone cannot express.

**Layered Graph Model**

Nodes carry explicit types (`article`, `source`, `topic`, `entity`, `claim`) while edges encode relationship semantics (`related`, `categorized_under`, `implies`). The `kind: "knowledge"` flag signals the dashboard to render a force-directed layout rather than the default hierarchical DAG visualization implemented in [`KnowledgeGraphView.tsx`](https://github.com/Lum1104/Understand-Anything/blob/main/KnowledgeGraphView.tsx).

**Parallel Execution**

The batch processing architecture allows up to three analysis batches to run concurrently, significantly reducing processing time for wikis containing hundreds of articles.

## Running the Pipeline

### Execute the Complete Skill

```bash
/understand-knowledge /path/to/wiki

```

If no path is provided, the command uses the current working directory. This triggers all five phases sequentially.

### Run the Deterministic Parser Manually

```bash
python3 understand-anything-plugin/skills/understand-knowledge/parse-knowledge-base.py /path/to/wiki

```

Output writes to [`/path/to/wiki/.understand-anything/intermediate/scan-manifest.json`](https://github.com/Lum1104/Understand-Anything/blob/main//path/to/wiki/.understand-anything/intermediate/scan-manifest.json).

### Dispatch Analysis Batches

The sub-agent receives input structured as:

```json
{
  "name": "article-analyzer",
  "input": {
    "batchId": 1,
    "articles": [
      { 
        "id": "a1", 
        "name": "Intro", 
        "summary": "...", 
        "wikilinks": ["b2"], 
        "category": "Basics", 
        "content": "..."
      }
    ],
    "existingNodeIds": ["a1","b2","t1"]
  }
}

```

Results write to [`analysis-batch-1.json`](https://github.com/Lum1104/Understand-Anything/blob/main/analysis-batch-1.json) in the intermediate folder.

### Merge and Finalize

```bash
python3 understand-anything-plugin/skills/understand-knowledge/merge-knowledge-graph.py /path/to/wiki

```

Produces [`assembled-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/assembled-graph.json) with deduplicated entities and normalized layers.

### Launch the Dashboard

```bash
/understand-dashboard /path/to/wiki

```

The dashboard reads [`.understand-anything/knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/.understand-anything/knowledge-graph.json) and renders the force-directed visualization with community clustering.

## Summary

- **[`parse-knowledge-base.py`](https://github.com/Lum1104/Understand-Anything/blob/main/parse-knowledge-base.py)** performs deterministic structural extraction, creating [`scan-manifest.json`](https://github.com/Lum1104/Understand-Anything/blob/main/scan-manifest.json) with explicit wikilinks and categories
- **[`article-analyzer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/article-analyzer.md)** defines LLM agents that process articles in batches of 10-15 to infer implicit semantic relationships
- **[`merge-knowledge-graph.py`](https://github.com/Lum1104/Understand-Anything/blob/main/merge-knowledge-graph.py)** combines deterministic and AI-generated data, deduplicating entities case-insensitively and normalizing node types
- **[`graph-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/graph-builder.ts)** validates the final graph structure before persisting [`knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/knowledge-graph.json) and triggering the dashboard
- **Parallel processing** of up to three batches simultaneously optimizes performance for large wiki repositories

## Frequently Asked Questions

### What is the "Karpathy-pattern" mentioned in the LLM wiki format?

The Karpathy-pattern refers to a markdown-based note-taking structure popularized by Andrej Karpathy, characterized by interlinked markdown files using `[[wikilink]]` syntax, hierarchical organization through [`index.md`](https://github.com/Lum1104/Understand-Anything/blob/main/index.md) files with categorized sections, and a `raw/` folder for source materials. The parser specifically recognizes these conventions to build the initial graph topology.

### How does the system handle duplicate entities across different articles?

During Phase 4, [`merge-knowledge-graph.py`](https://github.com/Lum1104/Understand-Anything/blob/main/merge-knowledge-graph.py) performs case-insensitive deduplication to identify when the same entity appears with different capitalizations or slight variations across articles. The merge script normalizes these into single canonical nodes while preserving all incoming relationship edges from the various mentions.

### Can I customize the batch size for the LLM analysis phase?

While the default configuration processes 10-15 articles per batch as defined in **SKILL.md** lines 54-68, the architecture supports batch size adjustments through the skill configuration. However, the three-batch parallel execution limit is designed to balance API rate limits with processing throughput for optimal performance.

### What distinguishes knowledge graphs from the standard DAG visualizations in the dashboard?

Knowledge graphs use the `kind: "knowledge"` flag to trigger a force-directed physics simulation layout in [`KnowledgeGraphView.tsx`](https://github.com/Lum1104/Understand-Anything/blob/main/KnowledgeGraphView.tsx), optimized for exploring dense, interconnected conceptual relationships. Standard DAG visualizations enforce hierarchical directionality, while knowledge graphs emphasize community clustering and bidirectional relationship exploration through interactive node dragging and zooming.