How /understand-knowledge Analyzes LLM Wikis into Knowledge Graphs: A 5-Phase Technical Breakdown
The /understand-knowledge skill converts Karpathy-pattern LLM wikis into interactive force-directed knowledge graphs through a deterministic five-phase pipeline that combines Python-based structural parsing with LLM-driven semantic analysis.
The Lum1104/Understand-Anything repository provides a Claude Code plugin that transforms markdown-based LLM wikis into structured knowledge graphs. The /understand-knowledge skill orchestrates this conversion through a hybrid architecture that merges deterministic file parsing with parallel LLM agent processing to extract both explicit wikilinks and implicit conceptual relationships.
The Five-Phase Pipeline
The skill processes wiki directories through five deterministic phases, each implemented by specialized Python scripts and sub-agents defined in understand-anything-plugin/skills/understand-knowledge/SKILL.md.
Phase 1: Detect and Parse
The pipeline begins by determining the target directory and executing the deterministic parser parse-knowledge-base.py. This script scans the wiki structure and writes a scan-manifest.json file that describes raw sources, wiki pages, topics, and wikilinks. According to SKILL.md lines 30-34, this phase establishes the foundational structure without LLM involvement, ensuring a baseline graph exists even if subsequent AI processing fails.
Phase 2: Scan and Structure
Phase 2 operates entirely within the manifest generation completed by Phase 1. The scan-manifest.json contains:
- Article nodes: One per markdown file with extracted wikilinks, headings, and front-matter
- Source nodes: Files located under the
raw/directory - Topic nodes: Extracted from
index.mdheadings - Edges:
relatedlinks (wikilinks) andcategorized_underrelationships (index sections)
As documented in SKILL.md lines 43-48, this manifest captures the explicit structural connections present in the markdown files.
Phase 3: Analyze with LLM Agents
The system dispatches the article-analyzer sub-agent (defined in understand-anything-plugin/agents/article-analyzer.md) to process batches of 10-15 articles concurrently. Each batch receives the article metadata plus the complete node-ID list, enabling the LLM to infer implicit relationships, entities, and claims that exist conceptually but not explicitly as wikilinks. The orchestration steps in SKILL.md lines 54-68 describe how up to three batches run in parallel to minimize latency for large wikis.
Each batch writes its results to analysis-batch-{N}.json files containing enriched semantic connections.
Phase 4: Merge and Normalize
The merge-knowledge-graph.py script (referenced in SKILL.md lines 74-85) consolidates the deterministic manifest with the LLM-generated analysis files. This phase performs:
- Entity deduplication: Case-insensitive matching to eliminate redundant nodes
- Type normalization: Standardizing
article,source,topic,entity, andclaimnode types - Layer construction: Building hierarchical layers from
index.mdcategories - Tour ordering: Creating navigational sequences through the graph
The output is assembled-graph.json, a unified representation combining explicit structural links with AI-inferred semantic connections.
Phase 5: Validate and Persist
The final phase validates the assembled graph using logic from understand-anything-plugin/packages/core/src/analyzer/graph-builder.ts. Validation checks that every edge's source and target nodes exist and that all required node fields are present. As detailed in SKILL.md lines 91-125, the system then:
- Writes the final
knowledge-graph.jsonto.<understand-anything>/ - Persists a small
meta.jsonwith processing metadata - Cleans up the intermediate folder
- Auto-triggers
/understand-dashboardto render the visualization
Architecture Highlights
Deterministic Pre-processing
The parse-knowledge-base.py script extracts only structural information—wikilinks, headings, front-matter, and categories—providing a guaranteed baseline graph independent of LLM availability or reliability.
LLM-Driven Enrichment
The article-analyzer agents receive the deterministic graph context, allowing them to add implicit edges such as entity-to-entity relationships or claim-based links (implies edges) that markdown syntax alone cannot express.
Layered Graph Model
Nodes carry explicit types (article, source, topic, entity, claim) while edges encode relationship semantics (related, categorized_under, implies). The kind: "knowledge" flag signals the dashboard to render a force-directed layout rather than the default hierarchical DAG visualization implemented in KnowledgeGraphView.tsx.
Parallel Execution
The batch processing architecture allows up to three analysis batches to run concurrently, significantly reducing processing time for wikis containing hundreds of articles.
Running the Pipeline
Execute the Complete Skill
/understand-knowledge /path/to/wiki
If no path is provided, the command uses the current working directory. This triggers all five phases sequentially.
Run the Deterministic Parser Manually
python3 understand-anything-plugin/skills/understand-knowledge/parse-knowledge-base.py /path/to/wiki
Output writes to /path/to/wiki/.understand-anything/intermediate/scan-manifest.json.
Dispatch Analysis Batches
The sub-agent receives input structured as:
{
"name": "article-analyzer",
"input": {
"batchId": 1,
"articles": [
{
"id": "a1",
"name": "Intro",
"summary": "...",
"wikilinks": ["b2"],
"category": "Basics",
"content": "..."
}
],
"existingNodeIds": ["a1","b2","t1"]
}
}
Results write to analysis-batch-1.json in the intermediate folder.
Merge and Finalize
python3 understand-anything-plugin/skills/understand-knowledge/merge-knowledge-graph.py /path/to/wiki
Produces assembled-graph.json with deduplicated entities and normalized layers.
Launch the Dashboard
/understand-dashboard /path/to/wiki
The dashboard reads .understand-anything/knowledge-graph.json and renders the force-directed visualization with community clustering.
Summary
parse-knowledge-base.pyperforms deterministic structural extraction, creatingscan-manifest.jsonwith explicit wikilinks and categoriesarticle-analyzer.mddefines LLM agents that process articles in batches of 10-15 to infer implicit semantic relationshipsmerge-knowledge-graph.pycombines deterministic and AI-generated data, deduplicating entities case-insensitively and normalizing node typesgraph-builder.tsvalidates the final graph structure before persistingknowledge-graph.jsonand triggering the dashboard- Parallel processing of up to three batches simultaneously optimizes performance for large wiki repositories
Frequently Asked Questions
What is the "Karpathy-pattern" mentioned in the LLM wiki format?
The Karpathy-pattern refers to a markdown-based note-taking structure popularized by Andrej Karpathy, characterized by interlinked markdown files using [[wikilink]] syntax, hierarchical organization through index.md files with categorized sections, and a raw/ folder for source materials. The parser specifically recognizes these conventions to build the initial graph topology.
How does the system handle duplicate entities across different articles?
During Phase 4, merge-knowledge-graph.py performs case-insensitive deduplication to identify when the same entity appears with different capitalizations or slight variations across articles. The merge script normalizes these into single canonical nodes while preserving all incoming relationship edges from the various mentions.
Can I customize the batch size for the LLM analysis phase?
While the default configuration processes 10-15 articles per batch as defined in SKILL.md lines 54-68, the architecture supports batch size adjustments through the skill configuration. However, the three-batch parallel execution limit is designed to balance API rate limits with processing throughput for optimal performance.
What distinguishes knowledge graphs from the standard DAG visualizations in the dashboard?
Knowledge graphs use the kind: "knowledge" flag to trigger a force-directed physics simulation layout in KnowledgeGraphView.tsx, optimized for exploring dense, interconnected conceptual relationships. Standard DAG visualizations enforce hierarchical directionality, while knowledge graphs emphasize community clustering and bidirectional relationship exploration through interactive node dragging and zooming.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →