How /understand-knowledge Analyzes LLM Wikis into Knowledge Graphs: A 5-Phase Technical Breakdown

The /understand-knowledge skill converts Karpathy-pattern LLM wikis into interactive force-directed knowledge graphs through a deterministic five-phase pipeline that combines Python-based structural parsing with LLM-driven semantic analysis.

The Lum1104/Understand-Anything repository provides a Claude Code plugin that transforms markdown-based LLM wikis into structured knowledge graphs. The /understand-knowledge skill orchestrates this conversion through a hybrid architecture that merges deterministic file parsing with parallel LLM agent processing to extract both explicit wikilinks and implicit conceptual relationships.

The Five-Phase Pipeline

The skill processes wiki directories through five deterministic phases, each implemented by specialized Python scripts and sub-agents defined in understand-anything-plugin/skills/understand-knowledge/SKILL.md.

Phase 1: Detect and Parse

The pipeline begins by determining the target directory and executing the deterministic parser parse-knowledge-base.py. This script scans the wiki structure and writes a scan-manifest.json file that describes raw sources, wiki pages, topics, and wikilinks. According to SKILL.md lines 30-34, this phase establishes the foundational structure without LLM involvement, ensuring a baseline graph exists even if subsequent AI processing fails.

Phase 2: Scan and Structure

Phase 2 operates entirely within the manifest generation completed by Phase 1. The scan-manifest.json contains:

  • Article nodes: One per markdown file with extracted wikilinks, headings, and front-matter
  • Source nodes: Files located under the raw/ directory
  • Topic nodes: Extracted from index.md headings
  • Edges: related links (wikilinks) and categorized_under relationships (index sections)

As documented in SKILL.md lines 43-48, this manifest captures the explicit structural connections present in the markdown files.

Phase 3: Analyze with LLM Agents

The system dispatches the article-analyzer sub-agent (defined in understand-anything-plugin/agents/article-analyzer.md) to process batches of 10-15 articles concurrently. Each batch receives the article metadata plus the complete node-ID list, enabling the LLM to infer implicit relationships, entities, and claims that exist conceptually but not explicitly as wikilinks. The orchestration steps in SKILL.md lines 54-68 describe how up to three batches run in parallel to minimize latency for large wikis.

Each batch writes its results to analysis-batch-{N}.json files containing enriched semantic connections.

Phase 4: Merge and Normalize

The merge-knowledge-graph.py script (referenced in SKILL.md lines 74-85) consolidates the deterministic manifest with the LLM-generated analysis files. This phase performs:

  1. Entity deduplication: Case-insensitive matching to eliminate redundant nodes
  2. Type normalization: Standardizing article, source, topic, entity, and claim node types
  3. Layer construction: Building hierarchical layers from index.md categories
  4. Tour ordering: Creating navigational sequences through the graph

The output is assembled-graph.json, a unified representation combining explicit structural links with AI-inferred semantic connections.

Phase 5: Validate and Persist

The final phase validates the assembled graph using logic from understand-anything-plugin/packages/core/src/analyzer/graph-builder.ts. Validation checks that every edge's source and target nodes exist and that all required node fields are present. As detailed in SKILL.md lines 91-125, the system then:

  • Writes the final knowledge-graph.json to .<understand-anything>/
  • Persists a small meta.json with processing metadata
  • Cleans up the intermediate folder
  • Auto-triggers /understand-dashboard to render the visualization

Architecture Highlights

Deterministic Pre-processing

The parse-knowledge-base.py script extracts only structural information—wikilinks, headings, front-matter, and categories—providing a guaranteed baseline graph independent of LLM availability or reliability.

LLM-Driven Enrichment

The article-analyzer agents receive the deterministic graph context, allowing them to add implicit edges such as entity-to-entity relationships or claim-based links (implies edges) that markdown syntax alone cannot express.

Layered Graph Model

Nodes carry explicit types (article, source, topic, entity, claim) while edges encode relationship semantics (related, categorized_under, implies). The kind: "knowledge" flag signals the dashboard to render a force-directed layout rather than the default hierarchical DAG visualization implemented in KnowledgeGraphView.tsx.

Parallel Execution

The batch processing architecture allows up to three analysis batches to run concurrently, significantly reducing processing time for wikis containing hundreds of articles.

Running the Pipeline

Execute the Complete Skill

/understand-knowledge /path/to/wiki

If no path is provided, the command uses the current working directory. This triggers all five phases sequentially.

Run the Deterministic Parser Manually

python3 understand-anything-plugin/skills/understand-knowledge/parse-knowledge-base.py /path/to/wiki

Output writes to /path/to/wiki/.understand-anything/intermediate/scan-manifest.json.

Dispatch Analysis Batches

The sub-agent receives input structured as:

{
  "name": "article-analyzer",
  "input": {
    "batchId": 1,
    "articles": [
      { 
        "id": "a1", 
        "name": "Intro", 
        "summary": "...", 
        "wikilinks": ["b2"], 
        "category": "Basics", 
        "content": "..."
      }
    ],
    "existingNodeIds": ["a1","b2","t1"]
  }
}

Results write to analysis-batch-1.json in the intermediate folder.

Merge and Finalize

python3 understand-anything-plugin/skills/understand-knowledge/merge-knowledge-graph.py /path/to/wiki

Produces assembled-graph.json with deduplicated entities and normalized layers.

Launch the Dashboard

/understand-dashboard /path/to/wiki

The dashboard reads .understand-anything/knowledge-graph.json and renders the force-directed visualization with community clustering.

Summary

  • parse-knowledge-base.py performs deterministic structural extraction, creating scan-manifest.json with explicit wikilinks and categories
  • article-analyzer.md defines LLM agents that process articles in batches of 10-15 to infer implicit semantic relationships
  • merge-knowledge-graph.py combines deterministic and AI-generated data, deduplicating entities case-insensitively and normalizing node types
  • graph-builder.ts validates the final graph structure before persisting knowledge-graph.json and triggering the dashboard
  • Parallel processing of up to three batches simultaneously optimizes performance for large wiki repositories

Frequently Asked Questions

What is the "Karpathy-pattern" mentioned in the LLM wiki format?

The Karpathy-pattern refers to a markdown-based note-taking structure popularized by Andrej Karpathy, characterized by interlinked markdown files using [[wikilink]] syntax, hierarchical organization through index.md files with categorized sections, and a raw/ folder for source materials. The parser specifically recognizes these conventions to build the initial graph topology.

How does the system handle duplicate entities across different articles?

During Phase 4, merge-knowledge-graph.py performs case-insensitive deduplication to identify when the same entity appears with different capitalizations or slight variations across articles. The merge script normalizes these into single canonical nodes while preserving all incoming relationship edges from the various mentions.

Can I customize the batch size for the LLM analysis phase?

While the default configuration processes 10-15 articles per batch as defined in SKILL.md lines 54-68, the architecture supports batch size adjustments through the skill configuration. However, the three-batch parallel execution limit is designed to balance API rate limits with processing throughput for optimal performance.

What distinguishes knowledge graphs from the standard DAG visualizations in the dashboard?

Knowledge graphs use the kind: "knowledge" flag to trigger a force-directed physics simulation layout in KnowledgeGraphView.tsx, optimized for exploring dense, interconnected conceptual relationships. Standard DAG visualizations enforce hierarchical directionality, while knowledge graphs emphasize community clustering and bidirectional relationship exploration through interactive node dragging and zooming.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →