How the /understand-knowledge Endpoint Analyzes Karpathy-Pattern LLM Wikis

The /understand-knowledge endpoint transforms Karpathy-pattern LLM wikis into interactive knowledge graphs through a five-phase pipeline that combines deterministic parsing with LLM-driven semantic analysis.

The Egonex-AI/Understand-Anything repository provides a skill that ingests three-layer knowledge bases—consisting of raw sources, markdown wiki pages, and schema files—and produces a structured knowledge graph. When you invoke the /understand-knowledge endpoint, it executes a deterministic parsing phase followed by LLM-powered semantic enrichment to generate the final graph used by the dashboard.

Phase 1: Format Detection

The pipeline begins by verifying the directory structure adheres to the Karpathy-pattern specification. The detect_format() function in understand-anything-plugin/skills/understand-knowledge/parse-knowledge-base.py (lines 38-64) examines the target directory for the three-layer structure: an index.md file, a sufficient number of .md files, and optional raw/ and schema directories.

This deterministic detector returns a JSON manifest describing the wiki layout, ensuring downstream phases only process valid knowledge bases.

Phase 2: Deterministic Parsing and Node Extraction

Once validated, the parse_wiki() function (lines 73-78, 124-191, and 227-278 of parse-knowledge-base.py) performs a comprehensive scan without LLM involvement. This phase extracts:

  • Front-matter metadata and heading structures
  • Wikilinks in the format [[target]]
  • Code-block languages and first paragraphs
  • Category relationships from index.md sections

The parser constructs article nodes, topic nodes derived from index.md sections, and source nodes from the raw/ directory. It generates edges for related connections (from wikilinks) and categorized_under relationships. The system also computes backlinks for each article (lines 50-61 of parse_wiki) to support bidirectional navigation.

This phase produces scan-manifest.json, containing the complete deterministic graph structure.

Phase 3: LLM-Driven Semantic Analysis

With the structural manifest complete, the skill batches articles (approximately 10-15 per batch) and dispatches specialized article-analyzer sub-agents. According to SKILL.md (lines 52-70), these agents process the raw content to extract:

  • Named entities not captured by pattern matching
  • Verifiable claims within the text
  • Implicit relationships undetectable through static parsing

Each batch writes its results to analysis-batch-N.json files. This separation ensures that LLM variability affects only semantic enrichment, not the stable structural backbone.

Phase 4: Knowledge Graph Merging

The merge-knowledge-graph.py script combines deterministic structure with LLM analysis through its merge() function. This phase spans three critical operations:

Deduplication and Normalization (lines 92-176): The system normalizes node types using the NODE_TYPE_ALIASES dictionary (lines 50-66), mapping variants like note, page, or tag to canonical types (article, entity, topic). It performs case-insensitive name matching to consolidate duplicate entities.

Layer Construction (lines 186-236): Categories from index.md become topic nodes driving the hierarchical layer:<slug> structure.

Tour Generation (lines 260-322): The system assembles a guided tour based on category ordering, creating walkthrough steps with order, title, and nodeIds fields.

The final assembled-graph.json includes validated edges ensuring every reference points to an existing node.

Phase 5: Persistence and Dashboard Rendering

The completed graph is copied to <wiki-dir>/.understand-anything/knowledge-graph.json as documented in SKILL.md (lines 91-118). Validation ensures structural integrity before the skill auto-triggers the /understand-dashboard endpoint.

The dashboard renders the graph using a force-directed layout with community clustering, consuming the JSON file via client-side fetch operations.

Practical Usage Examples

Execute the complete pipeline on a Karpathy-pattern wiki:

/understand-knowledge ~/my-wiki

Run the deterministic parser independently to generate the structural manifest:

python ./understand-anything-plugin/skills/understand-knowledge/parse-knowledge-base.py ~/my-wiki

# Creates scan-manifest.json in ~/.understand-anything/intermediate/

Merge deterministic structure with LLM analysis batches manually:

python ./understand-anything-plugin/skills/understand-knowledge/merge-knowledge-graph.py ~/my-wiki

# Produces assembled-graph.json and copies to ~/my-wiki/.understand-anything/knowledge-graph.json

Load the graph in a custom dashboard implementation:

fetch("/file-content.json?path=.understand-anything/knowledge-graph.json")
  .then(r => r.json())
  .then(graph => renderKnowledgeGraph(graph));

Key Architectural Principles

Deterministic First Pass: All structural extraction in parse-knowledge-base.py uses pure Python with regular expressions, guaranteeing reproducible outputs regardless of LLM temperature or version changes.

LLM-Only for Semantics: The article-analyzer agents (specified in understand-anything-plugin/agents/article-analyzer.md) handle only entity extraction and implicit relationship detection, keeping the core graph stable while adding semantic depth.

Alias Resolution: The merge script maintains strict type consistency through canonical mappings, preventing fragmentation where different authors use varying tags for equivalent concepts.

Summary

  • The /understand-knowledge endpoint processes Karpathy-pattern wikis through five distinct phases: Detection, Scanning, Analysis, Merging, and Display.
  • Deterministic parsing in parse-knowledge-base.py extracts structure, wikilinks, and categories without LLM involvement.
  • LLM sub-agents enrich the graph with entities, claims, and implicit relationships in batched operations.
  • Graph merging in merge-knowledge-graph.py deduplicates nodes, normalizes types via NODE_TYPE_ALIASES, and constructs navigable layers and guided tours.
  • Final output is written to .understand-anything/knowledge-graph.json and auto-rendered by the dashboard.

Frequently Asked Questions

What defines a Karpathy-pattern LLM wiki?

A Karpathy-pattern wiki consists of three distinct layers: a raw/ directory containing source materials, markdown files with wikilink syntax ([[target]]) representing the knowledge base, and schema files including index.md that defines categories and structure. The /understand-knowledge endpoint specifically validates this structure during Phase 1 detection before processing.

Why does the system separate deterministic parsing from LLM analysis?

The separation ensures reproducibility and stability. Structural elements like headings, wikilinks, and categories are extracted via deterministic regex in parse-knowledge-base.py, guaranteeing identical outputs across runs. LLM variability is isolated to semantic enrichment (entity extraction and implicit relationships), allowing the knowledge graph foundation to remain constant while supporting iterative semantic refinement.

How does the endpoint handle duplicate entities across batches?

During Phase 4, the merge() function in merge-knowledge-graph.py performs case-insensitive name matching to identify duplicates. It normalizes node types using the NODE_TYPE_ALIASES dictionary (lines 50-66), consolidating variants like note or page into canonical types (article, entity). Edge targets are validated to ensure every reference resolves to an existing node.

Can I execute individual pipeline phases without running the full skill?

Yes. You can manually invoke the deterministic parser to generate scan-manifest.json using parse-knowledge-base.py, and separately run merge-knowledge-graph.py to combine the manifest with LLM analysis batches. This modular approach supports debugging individual phases or integrating specific components into custom workflows.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →