How the Knowledge Base Analyzer Processes Karpathy-Pattern Wikis into Knowledge Graphs
The knowledge base analyzer converts Markdown wikis into structured knowledge graphs by parsing ATX headings and local references into hierarchical nodes, enriching them with LLM-generated wikilinks and semantic metadata, and validating the output against strict schemas before persisting to knowledge-graph.json.
The Understand-Anything repository provides a sophisticated pipeline that transforms unstructured Markdown documentation—commonly known as "Karpathy-pattern" wikis after the style popularized by Andrej Karpathy—into fully-typed entities within a unified knowledge graph. This system treats wiki pages as first-class citizens alongside source code, enabling cross-referencing between prose explanations and implementation details.
Step 1: Language Registration and File Discovery
The ingestion process begins in packages/core/src/languages/language-registry.ts, where the LanguageRegistry maintains a mapping between file extensions and analysis plugins. When the project scanner encounters files ending in .md, it routes them to the Markdown language handler, distinguishing them from source code files that require syntactic parsing. This registration ensures that wiki files trigger the specialized MarkdownParser rather than falling through to generic binary file handling.
Step 2: Structural Parsing with MarkdownParser
Located at understand-anything-plugin/packages/core/src/plugins/parsers/markdown-parser.ts, the MarkdownParser performs two critical extraction operations that transform flat text into structured data.
Extracting Hierarchical Sections
The extractSections method scans for ATX headings (# through ######) and generates SectionInfo objects containing the heading name, hierarchy level, and precise line ranges. These objects preserve the document's outline structure, capturing the semantic nesting of concepts without processing the actual sentence content.
Resolving Internal References
Using extractReferences, the parser identifies inline links ([text](target)) and creates ReferenceResolution records that capture the source file path, target path, reference type (distinguishing between local files and images), and the exact line number. The parser deliberately ignores content inside code fences and filters out external URLs, focusing exclusively on internal wiki navigation paths that contribute to the knowledge graph topology.
import { MarkdownParser } from "./plugins/parsers/markdown-parser";
const parser = new MarkdownParser();
const content = await Deno.readTextFile("docs/intro.md");
// Structural analysis (sections only)
const analysis = parser.analyzeFile("docs/intro.md", content);
console.log(analysis.sections);
// Reference resolution (local file and image links)
const refs = parser.extractReferences("docs/intro.md", content);
console.log(refs);
Step 3: Constructing Graph Nodes via GraphBuilder
The GraphBuilder class in packages/core/src/analyzer/graph-builder.ts ingests parsed wiki data through the addNonCodeFileWithAnalysis method. This function creates a file-type node (type: "file") that stores the file path, a generated summary, classification tags, and a complexity rating. Each extracted section spawns a child node connected by contains edges, transforming the linear markdown into a hierarchical graph structure where headings become distinct navigable entities.
import { GraphBuilder } from "./analyzer/graph-builder";
const gb = new GraphBuilder("MyProject", "a1b2c3d");
// Assume we already have `analysis` from the parser
gb.addNonCodeFileWithAnalysis("docs/intro.md", {
nodeType: "article", // LLM may decide the canonical type
summary: "Overview of the project",
tags: ["intro", "overview"],
complexity: "simple",
sections: analysis.sections, // sections become child nodes
// KnowledgeMeta will be filled later by the LLM analyzer
});
Step 4: LLM-Driven Semantic Enrichment
After structural nodes exist, the llm-analyzer (located in packages/core/src/analyzer/llm-analyzer.ts) processes the raw markdown text to generate semantic metadata defined by KnowledgeMetaSchema in packages/core/src/schema.ts (lines 61-66).
Generating Wikilinks and Knowledge Relationships
The LLM returns a structured payload containing wikilinks (array of referenced wiki pages), optional backlinks, category classification, and full content summaries. The analyzer writes this metadata into the node's knowledgeMeta field and instantiates semantic edges—such as cites, builds_on, and exemplifies (defined in EDGE_TYPE_ALIASES)—connecting related concepts across the documentation corpus.
import { llmAnalyze } from "./analyzer/llm-analyzer";
const wikiNode = gb.nodes.find(n => n.filePath === "docs/intro.md");
if (wikiNode) {
const llmResult = await llmAnalyze(wikiNode.summary);
// llmResult.wikilinks is an array like ["docs/architecture.md", "docs/usage.md"]
wikiNode.knowledgeMeta = {
wikilinks: llmResult.wikilinks,
category: "documentation",
content: llmResult.fullText,
};
}
Step 5: Validation and Persistence
Before serialization, the validateGraph function (in packages/core/src/schema.ts, lines 498-562) sanitizes the data, resolves NODE_TYPE_ALIASES, and auto-fixes missing fields to ensure strict conformance with GraphNodeSchema. This validator guarantees that wiki-derived nodes include the required wikilinks array and that any missing edges are automatically added as cites-type relationships.
The final KnowledgeGraph object produced by GraphBuilder.build is written to knowledge-graph.json via packages/core/src/persistence/index.ts. The dashboard consumes this file to render unified visualizations where wiki nodes appear alongside code symbols, enabling navigation across both documentation and implementation.
import { validateGraph } from "./schema";
import { writeFile } from "fs/promises";
const graph = gb.build();
const { success, data, issues } = validateGraph(graph);
if (success && data) {
await writeFile("./.understand-anything/knowledge-graph.json", JSON.stringify(data, null, 2));
}
Summary
- The
LanguageRegistryroutes.mdfiles to the specializedMarkdownParser, treating wikis as distinct from source code. MarkdownParser.extractSectionsandextractReferencescapture document hierarchy and internal links while ignoring code fences and external URLs.GraphBuilder.addNonCodeFileWithAnalysiscreates file nodes andcontainsedges for sections, establishing the structural backbone.- The
llm-analyzerpopulatesKnowledgeMetaSchemawithwikilinksand semantic categories, generating knowledge edges likebuilds_onandexemplifies. validateGraphensures schema compliance throughNODE_TYPE_ALIASESresolution and auto-fixing, persisting the result toknowledge-graph.json.
Frequently Asked Questions
What is a Karpathy-pattern wiki?
A Karpathy-pattern wiki refers to Markdown-based documentation organized with clear ATX headings (# Title, ## Section) and dense internal linking via relative paths, popularized by Andrej Karpathy's educational repositories. The analyzer specifically targets this structure because it emphasizes hierarchical organization and cross-referencing between concepts.
How does the analyzer distinguish between code files and wiki files?
The system relies on the LanguageRegistry in language-registry.ts to map file extensions to language types. Files ending in .md are tagged as the markdown language and routed to the MarkdownParser, while .ts, .js, .py, and other extensions trigger language-specific code parsers that extract classes, functions, and imports rather than headings and prose sections.
What types of relationships exist between wiki nodes?
The graph creates two primary relationship categories: structural edges (contains) generated by the GraphBuilder to represent heading hierarchies, and semantic edges (cites, builds_on, exemplifies) generated by the LLM analyzer based on EDGE_TYPE_ALIASES. These relationships allow the system to distinguish between a document simply including a section and a concept semantically depending on another concept.
Can the analyzer process external URLs in markdown?
No. The MarkdownParser deliberately filters out external URLs during the extractReferences phase, focusing exclusively on local file and image references that can be resolved to nodes within the project's knowledge graph. This design choice ensures that the graph remains self-contained and represents only navigable, project-local knowledge.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →