How the Knowledge Base Analyzer Extracts Entities from Karpathy-Pattern Wikis
The Knowledge Base analyzer extracts entities by parsing markdown headings into knowledge nodes and resolving internal links into wikilink edges using the MarkdownParser and GraphBuilder classes, storing the results in a unified knowledge-graph.json file.
The Egonex-AI/Understand-Anything repository includes a knowledge base analyzer that processes markdown wiki pages following the Karpathy pattern. When the /understand pipeline runs, the core package loads every .md file through the Markdown parser, transforming flat documentation into a structured knowledge graph. This process enables the dashboard to display wiki entities alongside code entities in a unified visualization.
The Five-Stage Entity Extraction Pipeline
1. Section Detection via MarkdownParser
In packages/core/src/plugins/parsers/markdown-parser.ts, the extractSections method (lines 41-76) scans each wiki file line-by-line while ignoring fenced code blocks. The parser records every ATX heading from # to ######, creating a knowledge node for each with kind: "knowledge". The heading text becomes the node's name field, while the lineRange property captures the section's content boundaries.
2. Reference Resolution for Internal Links
The same file implements extractReferences (lines 23-38) to identify markdown links using regular expressions. Only internal links are retained; external URLs are discarded. For each valid link, the parser records the source file, target path, line number, and whether the target is an image or plain file, creating ReferenceResolution objects that track relationships between wiki pages.
3. Graph Node Construction
The GraphBuilder class (located in packages/core/src/analyzer/graph-builder.ts) consumes the sections and references produced by the parser. For every heading detected, it instantiates a GraphNode with type knowledge, attaching optional metadata through the knowledgeMeta field. This field, defined in packages/core/src/types.ts (lines 21-40), can include properties like wikilinks and sourceUrl to enrich the node with contextual information.
4. Wikilink Edge Creation
When a reference points to another markdown file, the builder resolves the target node and creates a directional edge with edgeType: "wikilink". These edges represent semantic relationships between knowledge nodes. The edge categories, including wikilinks, are enumerated in packages/dashboard/src/store.ts (lines 39-52), enabling the dashboard to filter and render these connections specifically.
5. Knowledge Graph Persistence
The final assembly is written to knowledge-graph.json by packages/core/src/persistence/index.ts. This output contains a mixture of code nodes (representing files, classes, and functions) and knowledge nodes derived from wiki headings, interconnected by wikilink edges. The dashboard displays these entities with the label "knowledge" as specified in packages/dashboard/src/locales/en.ts (line 279), describing them as "concept, entity, or claim from a knowledge wiki".
Implementation Code Examples
The following TypeScript examples demonstrate how to interact with the knowledge extraction system programmatically:
// Example: extracting sections from a wiki page
import { MarkdownParser } from '@understand-anything/core';
import { readFile } from 'fs/promises';
const parser = new MarkdownParser();
const content = await readFile('docs/architecture.md', 'utf-8');
const sections = parser.extractSections(content);
// sections now holds each heading with its line range
// Example: creating graph nodes from a markdown file
import { GraphBuilder, MarkdownParser } from '@understand-anything/core';
import { readFile } from 'fs/promises';
const parser = new MarkdownParser();
const file = 'docs/architecture.md';
const src = await readFile(file, 'utf-8');
const analysis = parser.analyzeFile(file, src);
const refs = parser.extractReferences(file, src);
const builder = new GraphBuilder();
builder.addFile(file, analysis.sections, refs);
await builder.finalize(); // writes knowledge-graph.json
// Example: accessing wikilink edges in the dashboard store
import { useStore } from '@/store';
const wikilinks = useStore.getState().edges.filter(e => e.type === 'wikilink');
console.log(wikilinks);
Summary
- The
MarkdownParser.extractSectionsmethod (lines 41-76) converts ATX headings into knowledge nodes withkind: "knowledge". - Internal markdown links are extracted via
extractReferences(lines 23-38) and transformed intowikilinkedges. GraphBuilderconstructs the unified graph by combining knowledge nodes from wikis with code nodes from source files.- The resulting
knowledge-graph.jsonpersists both node types and their relationships for dashboard visualization.
Frequently Asked Questions
What distinguishes a Karpathy-pattern wiki from standard markdown?
Karpathy-pattern wikis organize knowledge through hierarchical ATX headings and extensive internal linking. The analyzer specifically targets these patterns by extracting heading-based sections and resolving markdown links to build navigable webs of connected concepts, as opposed to treating the document as a single monolithic block.
How does the analyzer distinguish between internal and external links?
The extractReferences method in markdown-parser.ts applies validation logic to filter link targets. It retains only relative paths and internal references while discarding absolute URLs (http/https) and external domains, ensuring the knowledge graph contains only project-relevant relationships.
What is the difference between knowledge nodes and code nodes in the graph?
Knowledge nodes (with kind: "knowledge") represent conceptual entities extracted from wiki headings, while code nodes represent structural elements like files, classes, and functions. Both share the same GraphNode interface defined in types.ts but differ in their kind property and metadata schemas.
Can the analyzer handle deeply nested heading structures?
Yes. The extractSections method processes all six levels of ATX headings (# through ######), creating hierarchical knowledge nodes. Each heading level becomes a distinct node with its own lineRange spanning the section content until the next heading or file end.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →