How to Analyze a Wiki with Understand Anything: From Markdown Files to an Interactive Knowledge Graph
Understand Anything converts a Karpathy-pattern LLM wiki into an interactive knowledge graph by detecting Markdown files with [[wikilinks]], extracting front-matter and headings, and merging explicit links with implicit LLM-generated relationships.
To analyze a wiki with Understand Anything, you run the knowledge-base parser against a directory that follows the Karpathy wiki convention—an index.md table of contents plus a collection of interlinked Markdown articles. The Lum1104/Understand-Anything repository provides the parse-knowledge-base.py and merge-knowledge-graph.py scripts that transform those files into a unified graph you can explore in the web dashboard, and the full skill specification is documented in understand-anything-plugin/skills/understand-knowledge/SKILL.md.
How the Parser Detects a Karpathy-Pattern Wiki
The pipeline begins in understand-anything-plugin/skills/understand-knowledge/parse-knowledge-base.py. This script identifies a Karpathy wiki by scanning for an index.md (or wiki/index.md), enforcing a minimum number of .md files, and checking for optional raw/ and schema files. When these signals align, the directory is flagged as a wiki knowledge base and the extraction phase starts.
Extracting Front-Matter, Headings, and Wikilinks
Once detected, every Markdown file is processed for three key data sources:
- Front-matter — Parsed with a lightweight YAML regex at the top of each article.
- Wikilinks — Explicit links in the form
[[target]]or[[target|display]]are captured by theWIKILINK_REpattern defined early in the parser. - Headings — The category structure is read from
index.md, where each section begins with a level-two heading (##).
These extractions supply the raw material for the nodes and edges that form the knowledge graph.
Creating Typed Nodes for Wiki Articles
In the core schema defined in understand-anything-plugin/packages/core/src/schema.ts, every parsed Markdown article becomes a node with type: "wiki_page". Each node stores:
name— The file stem (for example,my-articlefrommy-article.md)content— The raw article textwikilinks— An array of explicit links extracted from the bodycategory— The section heading derived fromindex.md
This typed representation ensures that wiki concepts live alongside code entities—files, classes, and functions—inside the same graph.
Building Explicit and Implicit Edges
The graph connects wiki articles through two edge strategies:
- Explicit edges (
related) — Generated directly from each wikilink found in the article body. Ifarticle-a.mdcontains[[article-b]], the parser creates arelatededge between the two nodes. - Implicit edges — Added later by the article-analyzer agent. This agent reads the full article text and injects entities, claims, and relationships that are not already covered by an explicit wikilink.
This two-layer approach captures both the deliberate link structure of the wiki and the latent semantic connections extracted by the LLM.
Merging and Assembling the Knowledge Graph
After parsing, understand-anything-plugin/skills/understand-knowledge/merge-knowledge-graph.py combines the article nodes with any additional knowledge produced by the LLM agents. The script writes the result to .understand-anything/intermediate/assembled-graph.json inside the project folder. This file is the canonical graph representation consumed by the dashboard and the core query API.
Visualizing Wikilinks in the Dashboard
When you inspect a wiki node in the web UI, the NodeInfo component renders its outgoing wikilinks as a dedicated field (for example, “Wikilinks (3)”). The dashboard’s language packs label this field consistently across locales, making it easy to spot densely linked articles and navigate the conceptual structure of the wiki.
How to Run the Pipeline to Analyze a Wiki with Understand Anything
To analyze a wiki with Understand Anything, run the two-stage pipeline from the root of the target repository.
Step 1 — Detect and parse the wiki.
python understand-anything-plugin/skills/understand-knowledge/parse-knowledge-base.py ./my-wiki
This creates .understand-anything/intermediate/scan-manifest.json.
Step 2 — Merge explicit and implicit knowledge.
python understand-anything-plugin/skills/understand-knowledge/merge-knowledge-graph.py ./my-wiki
This produces .understand-anything/intermediate/assembled-graph.json.
After both steps finish, start the dashboard:
pnpm dev:dashboard
Open the web UI and the graph will show both code entities and wiki concepts—articles, wikilinks, and inferred relationships—inside the same interactive view.
If you prefer to query the graph programmatically, load the assembled output via the core TypeScript package:
import { loadGraph } from '@understand-anything/core';
const graph = await loadGraph('./my-wiki/.understand-anything/intermediate/assembled-graph.json');
const article = graph.nodes.find(n => n.type === 'wiki_page' && n.name === 'my-article');
console.log('Wikilinks:', article?.wikilinks);
Summary
- Wiki detection relies on
parse-knowledge-base.pyscanning forindex.md, a minimum.mdfile count, and optionalraw/and schema files. - Extraction captures front-matter,
[[wikilinks]], and##headings from the wiki source. - Nodes are typed as
wiki_pagein the core schema and carryname,content,wikilinks, andcategory. - Edges come from explicit wikilinks (
related) and from implicit relationships added by the article-analyzer agent. - Assembly happens in
merge-knowledge-graph.py, which outputsassembled-graph.jsonfor the dashboard. - Visualization renders wikilink counts in the dashboard’s
NodeInfocomponent.
Frequently Asked Questions
What file structure does Understand Anything need to detect a wiki?
The parser expects a Karpathy-pattern wiki: an index.md (or wiki/index.md) serving as a table of contents, a minimum number of .md files, and optional raw/ or schema files. When these signals are present, parse-knowledge-base.py flags the folder as a wiki and begins extraction.
What is the difference between explicit and implicit edges in the wiki graph?
Explicit edges are generated from wikilinks written directly in the Markdown, such as [[target]]. Implicit edges are created by the article-analyzer agent, which reads the article text and adds semantic relationships, entities, and claims that the author did not explicitly link. Together they produce a dense, navigable knowledge graph.
Can I query the wiki graph without using the dashboard?
Yes. After merge-knowledge-graph.py writes assembled-graph.json, you can load it programmatically using @understand-anything/core. Filter for type === 'wiki_page' to access article nodes and their wikilinks arrays directly in TypeScript or JavaScript.
Where are the intermediate files stored during wiki analysis?
Both scan-manifest.json and assembled-graph.json are written to the .understand-anything/intermediate/ directory inside the target wiki project. These files are consumed by the dashboard and the core graph-loading utilities.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →