# How the Knowledge Base Analyzer Extracts Entities from Karpathy-Pattern Wikis

> Learn how the Knowledge Base analyzer extracts entities from Karpathy-pattern wikis by parsing markdown and resolving internal links into a knowledge graph.

- Repository: [Egonex/Understand-Anything](https://github.com/Egonex-AI/Understand-Anything)
- Tags: how-to-guide
- Published: 2026-06-18

---

**The Knowledge Base analyzer extracts entities by parsing markdown headings into knowledge nodes and resolving internal links into wikilink edges using the `MarkdownParser` and `GraphBuilder` classes, storing the results in a unified [`knowledge-graph.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/knowledge-graph.json) file.**

The Egonex-AI/Understand-Anything repository includes a knowledge base analyzer that processes markdown wiki pages following the Karpathy pattern. When the `/understand` pipeline runs, the core package loads every `.md` file through the **Markdown parser**, transforming flat documentation into a structured knowledge graph. This process enables the dashboard to display wiki entities alongside code entities in a unified visualization.

## The Five-Stage Entity Extraction Pipeline

### 1. Section Detection via MarkdownParser

In [`packages/core/src/plugins/parsers/markdown-parser.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/plugins/parsers/markdown-parser.ts), the `extractSections` method (lines 41-76) scans each wiki file line-by-line while ignoring fenced code blocks. The parser records every ATX heading from `#` to `######`, creating a **knowledge node** for each with `kind: "knowledge"`. The heading text becomes the node's `name` field, while the `lineRange` property captures the section's content boundaries.

### 2. Reference Resolution for Internal Links

The same file implements `extractReferences` (lines 23-38) to identify markdown links using regular expressions. **Only internal links are retained**; external URLs are discarded. For each valid link, the parser records the source file, target path, line number, and whether the target is an image or plain file, creating `ReferenceResolution` objects that track relationships between wiki pages.

### 3. Graph Node Construction

The `GraphBuilder` class (located in [`packages/core/src/analyzer/graph-builder.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/analyzer/graph-builder.ts)) consumes the sections and references produced by the parser. For every heading detected, it instantiates a `GraphNode` with type **knowledge**, attaching optional metadata through the `knowledgeMeta` field. This field, defined in [`packages/core/src/types.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/types.ts) (lines 21-40), can include properties like `wikilinks` and `sourceUrl` to enrich the node with contextual information.

### 4. Wikilink Edge Creation

When a reference points to another markdown file, the builder resolves the target node and creates a directional edge with `edgeType: "wikilink"`. These edges represent semantic relationships between knowledge nodes. The edge categories, including `wikilinks`, are enumerated in [`packages/dashboard/src/store.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/dashboard/src/store.ts) (lines 39-52), enabling the dashboard to filter and render these connections specifically.

### 5. Knowledge Graph Persistence

The final assembly is written to [`knowledge-graph.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/knowledge-graph.json) by [`packages/core/src/persistence/index.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/persistence/index.ts). This output contains a mixture of **code nodes** (representing files, classes, and functions) and **knowledge nodes** derived from wiki headings, interconnected by wikilink edges. The dashboard displays these entities with the label "knowledge" as specified in [`packages/dashboard/src/locales/en.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/dashboard/src/locales/en.ts) (line 279), describing them as "concept, entity, or claim from a knowledge wiki".

## Implementation Code Examples

The following TypeScript examples demonstrate how to interact with the knowledge extraction system programmatically:

```typescript
// Example: extracting sections from a wiki page
import { MarkdownParser } from '@understand-anything/core';
import { readFile } from 'fs/promises';

const parser = new MarkdownParser();
const content = await readFile('docs/architecture.md', 'utf-8');
const sections = parser.extractSections(content);
// sections now holds each heading with its line range

```

```typescript
// Example: creating graph nodes from a markdown file
import { GraphBuilder, MarkdownParser } from '@understand-anything/core';
import { readFile } from 'fs/promises';

const parser = new MarkdownParser();
const file = 'docs/architecture.md';
const src = await readFile(file, 'utf-8');

const analysis = parser.analyzeFile(file, src);
const refs = parser.extractReferences(file, src);

const builder = new GraphBuilder();
builder.addFile(file, analysis.sections, refs);
await builder.finalize();   // writes knowledge-graph.json

```

```typescript
// Example: accessing wikilink edges in the dashboard store
import { useStore } from '@/store';

const wikilinks = useStore.getState().edges.filter(e => e.type === 'wikilink');
console.log(wikilinks);

```

## Summary

- The `MarkdownParser.extractSections` method (lines 41-76) converts ATX headings into knowledge nodes with `kind: "knowledge"`.
- Internal markdown links are extracted via `extractReferences` (lines 23-38) and transformed into `wikilink` edges.
- `GraphBuilder` constructs the unified graph by combining knowledge nodes from wikis with code nodes from source files.
- The resulting [`knowledge-graph.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/knowledge-graph.json) persists both node types and their relationships for dashboard visualization.

## Frequently Asked Questions

### What distinguishes a Karpathy-pattern wiki from standard markdown?

Karpathy-pattern wikis organize knowledge through hierarchical ATX headings and extensive internal linking. The analyzer specifically targets these patterns by extracting heading-based sections and resolving markdown links to build navigable webs of connected concepts, as opposed to treating the document as a single monolithic block.

### How does the analyzer distinguish between internal and external links?

The `extractReferences` method in [`markdown-parser.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/markdown-parser.ts) applies validation logic to filter link targets. It retains only relative paths and internal references while discarding absolute URLs (http/https) and external domains, ensuring the knowledge graph contains only project-relevant relationships.

### What is the difference between knowledge nodes and code nodes in the graph?

Knowledge nodes (with `kind: "knowledge"`) represent conceptual entities extracted from wiki headings, while code nodes represent structural elements like files, classes, and functions. Both share the same `GraphNode` interface defined in [`types.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/types.ts) but differ in their `kind` property and metadata schemas.

### Can the analyzer handle deeply nested heading structures?

Yes. The `extractSections` method processes all six levels of ATX headings (`#` through `######`), creating hierarchical knowledge nodes. Each heading level becomes a distinct node with its own `lineRange` spanning the section content until the next heading or file end.