# How the Knowledge Base Analyzer Extracts Entities from Wiki Articles in Egonex-AI

> Discover how the Egonex-AI knowledge base analyzer extracts entities from wiki articles using a three-stage pipeline: parsing, LLM recognition, and deduplication.

- Repository: [Egonex/Understand-Anything](https://github.com/Egonex-AI/Understand-Anything)
- Tags: how-to-guide
- Published: 2026-06-09

---

**The Understand-Anything knowledge base analyzer extracts entities from wiki articles through a three-stage pipeline: deterministic parsing of markdown files and explicit wikilinks, LLM-driven recognition of implicit entities mentioned in text but lacking dedicated pages, and a final deduplication pass that normalizes IDs and merges duplicates into the knowledge graph.**

The Egonex-AI/Understand-Anything project transforms unstructured markdown wiki articles into structured knowledge graphs. Its **entity extraction** pipeline specifically targets both explicit references found in `[[wikilink]]` syntax and implicit mentions of people, tools, and organizations buried in article prose.

## The Three-Stage Entity Extraction Pipeline

### Stage 1: Deterministic Wiki Parsing with [`parse-knowledge-base.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/parse-knowledge-base.py)

The pipeline begins in [`understand-anything-plugin/skills/understand-knowledge/parse-knowledge-base.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/skills/understand-knowledge/parse-knowledge-base.py), which implements deterministic parsing of Karpathy-pattern wikis. The `detect_format` function (lines 38-66) identifies a valid wiki by verifying the presence of [`index.md`](https://github.com/Egonex-AI/Understand-Anything/blob/main/index.md) and a minimum threshold of markdown files.

Once validated, the parser iterates through every markdown file to extract front-matter, headings, and explicit wikilinks using the `extract_wikilinks` helper (lines 89-97). Each article becomes a structured object containing `id`, `name`, `summary`, `wikilinks`, `category`, and `truncated content`. The parser emits these objects into a **scan-manifest.json** file (lines 278-301), which serves as the input for the next stage.

### Stage 2: LLM-Driven Implicit Entity Recognition

The *Article Analyzer* agent processes the scan manifest to surface entities mentioned in article text that do not have their own wiki pages. According to the prompt defined in [`understand-anything-plugin/agents/article-analyzer.md`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/agents/article-analyzer.md) (lines 27-35), the LLM creates `entity` nodes with a normalized ID, name, short summary, and tags.

**Entity ID normalization** follows strict rules: convert to lower-case and replace spaces with hyphens. For example, *"Andrej Karpathy"* becomes `entity:andrej-karpathy`. The LLM explicitly filters out any entities that already exist as articles or wikilinks in the manifest, ensuring only new, implicit entities are returned.

### Stage 3: Deduplication and Graph Integration via [`merge-knowledge-graph.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/merge-knowledge-graph.py)

The final stage occurs in [`understand-anything-plugin/skills/understand-knowledge/merge-knowledge-graph.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/skills/understand-knowledge/merge-knowledge-graph.py). The `normalize_entity_name` function (lines 83-84) ensures consistent ID formatting across all candidate entities. A deduplication loop (lines 135-142) collapses duplicate IDs and updates edge references, integrating the new entity nodes into the existing knowledge graph without duplication.

## Implementation Example

The following TypeScript demonstrates how to orchestrate the three-stage extraction flow:

```typescript
// 1️⃣ Run the deterministic parser (Node.js wrapper calls the Python script)
import { execSync } from "child_process";

function parseWiki(wikiRoot: string) {
  execSync(`python parse-knowledge-base.py ${wikiRoot}`, { stdio: "inherit" });
}

// 2️⃣ Invoke the Article Analyzer skill (via the Understand-Anything CLI)
import { runSkill } from "@understand-anything/skill";

async function extractEntities(manifestPath: string) {
  const result = await runSkill("understand-knowledge", [
    "--input", manifestPath,
    "--batch-size", "15",
  ]);
  console.log("New entities:", result.nodes.filter(n => n.type === "entity"));
}

// 3️⃣ Merge the new entities into the graph
import { mergeGraph } from "@understand-anything/core";

async function finalizeGraph(intermediateDir: string) {
  await mergeGraph(intermediateDir);
}

```

## Summary

- **Deterministic parsing** in [`parse-knowledge-base.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/parse-knowledge-base.py) extracts explicit wikilinks and article metadata from Karpathy-pattern wikis into a scan manifest.
- **LLM-driven extraction** via the Article Analyzer identifies implicit entities mentioned in prose but lacking dedicated pages, applying strict normalization rules (lower-case, hyphenated IDs).
- **Graph merging** in [`merge-knowledge-graph.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/merge-knowledge-graph.py) deduplicates entities using `normalize_entity_name` and integrates them into the final knowledge graph without duplication.
- The pipeline distinguishes between **article nodes** (pages that exist) and **entity nodes** (concepts mentioned but unlinked), creating a comprehensive knowledge representation.

## Frequently Asked Questions

### How does the parser distinguish between a Karpathy wiki and other markdown collections?

The `detect_format` function in [`parse-knowledge-base.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/parse-knowledge-base.py) (lines 38-66) checks for the presence of [`index.md`](https://github.com/Egonex-AI/Understand-Anything/blob/main/index.md) and validates that the directory contains a minimum number of markdown files, ensuring it follows the expected Karpathy wiki structure before processing.

### What normalization rules apply to entity IDs?

Entity IDs follow the pattern `entity:andrej-karpathy`, where the name is converted to lower-case and spaces are replaced with hyphens. This normalization occurs in [`merge-knowledge-graph.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/merge-knowledge-graph.py) (lines 83-84) and ensures consistent referencing across the knowledge graph.

### Can the LLM extract entities that already have wiki pages?

No. The Article Analyzer prompt explicitly instructs the LLM to return only **new** entity nodes that do not already appear as articles or wikilinks in the scan manifest. This prevents duplication between explicit `[[wikilink]]` references and implicit textual mentions.

### Where does the deduplication logic reside?

Deduplication and final graph integration occur in [`merge-knowledge-graph.py`](https://github.com/Egonex-AI/Understand-Anything/blob/main/merge-knowledge-graph.py) (lines 135-142), where the system collapses duplicate entity IDs and updates all edge references to point to the canonical node, ensuring each unique entity appears only once in the output graph.