How the Knowledge Base Analyzer Extracts Entities from Wiki Articles in Egonex-AI

The Understand-Anything knowledge base analyzer extracts entities from wiki articles through a three-stage pipeline: deterministic parsing of markdown files and explicit wikilinks, LLM-driven recognition of implicit entities mentioned in text but lacking dedicated pages, and a final deduplication pass that normalizes IDs and merges duplicates into the knowledge graph.

The Egonex-AI/Understand-Anything project transforms unstructured markdown wiki articles into structured knowledge graphs. Its entity extraction pipeline specifically targets both explicit references found in [[wikilink]] syntax and implicit mentions of people, tools, and organizations buried in article prose.

The Three-Stage Entity Extraction Pipeline

Stage 1: Deterministic Wiki Parsing with parse-knowledge-base.py

The pipeline begins in understand-anything-plugin/skills/understand-knowledge/parse-knowledge-base.py, which implements deterministic parsing of Karpathy-pattern wikis. The detect_format function (lines 38-66) identifies a valid wiki by verifying the presence of index.md and a minimum threshold of markdown files.

Once validated, the parser iterates through every markdown file to extract front-matter, headings, and explicit wikilinks using the extract_wikilinks helper (lines 89-97). Each article becomes a structured object containing id, name, summary, wikilinks, category, and truncated content. The parser emits these objects into a scan-manifest.json file (lines 278-301), which serves as the input for the next stage.

Stage 2: LLM-Driven Implicit Entity Recognition

The Article Analyzer agent processes the scan manifest to surface entities mentioned in article text that do not have their own wiki pages. According to the prompt defined in understand-anything-plugin/agents/article-analyzer.md (lines 27-35), the LLM creates entity nodes with a normalized ID, name, short summary, and tags.

Entity ID normalization follows strict rules: convert to lower-case and replace spaces with hyphens. For example, "Andrej Karpathy" becomes entity:andrej-karpathy. The LLM explicitly filters out any entities that already exist as articles or wikilinks in the manifest, ensuring only new, implicit entities are returned.

Stage 3: Deduplication and Graph Integration via merge-knowledge-graph.py

The final stage occurs in understand-anything-plugin/skills/understand-knowledge/merge-knowledge-graph.py. The normalize_entity_name function (lines 83-84) ensures consistent ID formatting across all candidate entities. A deduplication loop (lines 135-142) collapses duplicate IDs and updates edge references, integrating the new entity nodes into the existing knowledge graph without duplication.

Implementation Example

The following TypeScript demonstrates how to orchestrate the three-stage extraction flow:

// 1️⃣ Run the deterministic parser (Node.js wrapper calls the Python script)
import { execSync } from "child_process";

function parseWiki(wikiRoot: string) {
  execSync(`python parse-knowledge-base.py ${wikiRoot}`, { stdio: "inherit" });
}

// 2️⃣ Invoke the Article Analyzer skill (via the Understand-Anything CLI)
import { runSkill } from "@understand-anything/skill";

async function extractEntities(manifestPath: string) {
  const result = await runSkill("understand-knowledge", [
    "--input", manifestPath,
    "--batch-size", "15",
  ]);
  console.log("New entities:", result.nodes.filter(n => n.type === "entity"));
}

// 3️⃣ Merge the new entities into the graph
import { mergeGraph } from "@understand-anything/core";

async function finalizeGraph(intermediateDir: string) {
  await mergeGraph(intermediateDir);
}

Summary

  • Deterministic parsing in parse-knowledge-base.py extracts explicit wikilinks and article metadata from Karpathy-pattern wikis into a scan manifest.
  • LLM-driven extraction via the Article Analyzer identifies implicit entities mentioned in prose but lacking dedicated pages, applying strict normalization rules (lower-case, hyphenated IDs).
  • Graph merging in merge-knowledge-graph.py deduplicates entities using normalize_entity_name and integrates them into the final knowledge graph without duplication.
  • The pipeline distinguishes between article nodes (pages that exist) and entity nodes (concepts mentioned but unlinked), creating a comprehensive knowledge representation.

Frequently Asked Questions

How does the parser distinguish between a Karpathy wiki and other markdown collections?

The detect_format function in parse-knowledge-base.py (lines 38-66) checks for the presence of index.md and validates that the directory contains a minimum number of markdown files, ensuring it follows the expected Karpathy wiki structure before processing.

What normalization rules apply to entity IDs?

Entity IDs follow the pattern entity:andrej-karpathy, where the name is converted to lower-case and spaces are replaced with hyphens. This normalization occurs in merge-knowledge-graph.py (lines 83-84) and ensures consistent referencing across the knowledge graph.

Can the LLM extract entities that already have wiki pages?

No. The Article Analyzer prompt explicitly instructs the LLM to return only new entity nodes that do not already appear as articles or wikilinks in the scan manifest. This prevents duplication between explicit [[wikilink]] references and implicit textual mentions.

Where does the deduplication logic reside?

Deduplication and final graph integration occur in merge-knowledge-graph.py (lines 135-142), where the system collapses duplicate entity IDs and updates all edge references to point to the canonical node, ensuring each unique entity appears only once in the output graph.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →