How the Article Analyzer Extracts Implicit Relationships from Wikis in Understand-Anything

The article analyzer in Egonex-AI/Understand-Anything combines deterministic markdown parsing with LLM-based inference to extract both explicit wikilinks and implicit semantic relationships from wiki-style knowledge bases.

The Understand-Anything platform transforms static markdown wikis into navigable knowledge graphs. The article analyzer serves as the core intelligence layer that reads wiki files, discovers hidden connections between entities, and constructs a structured graph representation. This process enables the dashboard to surface concepts and relationships that exist only in the contextual meaning of the text, not just in explicit hyperlinks.

The Wiki Analysis Pipeline

The extraction process follows a six-stage pipeline implemented across the agent architecture.

Step 1: File Discovery and Loading

When you invoke the /understand-knowledge command, the project-scanner agent walks the target directory tree and streams every *.md file to the article analyzer. This initial ingestion captures the raw markdown content while preserving file hierarchy and relative paths.

Step 2: Deterministic Parsing

Before any LLM processing, the markdown-parser.ts utility performs deterministic extraction of explicit structure:

  • Wikilinks ([[Entity Name]]) → Direct edges between pages
  • Categories and front-matter → Grouping nodes
  • Header hierarchy → Section and subsection scopes

This deterministic phase ensures reproducible baseline relationships from the wiki syntax itself.

Step 3: LLM Prompt Generation

For each article, the analyzer constructs a structured prompt that includes the preserved source text, the set of already-discovered explicit links, and specific instructions for relationship extraction. The prompt directs the LLM to identify facts, claims, and implications connecting two or more entities, returning structured triples in JSON format.

Step 4: Implicit Relationship Mining

The LLM identifies implicit connections through contextual inference—relationships implied by phrases like "X improves Y" or "X is used to speed up Y" that don't appear as explicit wikilinks. The model also performs cross-article reasoning using the shared link context, enabling indirect relationship discovery across the wiki graph.

Step 5: Graph Enrichment

The extracted triples (subject, predicate, object) merge into the master knowledge graph. The system creates nodes for entities and labels edges with predicates like depends_on, produces, or is_a, attaching metadata such as confidence scores and an implicit flag for LLM-derived edges.

Step 6: Validation and Normalization

The graph-reviewer agent executes a final validation pass that normalizes predicate vocabularies (mapping synonyms like "uses" and "utilises" to canonical forms), detects orphan nodes, and flags contradictory edges before persistence.

Core Implementation Files

The architecture spans several key files in the understand-anything-plugin directory:

Practical Usage Examples

You can trigger the wiki analysis pipeline through the CLI or programmatically via the core API.

Command Line Execution

Run the article analyzer against a local wiki directory:


# Process all markdown files in the wiki

/understand-knowledge ~/my-wiki

This generates ./.understand-anything/knowledge-graph.json. Launch the dashboard to explore the results:

/understand-dashboard

Programmatic API Access

Import the core functions to analyze wikis within your Node.js applications:

import { analyzeWiki } from '@understand-anything/core/knowledge';

// Analyze a directory containing markdown files
const graph = await analyzeWiki({ root: '/home/user/knowledge/wiki' });

// Filter for implicitly extracted relationships
graph.edges
  .filter(e => e.metadata?.implicit)
  .forEach(e => console.log(`${e.source} → ${e.predicate} → ${e.target}`));

Single Article Analysis

Process individual markdown files to extract relationships without full graph construction:

import { extractArticleRelations } from '@understand-anything/core/agents/article-analyzer';

const markdown = await readFile('README.md', 'utf-8');
const relations = await extractArticleRelations(markdown);

console.log(JSON.stringify(relations, null, 2));

Summary

  • The article analyzer combines deterministic parsing and LLM inference to extract both explicit wikilinks and implicit semantic relationships from markdown wikis.
  • Deterministic extraction in markdown-parser.ts captures wikilinks and structure before LLM processing, ensuring reproducible baseline relationships.
  • LLM prompting generates structured triples representing implicit connections inferred from contextual text analysis and cross-article reasoning.
  • Graph validation in graph-reviewer.ts normalizes predicates and validates edges before persistence to knowledge-graph.json.
  • The /understand-knowledge CLI command and analyzeWiki() API provide both interactive and programmatic access to the extraction pipeline.

Frequently Asked Questions

How does the article analyzer distinguish between explicit and implicit relationships?

The analyzer uses deterministic parsing to identify explicit wikilinks ([[Entity]]) and categorization tags, storing these as baseline graph edges. Implicit relationships are identified through LLM analysis of contextual text passages, marked with an implicit metadata flag and typically assigned confidence scores based on the model's certainty about the inferred connection.

Can I customize the relationship types extracted from my wiki?

Yes. The extraction behavior is controlled through the prompt templates in article-analyzer.ts. You can modify the relationship extraction instructions to target specific predicate vocabularies or domain-specific connection types. The graph-reviewer.ts agent then maps these custom predicates to your preferred canonical forms during the normalization phase.

What file formats does the wiki analyzer support?

The current implementation focuses on markdown files (*.md) following the Karpathy-pattern wiki structure. The markdown-parser.ts utility specifically handles wikilink syntax, front-matter metadata, and header hierarchies standard to markdown-based knowledge bases.

Where does the article analyzer store the extracted knowledge graph?

By default, the pipeline outputs to ./.understand-anything/knowledge-graph.json in the root of the analyzed directory. This JSON file contains the complete node and edge structure, including metadata flags distinguishing implicit from explicit relationships, enabling the dashboard to render the semantic graph.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →