How the Article Analyzer Extracts Implicit Relationships from Wikis in Understand-Anything
The article analyzer in Egonex-AI/Understand-Anything combines deterministic markdown parsing with LLM-based inference to extract both explicit wikilinks and implicit semantic relationships from wiki-style knowledge bases.
The Understand-Anything platform transforms static markdown wikis into navigable knowledge graphs. The article analyzer serves as the core intelligence layer that reads wiki files, discovers hidden connections between entities, and constructs a structured graph representation. This process enables the dashboard to surface concepts and relationships that exist only in the contextual meaning of the text, not just in explicit hyperlinks.
The Wiki Analysis Pipeline
The extraction process follows a six-stage pipeline implemented across the agent architecture.
Step 1: File Discovery and Loading
When you invoke the /understand-knowledge command, the project-scanner agent walks the target directory tree and streams every *.md file to the article analyzer. This initial ingestion captures the raw markdown content while preserving file hierarchy and relative paths.
Step 2: Deterministic Parsing
Before any LLM processing, the markdown-parser.ts utility performs deterministic extraction of explicit structure:
- Wikilinks (
[[Entity Name]]) → Direct edges between pages - Categories and front-matter → Grouping nodes
- Header hierarchy → Section and subsection scopes
This deterministic phase ensures reproducible baseline relationships from the wiki syntax itself.
Step 3: LLM Prompt Generation
For each article, the analyzer constructs a structured prompt that includes the preserved source text, the set of already-discovered explicit links, and specific instructions for relationship extraction. The prompt directs the LLM to identify facts, claims, and implications connecting two or more entities, returning structured triples in JSON format.
Step 4: Implicit Relationship Mining
The LLM identifies implicit connections through contextual inference—relationships implied by phrases like "X improves Y" or "X is used to speed up Y" that don't appear as explicit wikilinks. The model also performs cross-article reasoning using the shared link context, enabling indirect relationship discovery across the wiki graph.
Step 5: Graph Enrichment
The extracted triples (subject, predicate, object) merge into the master knowledge graph. The system creates nodes for entities and labels edges with predicates like depends_on, produces, or is_a, attaching metadata such as confidence scores and an implicit flag for LLM-derived edges.
Step 6: Validation and Normalization
The graph-reviewer agent executes a final validation pass that normalizes predicate vocabularies (mapping synonyms like "uses" and "utilises" to canonical forms), detects orphan nodes, and flags contradictory edges before persistence.
Core Implementation Files
The architecture spans several key files in the understand-anything-plugin directory:
understand-anything-plugin/agents/article-analyzer.ts– Orchestrates the LLM prompting, parses JSON responses into triples, and injects them into the knowledge graph.understand-anything-plugin/agents/graph-reviewer.ts– Validates the constructed graph, normalizes predicates, and ensures semantic consistency.understand-anything-plugin/skills/understand-knowledge.ts– CLI entry point that wires the scanner, analyzer, and reviewer into the/understand-knowledgecommand.understand-anything-plugin/src/utils/markdown-parser.ts– Extracts wikilinks, categories, and heading hierarchies before LLM processing.
Practical Usage Examples
You can trigger the wiki analysis pipeline through the CLI or programmatically via the core API.
Command Line Execution
Run the article analyzer against a local wiki directory:
# Process all markdown files in the wiki
/understand-knowledge ~/my-wiki
This generates ./.understand-anything/knowledge-graph.json. Launch the dashboard to explore the results:
/understand-dashboard
Programmatic API Access
Import the core functions to analyze wikis within your Node.js applications:
import { analyzeWiki } from '@understand-anything/core/knowledge';
// Analyze a directory containing markdown files
const graph = await analyzeWiki({ root: '/home/user/knowledge/wiki' });
// Filter for implicitly extracted relationships
graph.edges
.filter(e => e.metadata?.implicit)
.forEach(e => console.log(`${e.source} → ${e.predicate} → ${e.target}`));
Single Article Analysis
Process individual markdown files to extract relationships without full graph construction:
import { extractArticleRelations } from '@understand-anything/core/agents/article-analyzer';
const markdown = await readFile('README.md', 'utf-8');
const relations = await extractArticleRelations(markdown);
console.log(JSON.stringify(relations, null, 2));
Summary
- The article analyzer combines deterministic parsing and LLM inference to extract both explicit wikilinks and implicit semantic relationships from markdown wikis.
- Deterministic extraction in
markdown-parser.tscaptures wikilinks and structure before LLM processing, ensuring reproducible baseline relationships. - LLM prompting generates structured triples representing implicit connections inferred from contextual text analysis and cross-article reasoning.
- Graph validation in
graph-reviewer.tsnormalizes predicates and validates edges before persistence toknowledge-graph.json. - The
/understand-knowledgeCLI command andanalyzeWiki()API provide both interactive and programmatic access to the extraction pipeline.
Frequently Asked Questions
How does the article analyzer distinguish between explicit and implicit relationships?
The analyzer uses deterministic parsing to identify explicit wikilinks ([[Entity]]) and categorization tags, storing these as baseline graph edges. Implicit relationships are identified through LLM analysis of contextual text passages, marked with an implicit metadata flag and typically assigned confidence scores based on the model's certainty about the inferred connection.
Can I customize the relationship types extracted from my wiki?
Yes. The extraction behavior is controlled through the prompt templates in article-analyzer.ts. You can modify the relationship extraction instructions to target specific predicate vocabularies or domain-specific connection types. The graph-reviewer.ts agent then maps these custom predicates to your preferred canonical forms during the normalization phase.
What file formats does the wiki analyzer support?
The current implementation focuses on markdown files (*.md) following the Karpathy-pattern wiki structure. The markdown-parser.ts utility specifically handles wikilink syntax, front-matter metadata, and header hierarchies standard to markdown-based knowledge bases.
Where does the article analyzer store the extracted knowledge graph?
By default, the pipeline outputs to ./.understand-anything/knowledge-graph.json in the root of the analyzed directory. This JSON file contains the complete node and edge structure, including metadata flags distinguishing implicit from explicit relationships, enabling the dashboard to render the semantic graph.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →