Understanding the Knowledge Graph Relevance Model in LLM Wiki
The knowledge graph relevance model in LLM Wiki calculates semantic and structural relevance scores between connected markdown nodes using edge weights, node degree penalties, and embedding cosine similarity to surface the most pertinent documents for retrieval-augmented responses.
The nashsu/llm_wiki repository implements a retrieval-oriented knowledge graph that transforms a collection of markdown files into a connected semantic network. At the heart of this system lies the knowledge graph relevance model, which quantifies the relationship strength between documents to prioritize the most contextually useful information during queries.
How the Knowledge Graph Relevance Model Works
The relevance model operates as a multi-factor scoring system that evaluates both the structural connectivity and semantic content of your wiki. According to the source code in src/lib/graph-relevance.ts, the implementation follows a three-phase pipeline.
Graph Construction via buildRetrievalGraph()
The system first constructs a directed graph where each node represents a markdown file and each edge represents a citation or wikilink between documents. The buildRetrievalGraph() function parses frontmatter and link syntax to establish these connections, creating the foundation upon which relevance scores are calculated.
Relevance Scoring with calculateRelevance()
For any pair of connected nodes, the calculateRelevance(sourceNode, targetNode, graph) function computes a composite score based on three distinct factors:
- Edge weight – The base significance of the link type, where explicit citations receive higher weights than incidental references.
- Node degree penalty –Hub nodes with many connections receive downward adjustments to prevent overly generic documents from dominating results.
- Embedding similarity – The cosine similarity between vector embeddings of the source and target node content, capturing semantic closeness even when explicit links are sparse.
The final relevance score is the product of the edge weight and embedding similarity, adjusted by the node degree factor.
Node Ranking via getRelatedNodes()
When retrieving context for a specific query, getRelatedNodes(sourceNode, graph, limit?) aggregates all reachable nodes, invokes calculateRelevance() for each candidate, filters out scores ≤ 0, and returns the top-K results sorted in descending order. This ranking directly drives which documents the LLM receives as context.
Core Implementation Files
The knowledge graph relevance model spans two primary locations in the codebase:
-
src/lib/graph-relevance.ts– Contains the implementation ofbuildRetrievalGraph(),calculateRelevance(), andgetRelatedNodes(), implementing the scoring logic and graph traversal algorithms. -
src/lib/wiki-graph.ts– Defines theRetrievalNodetype and maintains the graph data structure, storing nodes as a Map keyed by file identifiers and managing the invocation of relevance calculations during query processing.
These modules work together to transform static markdown content into a dynamic, queryable knowledge base.
Practical Code Examples
Building the Retrieval Graph
Initialize the graph structure by scanning your wiki directory:
import { buildRetrievalGraph } from '@/lib/graph-relevance'
const graph = await buildRetrievalGraph('/path/to/wiki')
console.log(`Indexed ${graph.nodes.size} documents with ${graph.edges.length} connections`)
Retrieving Top-K Relevant Nodes
Fetch the most relevant documents for a specific file:
import { getRelatedNodes } from '@/lib/graph-relevance'
const sourceId = 'README.md'
const topRelated = getRelatedNodes(sourceId, graph, 5)
topRelated.forEach(({ node, relevance }) => {
console.log(`${node.id}: ${relevance.toFixed(3)}`)
})
Direct Relevance Calculation
Calculate the specific relevance between two known nodes for custom ranking logic:
import { calculateRelevance } from '@/lib/graph-relevance'
const score = calculateRelevance(
graph.nodes.get('Architecture.md')!,
graph.nodes.get('API_Reference.md')!,
graph
)
console.log(`Semantic relevance: ${score}`)
Summary
- The knowledge graph relevance model combines structural link analysis with semantic embedding similarity to rank document relevance.
- Implementation resides primarily in
src/lib/graph-relevance.ts, with type definitions insrc/lib/wiki-graph.ts. - The scoring algorithm balances edge weights, node degree penalties, and cosine similarity of content embeddings.
- Use
getRelatedNodes()for standard retrieval orcalculateRelevance()for custom scoring workflows.
Frequently Asked Questions
What factors determine the relevance score in LLM Wiki's knowledge graph?
The relevance score derives from three components: the inherent weight of the link type (edge weight), a penalty applied to highly connected hub nodes (node degree), and the cosine similarity between vector embeddings of the source and target document contents. These factors combine to ensure both structurally important and semantically related documents achieve high rankings.
How does LLM Wiki handle semantic similarity between documents?
The model generates vector embeddings for each markdown node's textual content using the system's text encoder, then computes cosine similarity between the source node and candidate target nodes. This embedding-based similarity captures latent semantic relationships that explicit linking alone might miss, particularly between conceptually related but unlinked documents.
Where is the knowledge graph relevance model implemented in the codebase?
The core logic resides in src/lib/graph-relevance.ts, which exports buildRetrievalGraph(), calculateRelevance(), and getRelatedNodes(). The underlying data structures and RetrievalNode type definitions live in src/lib/wiki-graph.ts, while src/lib/reset-project-state.ts handles cache invalidation for the relevance module when projects reload.
How does the relevance model avoid generic hub nodes?
The calculateRelevance() function applies a node degree penalty during scoring, reducing the relevance of documents that maintain an unusually high number of connections. This prevents ubiquitous "hub" pages (like main index files) from monopolizing results, ensuring that more specific, contextually relevant documents surface instead.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →