# Understanding the Knowledge Graph Relevance Model in LLM Wiki

> Discover the knowledge graph relevance model in LLM Wiki. Learn how it calculates semantic and structural relevance using edge weights and embedding similarity for precise document retrieval.

- Repository: [nash_su/llm_wiki](https://github.com/nashsu/llm_wiki)
- Tags: deep-dive
- Published: 2026-09-12

---

**The knowledge graph relevance model in LLM Wiki calculates semantic and structural relevance scores between connected markdown nodes using edge weights, node degree penalties, and embedding cosine similarity to surface the most pertinent documents for retrieval-augmented responses.**

The `nashsu/llm_wiki` repository implements a retrieval-oriented knowledge graph that transforms a collection of markdown files into a connected semantic network. At the heart of this system lies the **knowledge graph relevance model**, which quantifies the relationship strength between documents to prioritize the most contextually useful information during queries.

## How the Knowledge Graph Relevance Model Works

The relevance model operates as a multi-factor scoring system that evaluates both the structural connectivity and semantic content of your wiki. According to the source code in [`src/lib/graph-relevance.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/graph-relevance.ts), the implementation follows a three-phase pipeline.

### Graph Construction via `buildRetrievalGraph()`

The system first constructs a directed graph where each node represents a markdown file and each edge represents a citation or wikilink between documents. The `buildRetrievalGraph()` function parses frontmatter and link syntax to establish these connections, creating the foundation upon which relevance scores are calculated.

### Relevance Scoring with `calculateRelevance()`

For any pair of connected nodes, the `calculateRelevance(sourceNode, targetNode, graph)` function computes a composite score based on three distinct factors:

- **Edge weight** – The base significance of the link type, where explicit citations receive higher weights than incidental references.
- **Node degree penalty** –Hub nodes with many connections receive downward adjustments to prevent overly generic documents from dominating results.
- **Embedding similarity** – The cosine similarity between vector embeddings of the source and target node content, capturing semantic closeness even when explicit links are sparse.

The final relevance score is the product of the edge weight and embedding similarity, adjusted by the node degree factor.

### Node Ranking via `getRelatedNodes()`

When retrieving context for a specific query, `getRelatedNodes(sourceNode, graph, limit?)` aggregates all reachable nodes, invokes `calculateRelevance()` for each candidate, filters out scores ≤ 0, and returns the top-K results sorted in descending order. This ranking directly drives which documents the LLM receives as context.

## Core Implementation Files

The knowledge graph relevance model spans two primary locations in the codebase:

- **[`src/lib/graph-relevance.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/graph-relevance.ts)** – Contains the implementation of `buildRetrievalGraph()`, `calculateRelevance()`, and `getRelatedNodes()`, implementing the scoring logic and graph traversal algorithms.

- **[`src/lib/wiki-graph.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/wiki-graph.ts)** – Defines the `RetrievalNode` type and maintains the graph data structure, storing nodes as a Map keyed by file identifiers and managing the invocation of relevance calculations during query processing.

These modules work together to transform static markdown content into a dynamic, queryable knowledge base.

## Practical Code Examples

### Building the Retrieval Graph

Initialize the graph structure by scanning your wiki directory:

```typescript
import { buildRetrievalGraph } from '@/lib/graph-relevance'

const graph = await buildRetrievalGraph('/path/to/wiki')
console.log(`Indexed ${graph.nodes.size} documents with ${graph.edges.length} connections`)

```

### Retrieving Top-K Relevant Nodes

Fetch the most relevant documents for a specific file:

```typescript
import { getRelatedNodes } from '@/lib/graph-relevance'

const sourceId = 'README.md'
const topRelated = getRelatedNodes(sourceId, graph, 5)

topRelated.forEach(({ node, relevance }) => {
  console.log(`${node.id}: ${relevance.toFixed(3)}`)
})

```

### Direct Relevance Calculation

Calculate the specific relevance between two known nodes for custom ranking logic:

```typescript
import { calculateRelevance } from '@/lib/graph-relevance'

const score = calculateRelevance(
  graph.nodes.get('Architecture.md')!,
  graph.nodes.get('API_Reference.md')!,
  graph
)

console.log(`Semantic relevance: ${score}`)

```

## Summary

- The **knowledge graph relevance model** combines structural link analysis with semantic embedding similarity to rank document relevance.
- Implementation resides primarily in [`src/lib/graph-relevance.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/graph-relevance.ts), with type definitions in [`src/lib/wiki-graph.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/wiki-graph.ts).
- The scoring algorithm balances **edge weights**, **node degree penalties**, and **cosine similarity** of content embeddings.
- Use `getRelatedNodes()` for standard retrieval or `calculateRelevance()` for custom scoring workflows.

## Frequently Asked Questions

### What factors determine the relevance score in LLM Wiki's knowledge graph?

The relevance score derives from three components: the inherent weight of the link type (edge weight), a penalty applied to highly connected hub nodes (node degree), and the cosine similarity between vector embeddings of the source and target document contents. These factors combine to ensure both structurally important and semantically related documents achieve high rankings.

### How does LLM Wiki handle semantic similarity between documents?

The model generates vector embeddings for each markdown node's textual content using the system's text encoder, then computes cosine similarity between the source node and candidate target nodes. This embedding-based similarity captures latent semantic relationships that explicit linking alone might miss, particularly between conceptually related but unlinked documents.

### Where is the knowledge graph relevance model implemented in the codebase?

The core logic resides in [`src/lib/graph-relevance.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/graph-relevance.ts), which exports `buildRetrievalGraph()`, `calculateRelevance()`, and `getRelatedNodes()`. The underlying data structures and `RetrievalNode` type definitions live in [`src/lib/wiki-graph.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/wiki-graph.ts), while [`src/lib/reset-project-state.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/reset-project-state.ts) handles cache invalidation for the relevance module when projects reload.

### How does the relevance model avoid generic hub nodes?

The `calculateRelevance()` function applies a node degree penalty during scoring, reducing the relevance of documents that maintain an unusually high number of connections. This prevents ubiquitous "hub" pages (like main index files) from monopolizing results, ensuring that more specific, contextually relevant documents surface instead.