# Implementing Knowledge Graphs for AI Agent Knowledge Bases: A GraphRAG Deep Dive

> Learn how to implement knowledge graphs for AI agent knowledge bases using GraphRAG. Transform documents into queryable structures with LLM entity extraction and graph traversal.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: deep-dive
- Published: 2026-08-18

---

**GraphRAG combines LLM-powered entity extraction with graph traversal and dense retrieval to transform unstructured documents into queryable knowledge structures for AI agents.**

AI agents need more than raw text—they need structured, traversable knowledge. In the **ai-agent-book** repository, the **GraphRAG** implementation provides a complete pipeline for building knowledge graphs from technical documentation, enabling hybrid retrieval that merges semantic similarity with graph-based exploration. This article walks through the architecture, data structures, and practical implementation based on the source code in `bojieli/ai-agent-book`.

## Why Knowledge Graphs Matter for AI Agents

Traditional **retrieval-augmented generation (RAG)** relies solely on vector similarity, which misses relational context. Knowledge graphs capture entities and their relationships explicitly, allowing agents to:

- Traverse connections between concepts (e.g., "which instructions affect register X?")
- Reason over community structures and hierarchical abstractions
- Combine dense retrieval with graph navigation for more accurate answers

The `GraphRAGIndexer` class in [`chapter3/structured-index/graphrag_indexer.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/graphrag_indexer.py) implements this approach end-to-end.

## GraphRAG Architecture Overview

The system divides work into six coordinated layers, each implemented as a method in `GraphRAGIndexer`:

| Layer | Responsibility | Key Method |
|-------|--------------|------------|
| **Input Processing** | Split documents into overlapping chunks | `chunk_text` ([L78-L95](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/graphrag_indexer.py#L78-L95)) |
| **Extraction** | LLM-powered entity and relationship identification | `extract_entities_relationships` ([L16-L45](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/graphrag_indexer.py#L16-L45)) |
| **Graph Construction** | Build NetworkX graph with embedded nodes | `build_knowledge_graph` ([L22-L30](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/graphrag_indexer.py#L22-L30)) |
| **Community Detection** | Cluster related entities using Leiden/Louvain | `detect_communities` ([L34-L62](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/graphrag_indexer.py#L34-L62)) |
| **Hierarchical Summarization** | Create multi-level community abstracts | `hierarchical_summarization` ([L12-L28](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/graphrag_indexer.py#L12-L28)) |
| **Search API** | Hybrid retrieval combining graph and vector similarity | `search` ([L89-L115](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/graphrag_indexer.py#L89-L115)) |

Configuration is centralized in `GraphRAGConfig` ([[`config.py`](https://github.com/bojieli/ai-agent-book/blob/main/config.py#L71-L84)](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/config.py#L71-L84)), which loads API keys, model identifiers, and hyperparameters from environment variables.

## Core Data Structures for Knowledge Graphs

Three dataclasses define the graph's building blocks, all defined in [`graphrag_indexer.py`](https://github.com/bojieli/ai-agent-book/blob/main/graphrag_indexer.py):

### Entity Nodes

```python
@dataclass
class Entity:
    id: str
    name: str
    type: str
    description: str
    embedding: Optional[np.ndarray]
    attributes: Dict[str, Any]

```

*Source*: [L26-L34](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/graphrag_indexer.py#L26-L34)

### Relationship Edges

```python
@dataclass
class Relationship:
    id: str
    source: str   # Entity ID

    target: str   # Entity ID

    type: str
    description: str
    weight: float = 1.0

```

*Source*: [L38-L45](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/graphrag_indexer.py#L38-L45)

### Community Clusters

```python
@dataclass
class Community:
    id: str
    entity_ids: List[str]
    summary: str
    embedding: Optional[np.ndarray]
    level: int

```

*Source*: [L48-L56](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/graphrag_indexer.py#L48-L56)

These structures enable **typed nodes**, **weighted edges**, and **hierarchical abstractions**—essential for complex domain modeling.

## Step-by-Step Indexing Workflow

Understanding the pipeline helps customize it for specific domains. Here is how `GraphRAGIndexer` processes raw text into a queryable knowledge graph:

1. **Initialization** ([L58-L66](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/graphrag_indexer.py#L58-L66))
   - Loads `GraphRAGConfig`
   - Creates OpenAI client for LLM calls
   - Initializes Sentence-Transformer embedding model
   - Ensures output directories exist

2. **Chunking** ([L78-L95](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/graphrag_indexer.py#L78-L95))
   - Splits text on sentence boundaries
   - Respects `chunk_size` and `chunk_overlap` parameters
   - Preserves context across chunk boundaries

3. **LLM Extraction** ([L16-L45](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/graphrag_indexer.py#L16-L45))
   - Sends structured prompt to LLM for each chunk
   - Parses JSON response into `Entity` and `Relationship` lists
   - Handles extraction failures gracefully

4. **Embedding Generation**
   - Encodes each entity description via `self.embedding_model.encode`
   - Stores vectors alongside metadata for similarity search

5. **Graph Assembly**
   - Populates NetworkX graph: `self.graph.add_node()` and `self.graph.add_edge()`
   - Attaches full `Entity` and `Relationship` objects as node/edge attributes

6. **Community Detection** ([L34-L62](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/graphrag_indexer.py#L34-L62))
   - Runs Leiden algorithm (falls back to Louvain)
   - Generates human-readable summaries via LLM for each community

7. **Hierarchical Summarization** ([L12-L28](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/graphrag_indexer.py#L12-L28))
   - Merges communities by embedding cosine similarity
   - Creates higher-level abstractions with incrementing `level` values

8. **Search Enablement** ([L89-L115](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/graphrag_indexer.py#L89-L115))
   - Implements hybrid retrieval: graph traversal first, then dense fallback

## Practical Code Examples for Knowledge Graph Implementation

### Building a Knowledge Graph from Documentation

This complete example shows how to process technical documentation into a structured graph:

```python
from chapter3.structured_index.graphrag_indexer import GraphRAGIndexer
from chapter3.structured_index.config import get_graphrag_config

# 1️⃣ Load configuration (reads API keys from .env)

cfg = get_graphrag_config()

# 2️⃣ Create the indexer

indexer = GraphRAGIndexer(cfg)

# 3️⃣ Load raw documentation (e.g., Intel x86 manual)

with open("tech_doc.txt", "r", encoding="utf-8") as f:
    raw_text = f.read()

# 4️⃣ Build the graph

indexer.build_knowledge_graph(raw_text)

# 5️⃣ Detect communities and create hierarchical summaries

indexer.detect_communities()
indexer.hierarchical_summarization()

```

**Key sequence**: initialization → `build_knowledge_graph()` → `detect_communities()` → `hierarchical_summarization()`.

### Querying the Knowledge Graph

Once built, the graph supports natural language queries with hybrid retrieval:

```python

# Search for "register renaming"

results = indexer.search("register renaming", top_k=5)

for r in results:
    print(f"Node: {r['entity_name']}")
    print(f"Score: {r['score']:.2f}")
    print("-" * 40)

```

The `search` method returns dictionaries containing:
- `entity_name`: matched node name
- `score`: relevance score from combined graph+vector ranking
- Additional metadata (snippet, community ID, etc.)

### Persisting and Reloading Indexes

Avoid reprocessing by saving built graphs:

```python

# Persist to disk (implementation at L523-L550)

indexer.save()

# Later—restore without re-processing source text

indexer = GraphRAGIndexer(cfg)
indexer.load()

```

Artifacts are stored under `indexes/graphrag/` as specified by `GraphRAGConfig.index_dir`.

## Key Source Files for Knowledge Graph Implementation

| File | Purpose | Direct Link |
|------|---------|-------------|
| [`graphrag_indexer.py`](https://github.com/bojieli/ai-agent-book/blob/main/graphrag_indexer.py) | Core implementation: dataclasses, chunking, graph building, community detection, search | [View source](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/graphrag_indexer.py) |
| [`config.py`](https://github.com/bojieli/ai-agent-book/blob/main/config.py) | Configuration dataclass and environment handling | [View source](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/config.py) |
| [`main.py`](https://github.com/bojieli/ai-agent-book/blob/main/main.py) | CLI entry point for running experiments | [View source](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/main.py) |
| [`test_indexing.py`](https://github.com/bojieli/ai-agent-book/blob/main/test_indexing.py) | Unit tests covering the full pipeline | [View source](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/test_indexing.py) |
| [`tools.py`](https://github.com/bojieli/ai-agent-book/blob/main/tools.py) (agentic-rag) | `search` wrapper for higher-level agent integration | [View source](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/agentic-rag/tools.py) |

## Summary

The **GraphRAG** implementation in `bojieli/ai-agent-book` demonstrates production-ready **knowledge graph construction for AI agents**:

- **LLM-driven extraction** converts unstructured text into typed entities and relationships
- **NetworkX-based storage** enables flexible graph traversal and analysis
- **Community detection** (Leiden/Louvain) identifies semantic clusters
- **Hierarchical summarization** creates multi-level knowledge abstractions
- **Hybrid search** combines graph navigation with dense vector retrieval

These components integrate into a cohesive system that outperforms pure vector RAG on relational reasoning tasks.

## Frequently Asked Questions

### What makes GraphRAG different from standard vector RAG?

**GraphRAG explicitly models relationships between entities**, enabling traversal-based reasoning that vector similarity alone cannot provide. While standard RAG retrieves chunks based on semantic similarity, GraphRAG can follow connection chains (e.g., "instructions that modify register X") and leverage community summaries for broader context. The hybrid `search` method in [`graphrag_indexer.py`](https://github.com/bojieli/ai-agent-book/blob/main/graphrag_indexer.py) uses both approaches for optimal recall.

### Which graph algorithms does the implementation use?

The system uses **Leiden community detection** as primary, with **Louvain as fallback** (see [`detect_communities`, L34-L62](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/graphrag_indexer.py#L34-L62)). These algorithms partition the graph into clusters of densely connected entities. The implementation also uses **cosine similarity** on community embeddings to drive hierarchical merging in `hierarchical_summarization`.

### How scalable is this knowledge graph approach?

The current implementation uses **NetworkX** for graph storage, which handles graphs with thousands of nodes efficiently. For larger scale, the data structures (`Entity`, `Relationship`, `Community`) can be adapted to **Neo4j** or **Apache AGE** without changing the core logic. The embedding-based retrieval in `search()` remains efficient via approximate nearest neighbor libraries.

### Can I customize the entity extraction for my domain?

Yes—modify the prompt in `extract_entities_relationships` ([L16-L45](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/graphrag_indexer.py#L16-L45)) to specify domain-relevant entity types and relationship categories. The JSON parsing logic expects consistent output format but is agnostic to the specific types extracted. Adjust `GraphRAGConfig.llm_model` to use specialized fine-tuned models if available.