Implementing Knowledge Graphs for AI Agent Knowledge Bases: A GraphRAG Deep Dive
GraphRAG combines LLM-powered entity extraction with graph traversal and dense retrieval to transform unstructured documents into queryable knowledge structures for AI agents.
AI agents need more than raw text—they need structured, traversable knowledge. In the ai-agent-book repository, the GraphRAG implementation provides a complete pipeline for building knowledge graphs from technical documentation, enabling hybrid retrieval that merges semantic similarity with graph-based exploration. This article walks through the architecture, data structures, and practical implementation based on the source code in bojieli/ai-agent-book.
Why Knowledge Graphs Matter for AI Agents
Traditional retrieval-augmented generation (RAG) relies solely on vector similarity, which misses relational context. Knowledge graphs capture entities and their relationships explicitly, allowing agents to:
- Traverse connections between concepts (e.g., "which instructions affect register X?")
- Reason over community structures and hierarchical abstractions
- Combine dense retrieval with graph navigation for more accurate answers
The GraphRAGIndexer class in chapter3/structured-index/graphrag_indexer.py implements this approach end-to-end.
GraphRAG Architecture Overview
The system divides work into six coordinated layers, each implemented as a method in GraphRAGIndexer:
| Layer | Responsibility | Key Method |
|---|---|---|
| Input Processing | Split documents into overlapping chunks | chunk_text (L78-L95) |
| Extraction | LLM-powered entity and relationship identification | extract_entities_relationships (L16-L45) |
| Graph Construction | Build NetworkX graph with embedded nodes | build_knowledge_graph (L22-L30) |
| Community Detection | Cluster related entities using Leiden/Louvain | detect_communities (L34-L62) |
| Hierarchical Summarization | Create multi-level community abstracts | hierarchical_summarization (L12-L28) |
| Search API | Hybrid retrieval combining graph and vector similarity | search (L89-L115) |
Configuration is centralized in GraphRAGConfig ([config.py](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/config.py#L71-L84)), which loads API keys, model identifiers, and hyperparameters from environment variables.
Core Data Structures for Knowledge Graphs
Three dataclasses define the graph's building blocks, all defined in graphrag_indexer.py:
Entity Nodes
@dataclass
class Entity:
id: str
name: str
type: str
description: str
embedding: Optional[np.ndarray]
attributes: Dict[str, Any]
Source: L26-L34
Relationship Edges
@dataclass
class Relationship:
id: str
source: str # Entity ID
target: str # Entity ID
type: str
description: str
weight: float = 1.0
Source: L38-L45
Community Clusters
@dataclass
class Community:
id: str
entity_ids: List[str]
summary: str
embedding: Optional[np.ndarray]
level: int
Source: L48-L56
These structures enable typed nodes, weighted edges, and hierarchical abstractions—essential for complex domain modeling.
Step-by-Step Indexing Workflow
Understanding the pipeline helps customize it for specific domains. Here is how GraphRAGIndexer processes raw text into a queryable knowledge graph:
-
Initialization (L58-L66)
- Loads
GraphRAGConfig - Creates OpenAI client for LLM calls
- Initializes Sentence-Transformer embedding model
- Ensures output directories exist
- Loads
-
Chunking (L78-L95)
- Splits text on sentence boundaries
- Respects
chunk_sizeandchunk_overlapparameters - Preserves context across chunk boundaries
-
LLM Extraction (L16-L45)
- Sends structured prompt to LLM for each chunk
- Parses JSON response into
EntityandRelationshiplists - Handles extraction failures gracefully
-
Embedding Generation
- Encodes each entity description via
self.embedding_model.encode - Stores vectors alongside metadata for similarity search
- Encodes each entity description via
-
Graph Assembly
- Populates NetworkX graph:
self.graph.add_node()andself.graph.add_edge() - Attaches full
EntityandRelationshipobjects as node/edge attributes
- Populates NetworkX graph:
-
Community Detection (L34-L62)
- Runs Leiden algorithm (falls back to Louvain)
- Generates human-readable summaries via LLM for each community
-
Hierarchical Summarization (L12-L28)
- Merges communities by embedding cosine similarity
- Creates higher-level abstractions with incrementing
levelvalues
-
Search Enablement (L89-L115)
- Implements hybrid retrieval: graph traversal first, then dense fallback
Practical Code Examples for Knowledge Graph Implementation
Building a Knowledge Graph from Documentation
This complete example shows how to process technical documentation into a structured graph:
from chapter3.structured_index.graphrag_indexer import GraphRAGIndexer
from chapter3.structured_index.config import get_graphrag_config
# 1️⃣ Load configuration (reads API keys from .env)
cfg = get_graphrag_config()
# 2️⃣ Create the indexer
indexer = GraphRAGIndexer(cfg)
# 3️⃣ Load raw documentation (e.g., Intel x86 manual)
with open("tech_doc.txt", "r", encoding="utf-8") as f:
raw_text = f.read()
# 4️⃣ Build the graph
indexer.build_knowledge_graph(raw_text)
# 5️⃣ Detect communities and create hierarchical summaries
indexer.detect_communities()
indexer.hierarchical_summarization()
Key sequence: initialization → build_knowledge_graph() → detect_communities() → hierarchical_summarization().
Querying the Knowledge Graph
Once built, the graph supports natural language queries with hybrid retrieval:
# Search for "register renaming"
results = indexer.search("register renaming", top_k=5)
for r in results:
print(f"Node: {r['entity_name']}")
print(f"Score: {r['score']:.2f}")
print("-" * 40)
The search method returns dictionaries containing:
entity_name: matched node namescore: relevance score from combined graph+vector ranking- Additional metadata (snippet, community ID, etc.)
Persisting and Reloading Indexes
Avoid reprocessing by saving built graphs:
# Persist to disk (implementation at L523-L550)
indexer.save()
# Later—restore without re-processing source text
indexer = GraphRAGIndexer(cfg)
indexer.load()
Artifacts are stored under indexes/graphrag/ as specified by GraphRAGConfig.index_dir.
Key Source Files for Knowledge Graph Implementation
| File | Purpose | Direct Link |
|---|---|---|
graphrag_indexer.py |
Core implementation: dataclasses, chunking, graph building, community detection, search | View source |
config.py |
Configuration dataclass and environment handling | View source |
main.py |
CLI entry point for running experiments | View source |
test_indexing.py |
Unit tests covering the full pipeline | View source |
tools.py (agentic-rag) |
search wrapper for higher-level agent integration |
View source |
Summary
The GraphRAG implementation in bojieli/ai-agent-book demonstrates production-ready knowledge graph construction for AI agents:
- LLM-driven extraction converts unstructured text into typed entities and relationships
- NetworkX-based storage enables flexible graph traversal and analysis
- Community detection (Leiden/Louvain) identifies semantic clusters
- Hierarchical summarization creates multi-level knowledge abstractions
- Hybrid search combines graph navigation with dense vector retrieval
These components integrate into a cohesive system that outperforms pure vector RAG on relational reasoning tasks.
Frequently Asked Questions
What makes GraphRAG different from standard vector RAG?
GraphRAG explicitly models relationships between entities, enabling traversal-based reasoning that vector similarity alone cannot provide. While standard RAG retrieves chunks based on semantic similarity, GraphRAG can follow connection chains (e.g., "instructions that modify register X") and leverage community summaries for broader context. The hybrid search method in graphrag_indexer.py uses both approaches for optimal recall.
Which graph algorithms does the implementation use?
The system uses Leiden community detection as primary, with Louvain as fallback (see detect_communities, L34-L62). These algorithms partition the graph into clusters of densely connected entities. The implementation also uses cosine similarity on community embeddings to drive hierarchical merging in hierarchical_summarization.
How scalable is this knowledge graph approach?
The current implementation uses NetworkX for graph storage, which handles graphs with thousands of nodes efficiently. For larger scale, the data structures (Entity, Relationship, Community) can be adapted to Neo4j or Apache AGE without changing the core logic. The embedding-based retrieval in search() remains efficient via approximate nearest neighbor libraries.
Can I customize the entity extraction for my domain?
Yes—modify the prompt in extract_entities_relationships (L16-L45) to specify domain-relevant entity types and relationship categories. The JSON parsing logic expects consistent output format but is agnostic to the specific types extracted. Adjust GraphRAGConfig.llm_model to use specialized fine-tuned models if available.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →