Implementing Knowledge Graphs for AI Agent Knowledge Bases: A GraphRAG Deep Dive

GraphRAG combines LLM-powered entity extraction with graph traversal and dense retrieval to transform unstructured documents into queryable knowledge structures for AI agents.

AI agents need more than raw text—they need structured, traversable knowledge. In the ai-agent-book repository, the GraphRAG implementation provides a complete pipeline for building knowledge graphs from technical documentation, enabling hybrid retrieval that merges semantic similarity with graph-based exploration. This article walks through the architecture, data structures, and practical implementation based on the source code in bojieli/ai-agent-book.

Why Knowledge Graphs Matter for AI Agents

Traditional retrieval-augmented generation (RAG) relies solely on vector similarity, which misses relational context. Knowledge graphs capture entities and their relationships explicitly, allowing agents to:

  • Traverse connections between concepts (e.g., "which instructions affect register X?")
  • Reason over community structures and hierarchical abstractions
  • Combine dense retrieval with graph navigation for more accurate answers

The GraphRAGIndexer class in chapter3/structured-index/graphrag_indexer.py implements this approach end-to-end.

GraphRAG Architecture Overview

The system divides work into six coordinated layers, each implemented as a method in GraphRAGIndexer:

Layer Responsibility Key Method
Input Processing Split documents into overlapping chunks chunk_text (L78-L95)
Extraction LLM-powered entity and relationship identification extract_entities_relationships (L16-L45)
Graph Construction Build NetworkX graph with embedded nodes build_knowledge_graph (L22-L30)
Community Detection Cluster related entities using Leiden/Louvain detect_communities (L34-L62)
Hierarchical Summarization Create multi-level community abstracts hierarchical_summarization (L12-L28)
Search API Hybrid retrieval combining graph and vector similarity search (L89-L115)

Configuration is centralized in GraphRAGConfig ([config.py](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/config.py#L71-L84)), which loads API keys, model identifiers, and hyperparameters from environment variables.

Core Data Structures for Knowledge Graphs

Three dataclasses define the graph's building blocks, all defined in graphrag_indexer.py:

Entity Nodes

@dataclass
class Entity:
    id: str
    name: str
    type: str
    description: str
    embedding: Optional[np.ndarray]
    attributes: Dict[str, Any]

Source: L26-L34

Relationship Edges

@dataclass
class Relationship:
    id: str
    source: str   # Entity ID

    target: str   # Entity ID

    type: str
    description: str
    weight: float = 1.0

Source: L38-L45

Community Clusters

@dataclass
class Community:
    id: str
    entity_ids: List[str]
    summary: str
    embedding: Optional[np.ndarray]
    level: int

Source: L48-L56

These structures enable typed nodes, weighted edges, and hierarchical abstractions—essential for complex domain modeling.

Step-by-Step Indexing Workflow

Understanding the pipeline helps customize it for specific domains. Here is how GraphRAGIndexer processes raw text into a queryable knowledge graph:

  1. Initialization (L58-L66)

    • Loads GraphRAGConfig
    • Creates OpenAI client for LLM calls
    • Initializes Sentence-Transformer embedding model
    • Ensures output directories exist
  2. Chunking (L78-L95)

    • Splits text on sentence boundaries
    • Respects chunk_size and chunk_overlap parameters
    • Preserves context across chunk boundaries
  3. LLM Extraction (L16-L45)

    • Sends structured prompt to LLM for each chunk
    • Parses JSON response into Entity and Relationship lists
    • Handles extraction failures gracefully
  4. Embedding Generation

    • Encodes each entity description via self.embedding_model.encode
    • Stores vectors alongside metadata for similarity search
  5. Graph Assembly

    • Populates NetworkX graph: self.graph.add_node() and self.graph.add_edge()
    • Attaches full Entity and Relationship objects as node/edge attributes
  6. Community Detection (L34-L62)

    • Runs Leiden algorithm (falls back to Louvain)
    • Generates human-readable summaries via LLM for each community
  7. Hierarchical Summarization (L12-L28)

    • Merges communities by embedding cosine similarity
    • Creates higher-level abstractions with incrementing level values
  8. Search Enablement (L89-L115)

    • Implements hybrid retrieval: graph traversal first, then dense fallback

Practical Code Examples for Knowledge Graph Implementation

Building a Knowledge Graph from Documentation

This complete example shows how to process technical documentation into a structured graph:

from chapter3.structured_index.graphrag_indexer import GraphRAGIndexer
from chapter3.structured_index.config import get_graphrag_config

# 1️⃣ Load configuration (reads API keys from .env)

cfg = get_graphrag_config()

# 2️⃣ Create the indexer

indexer = GraphRAGIndexer(cfg)

# 3️⃣ Load raw documentation (e.g., Intel x86 manual)

with open("tech_doc.txt", "r", encoding="utf-8") as f:
    raw_text = f.read()

# 4️⃣ Build the graph

indexer.build_knowledge_graph(raw_text)

# 5️⃣ Detect communities and create hierarchical summaries

indexer.detect_communities()
indexer.hierarchical_summarization()

Key sequence: initialization → build_knowledge_graph() → detect_communities() → hierarchical_summarization().

Querying the Knowledge Graph

Once built, the graph supports natural language queries with hybrid retrieval:


# Search for "register renaming"

results = indexer.search("register renaming", top_k=5)

for r in results:
    print(f"Node: {r['entity_name']}")
    print(f"Score: {r['score']:.2f}")
    print("-" * 40)

The search method returns dictionaries containing:

  • entity_name: matched node name
  • score: relevance score from combined graph+vector ranking
  • Additional metadata (snippet, community ID, etc.)

Persisting and Reloading Indexes

Avoid reprocessing by saving built graphs:


# Persist to disk (implementation at L523-L550)

indexer.save()

# Later—restore without re-processing source text

indexer = GraphRAGIndexer(cfg)
indexer.load()

Artifacts are stored under indexes/graphrag/ as specified by GraphRAGConfig.index_dir.

Key Source Files for Knowledge Graph Implementation

File Purpose Direct Link
graphrag_indexer.py Core implementation: dataclasses, chunking, graph building, community detection, search View source
config.py Configuration dataclass and environment handling View source
main.py CLI entry point for running experiments View source
test_indexing.py Unit tests covering the full pipeline View source
tools.py (agentic-rag) search wrapper for higher-level agent integration View source

Summary

The GraphRAG implementation in bojieli/ai-agent-book demonstrates production-ready knowledge graph construction for AI agents:

  • LLM-driven extraction converts unstructured text into typed entities and relationships
  • NetworkX-based storage enables flexible graph traversal and analysis
  • Community detection (Leiden/Louvain) identifies semantic clusters
  • Hierarchical summarization creates multi-level knowledge abstractions
  • Hybrid search combines graph navigation with dense vector retrieval

These components integrate into a cohesive system that outperforms pure vector RAG on relational reasoning tasks.

Frequently Asked Questions

What makes GraphRAG different from standard vector RAG?

GraphRAG explicitly models relationships between entities, enabling traversal-based reasoning that vector similarity alone cannot provide. While standard RAG retrieves chunks based on semantic similarity, GraphRAG can follow connection chains (e.g., "instructions that modify register X") and leverage community summaries for broader context. The hybrid search method in graphrag_indexer.py uses both approaches for optimal recall.

Which graph algorithms does the implementation use?

The system uses Leiden community detection as primary, with Louvain as fallback (see detect_communities, L34-L62). These algorithms partition the graph into clusters of densely connected entities. The implementation also uses cosine similarity on community embeddings to drive hierarchical merging in hierarchical_summarization.

How scalable is this knowledge graph approach?

The current implementation uses NetworkX for graph storage, which handles graphs with thousands of nodes efficiently. For larger scale, the data structures (Entity, Relationship, Community) can be adapted to Neo4j or Apache AGE without changing the core logic. The embedding-based retrieval in search() remains efficient via approximate nearest neighbor libraries.

Can I customize the entity extraction for my domain?

Yes—modify the prompt in extract_entities_relationships (L16-L45) to specify domain-relevant entity types and relationship categories. The JSON parsing logic expects consistent output format but is agnostic to the specific types extracted. Adjust GraphRAGConfig.llm_model to use specialized fine-tuned models if available.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →