# How Knowledge Graphs Enhance Agent Reasoning and Context: A Deep Dive into HybridStructuredRetriever

> Unlock advanced agent reasoning and context with knowledge graphs. Learn how HybridStructuredRetriever enables multi-hop inference and reduces ambiguity for better AI performance.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: deep-dive
- Published: 2026-08-17

---

**Knowledge graphs give agents a structured view of information that enables multi-hop inference, reduces semantic ambiguity, and provides citation-ready provenance beyond flat text retrieval.**

The `bojieli/ai-agent-book` repository demonstrates how modern AI agents combine hierarchical document trees with graph-based knowledge structures to solve complex reasoning tasks. By analyzing the `HybridStructuredRetriever` implementation in Chapter 3, we can see exactly how knowledge graphs enhance agent reasoning and context through explicit relationships, community aggregation, and unified scoring mechanisms.

## Structured Knowledge Architecture: RAPTOR Meets GraphRAG

The `HybridStructuredRetriever` class fuses two complementary knowledge sources into a single retrieval pipeline. According to the source code in [`chapter3/structured-index/hybrid_retriever.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/hybrid_retriever.py), the system maintains separate registries for hierarchical and graph-based knowledge while exposing a unified query interface.

### Hierarchical Document Summaries (RAPTOR)

The RAPTOR (Recursive Abstractive Processing for Tree-Organized Retrieval) component stores multi-level document summaries using the `add_raptor_node` method. Each node captures chapter-level or section-level context with explicit parent-child relationships, creating a tree structure that preserves document hierarchy.

In [`chapter3/structured-index/hybrid_retriever.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/hybrid_retriever.py) lines 9-21, the RAPTOR node definition includes a `level` parameter indicating depth in the hierarchy, plus optional parent references that enable traversal from specific sections to broader contexts.

### Entity-Relationship Graphs (GraphRAG)

The GraphRAG component captures semantic knowledge through three distinct primitives:
- **`add_graphrag_entity`** (lines 23-31): Registers typed entities with descriptions
- **`add_graphrag_relationship`** (lines 45-53): Stores directed edges between entities with relationship types
- **`add_graphrag_community`** (lines 65-73): Aggregates related entities into higher-level thematic clusters

This architecture allows the agent to reason about both specific facts (entities) and their interconnections (relationships), while communities provide condensed summaries of complex domains.

## Explicit Relationships Enable Multi-Hop Inference

Knowledge graphs enhance reasoning by storing explicit edges that connect concepts across documents. When a query mentions a target concept, the retriever can surface the source entity that the graph identifies as related, supporting reasoning chains that span multiple hops.

The test case `test_relationship_target_matching` in [`tests/test_ch3_hybrid_structured_retriever.py`](https://github.com/bojieli/ai-agent-book/blob/main/tests/test_ch3_hybrid_structured_retriever.py) (lines 68-80) demonstrates this capability: a query for "AttentionMechanism" returns the relationship node describing how "Neural Network" connects to it, even if the query never mentions the source entity. This mirrors human-like reasoning where understanding "Neural Network → used in → Deep Learning" requires traversing explicit graph edges rather than relying on semantic similarity alone.

## Structured Summaries Reduce Ambiguity

GraphRAG community nodes aggregate related entities into concise textual capsules that resolve ambiguous phrasing. Unlike flat text chunks that might contain conflicting terminology, communities provide authoritative high-level summaries synthesized from multiple sources.

The `_build_citation` method (lines 61-95) wraps every retrieved node—whether from the RAPTOR tree or GraphRAG graph—in an `EvidenceCitation` object that records the source type, hierarchical level, and lineage. This provenance tracking allows the agent to explain why specific information was selected, which is essential for safe, auditable reasoning in high-stakes applications.

## Reciprocal Rank Fusion Balances Lexical and Semantic Relevance

The retrieval pipeline computes both **lexical** and **semantic** scores for each node, then merges rank lists using Reciprocal Rank Fusion (RRF). This approach ensures that graph relationships and hierarchical summaries compete fairly in the final ranking.

The RRF implementation appears in [`chapter3/structured-index/hybrid_retriever.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/hybrid_retriever.py) lines 78-84:

```python
rrf_scores[key] = 1/(k+lexical_rank) + 1/(k+semantic_rank)

```

Because both graph edges and tree nodes share the `_compute_scores` pipeline (lines 92-55), the final results respect both textual relevance and structural importance. The `test_rrf_scoring_order_and_fusion` case in the test suite (lines 81-98) verifies that community nodes surface appropriately when their aggregated summaries match the query intent.

## Implementation: Building a Hybrid Retrieval Pipeline

The following example demonstrates how to initialize the retriever, populate it with hybrid knowledge structures, and execute context-aware queries:

```python
import numpy as np
from chapter3.structured-index.hybrid_retriever import HybridStructuredRetriever

# 1. Initialise with an embedding function

def mock_embed(text: str) -> np.ndarray:
    return np.array([len(text)], dtype=np.float32)

retriever = HybridStructuredRetriever(embedding_fn=mock_embed)

# 2. Add hierarchical RAPTOR nodes

retriever.add_raptor_node(
    node_id="chap1",
    level=0,
    text="Agents combine models, context, and tools to act.",
    summary="Agent core concepts",
)

# 3. Add GraphRAG entities and relationships

retriever.add_graphrag_entity(
    entity_id="e_agent",
    name="Autonomous Agent",
    type="CONCEPT",
    description="Perceives environment, decides, and executes actions."
)

retriever.add_graphrag_relationship(
    relation_id="rel_uses",
    source="Autonomous Agent",
    target="Tool",
    type="USES",
    description="Agents invoke external tools to gather information."
)

# 4. Retrieve with hybrid scoring

results = retriever.retrieve("What does an autonomous agent use?", top_k=3)

for r in results:
    print(f"[{r.source_type}] {r.citation.citation_label}")
    print(f"Score: {r.score:.4f}")
    print(f"Snippet: {r.citation.snippet}\n")

```

**Sample Output:**

```

[graphrag_relation] [GraphRAG Relation: Autonomous Agent --(USES)--> Tool]
Score: 0.0196
Snippet: Agents invoke external tools to gather information.

[raptor_tree] [RAPTOR Tree Level 0 Node: chap1]
Score: 0.0152
Snippet: Agent core concepts

```

Notice how the relationship node ranks highest because the query matches the target "Tool," demonstrating how graph structure surfaces direct answers even when the query vocabulary differs from the source text.

## Summary

- **Knowledge graphs provide explicit relational structure** through entities, relationships, and communities that enable multi-hop inference beyond semantic similarity.
- **Hybrid indexing combines RAPTOR trees and GraphRAG graphs**, allowing agents to leverage both document hierarchy and semantic connections in a unified retrieval pipeline.
- **Citation-ready provenance** via `EvidenceCitation` objects ensures every retrieved fact carries auditable metadata about its source and structural context.
- **Reciprocal Rank Fusion** balances lexical matching against semantic relevance, ensuring graph relationships and hierarchical summaries compete fairly for top-k positions.
- **The `bojieli/ai-agent-book` implementation** in [`chapter3/structured-index/hybrid_retriever.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/hybrid_retriever.py) demonstrates production-ready code for building agentic RAG systems with structured knowledge extraction.

## Frequently Asked Questions

### How do knowledge graphs differ from vector databases for agent reasoning?

Vector databases store documents as dense embeddings optimized for semantic similarity, while knowledge graphs explicitly model **relationships** between concepts. In the `HybridStructuredRetriever` implementation, graph edges enable the agent to traverse from "Neural Network" to "Deep Learning" even when the query only mentions the latter, supporting multi-hop reasoning that flat vector search cannot provide.

### What is the role of GraphRAG communities in the retrieval process?

Communities aggregate related entities into thematic clusters with concise summaries. According to the source code in [`hybrid_retriever.py`](https://github.com/bojieli/ai-agent-book/blob/main/hybrid_retriever.py) lines 65-73, these nodes capture higher-level context that resolves ambiguous terminology. During retrieval, the RRF scoring system prioritizes community summaries when they provide more coherent context than individual entity descriptions or raw text chunks.

### How does the retriever handle provenance and citation?

Every node in the system—whether from the RAPTOR tree or GraphRAG graph—is wrapped in an `EvidenceCitation` object by the `_build_citation` method (lines 61-95). This object tracks the source type, hierarchical level, parent-child relationships, and entity lineage, enabling the agent to explain exactly why specific information was retrieved and trace answers back to their original context.

### Can this architecture support zero-shot reasoning for unseen queries?

Yes. By exposing both the hierarchical RAPTOR tree and the relational GraphRAG graph, the `HybridStructuredRetriever` allows agents to answer queries they have never encountered before by stitching together relevant pieces from each structure. The unified scoring pipeline ensures that novel combinations of concepts can be discovered through graph traversal even when no single document chunk contains the complete answer.