Retrieving Context with Vector and Graph Traversal in AgentContext: A Complete Guide

AgentContext orchestrates hybrid retrieval by combining vector similarity search with knowledge graph traversal, automatically selecting the optimal strategy based on available stores and weighting results via the hybrid_alpha parameter.

The AgentContext class in the semantica-agi/semantica repository provides a high-level façade for autonomous agents that need both semantic similarity and structured reasoning capabilities. When building retrieval-augmented generation (RAG) systems, developers often face a choice between vector databases for similarity search and knowledge graphs for relationship traversal. AgentContext eliminates this trade-off by unifying both approaches into a single retrieval pipeline.

How AgentContext Orchestrates Hybrid Retrieval

The retrieval flow follows a five-stage pipeline that dynamically adapts based on the available storage backends and query parameters.

Auto-Detecting the Retrieval Strategy

When you invoke the retrieve() method, AgentContext first inspects the available storage configuration. In semantica/context/agent_context.py, the initialization logic checks for the presence of a knowledge graph (lines 94-101). If a graph is supplied, the system instantiates the hybrid ContextRetriever; otherwise, it falls back to pure vector retrieval from AgentMemory (lines 108-113). This automatic fallback ensures your agent degrades gracefully when graph infrastructure is unavailable.

The Three-Source Architecture

The ContextRetriever.retrieve method, implemented in semantica/context/context_retriever.py, aggregates results from three distinct sources:

  • Vector store: Performs similarity search against embedded query vectors using cosine similarity or Euclidean distance (lines 78-86, 90-102).
  • Knowledge graph: Executes GraphRAG-style traversal that first extracts intent (identifying entity and relationship verbs) from the query (lines 69-77, 1089-1099). It then matches entities and relationships using embeddings and expands the graph up to a configurable hop limit (lines 1225-1246, 1255-1272).
  • Memory store: Directly queries AgentMemory for recent conversational items to maintain short-term context continuity (lines 1045-1052).

Normalization and Weighting

Each source returns relevance scores that undergo normalization to a 0-1 scale before merging. The system applies the hybrid_alpha weighting parameter to balance contributions: a value of 0 yields pure vector retrieval, while 1 favors graph-only results (lines 63-70). Graph-derived results receive an additional boost proportional to the number of related entities and relationships discovered during traversal (lines 71-75).

A deduplication pass then merges results referring to the same graph node or containing identical content, applying a final boost for items found by multiple sources (lines 78-85, 86-94).

Proximity Enrichment with Anchor Nodes

When you supply an anchor_node parameter, AgentContext._apply_proximity_metadata calculates hop distance from the anchor, applies confidence decay based on graph distance, and generates a combined proximity score for each result (lines 1249-1295). This allows downstream components to re-rank results based on topological closeness to a specific entity, even if the semantic similarity score is lower.

Practical Implementation

The following example demonstrates configuring AgentContext with both FAISS vector storage and a knowledge graph, then executing a hybrid retrieval query:

from semantica.context import AgentContext
from semantica.vector_store.faiss_store import FaissStore
from semantica.kg import KnowledgeGraph

# Initialize storage backends

vs = FaissStore(dim=768)
kg = KnowledgeGraph()

# Create context with hybrid retrieval enabled

ctx = AgentContext(
    vector_store=vs,
    knowledge_graph=kg,
    hybrid_alpha=0.6,        # Favor graph-based retrieval (60% weight)

    graph_expansion=True
)

# Store memory with automatic entity/relationship extraction

mem_id = ctx.store(
    "User asked about loan approval criteria.",
    conversation_id="conv-42",
    extract_entities=True,
    extract_relationships=True,
)

# Retrieve using hybrid GraphRAG

results = ctx.retrieve(
    "What factors influence loan approval?",
    max_results=5,
    anchor_node="LoanApplication",
    proximity_weight=0.3,
)

for r in results:
    print(f"Score {r['combined_score']:.2f}: {r['content']}")
    print(f"  Related entities: {r.get('related_entities')}")
    print(f"  Graph distance: {r.get('hop_distance')}")

The system returns vector-derived memories, graph-derived entity descriptions, and any overlapping items, unified under a single ranking score.

Core Components and File Structure

Understanding the source layout helps debug retrieval behavior and extend functionality:

Summary

  • Hybrid retrieval in AgentContext automatically combines vector similarity, graph traversal, and memory lookup into a unified search result.
  • The hybrid_alpha parameter (0.0 to 1.0) controls the weighting between vector-only and graph-only retrieval modes.
  • GraphRAG-style traversal extracts entities and relationships from queries, then expands the graph up to configurable hop limits while applying relationship-based scoring boosts.
  • Anchor node proximity enables distance-based re-ranking from specific graph entities, supporting use cases where topological closeness matters more than pure semantic similarity.
  • Constructor flags like graph_expansion, decision_tracking, and kg_algorithms allow fine-grained control over retrieval behavior without modifying core logic.

Frequently Asked Questions

What does the hybrid_alpha parameter control in AgentContext?

The hybrid_alpha parameter determines the weighting between vector-based and graph-based retrieval scores during the merging phase. Set to 0, the system uses only vector similarity; set to 1, it uses only graph traversal results. Values between 0 and 1 blend both sources, with the default typically favoring a balanced approach. This normalization occurs in ContextRetriever before deduplication and proximity scoring are applied.

How does the anchor_node parameter improve retrieval accuracy?

When you specify an anchor_node, AgentContext._apply_proximity_metadata calculates the graph distance (hop count) between each result and the specified anchor entity. Results closer to the anchor receive higher proximity scores, which are combined with semantic relevance scores using the proximity_weight parameter. This is particularly effective for domain-specific queries where entities related to a known concept (like "LoanApplication") should be prioritized regardless of exact text matching.

How does AgentContext differ from standard vector-only RAG systems?

Unlike pure vector retrieval systems that rely solely on embedding similarity, AgentContext implements a GraphRAG architecture that understands entity relationships and traversal paths. When a knowledge graph is available, the system extracts intent from queries to identify relevant entities and relationships, then expands the search across the graph structure. According to the semantica source code, this hybrid approach captures relational context that vector similarity alone might miss, such as indirect connections between concepts.

How does the system prevent duplicate results from different sources?

During the scoring phase, ContextRetriever identifies results that refer to the same graph node or contain identical content across the vector store, graph, and memory sources. It merges these duplicates while preserving the highest individual score and applying a multi-source boost to items found by more than one retrieval method. This deduplication logic ensures the final result set in AgentContext.retrieve (lines 1025-1035) contains unique, high-quality context items ranked by their combined relevance.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →