How Semantica's Precedent Search Works: Hybrid Vector and Graph Retrieval Explained
Semantica performs precedent search using a hybrid approach that combines vector similarity search with graph-based relationship traversal through PRECEDENT_FOR edges, merging results from both sources to return ranked Precedent objects.
Semantica is an open-source AGI framework that stores decisions as structured nodes in a ContextGraph and retrieves them using a sophisticated dual-path retrieval system. Understanding how Semantica's precedent search works requires examining both the vector similarity layer and the explicit relationship graph that connects related decisions.
The Hybrid Retrieval Pipeline
Semantica's precedent search operates through a multi-stage pipeline that unifies semantic similarity with explicit relational links recorded at storage time.
Step 1: Scenario Embedding and Vector Lookup
When a user submits a scenario, the system first encodes the text into embeddings using the configured EmbeddingProvider (OpenAI, Cohere, HuggingFace, etc.). These embeddings query the VectorStoreFacade, which abstracts concrete backends like FAISS, Qdrant, Weaviate, or Pinecone. The vector store returns the k most similar decision records, each carrying a similarity score representing semantic proximity.
Step 2: Graph Traversal via PRECEDENT_FOR Edges
Simultaneously, the system queries the KnowledgeGraph via AgentContext.find_precedents in semantica/context_graph.py. This method walks PRECEDENT_FOR edges to locate precedents that share explicit links with candidate decisions. The traversal respects optional filters such as category and limit parameters, and can be persisted in Neo4j, ArangoDB, or in-memory storage.
Step 3: Score Reconciliation and Result Normalization
The DecisionQuery.find_precedents_hybrid method in semantica/decision_query.py orchestrates the merging phase. When the same decision appears in both vector and graph results, scores are reconciled—typically favoring higher similarity values—and sorted by descending relevance. The final list is wrapped in lightweight Precedent models exposing decision_id, content, score, and provenance metadata.
Core Components and Architecture
The precedent search relies on three injected abstractions that enable completely offline or remote deployment configurations without code changes.
AgentContext and Explicit Graph Queries
The AgentContext class provides the primary interface for graph-only precedent retrieval through find_precedents(). This method performs targeted lookups against the ContextGraph without vector similarity, useful when you need strictly explicit precedent chains rather than semantic matches.
DecisionQuery for Hybrid Orchestration
DecisionQuery serves as the high-level orchestrator implementing find_precedents_hybrid(). This class coordinates between embedding generation, vector store queries, and graph traversals, implementing the business logic for merging heterogeneous result sets into unified Precedent objects.
ContextRetriever and Integration Facade
For higher-level APIs and framework integrations (such as CrewAI), ContextRetriever in semantica/retriever.py provides find_precedents_hybrid() as a thin wrapper. This facade allows third-party tools to invoke the full hybrid search without managing underlying component dependencies directly.
Key Source Files and Implementation Details
The precedent search implementation spans five critical modules in the Semantica codebase:
semantica/context_graph.py: Implements theAgentContextclass andfind_precedents()method that traversesPRECEDENT_FORedges in the knowledge graph layer.semantica/decision_query.py: ContainsDecisionQuery.find_precedents_hybrid(), the main orchestration logic for combining vector and graph results.semantica/retriever.py: HousesContextRetriever.find_precedents_hybrid(), a facade used by integrations and external APIs.semantica/vector_store.py: DefinesVectorStoreFacadeand concrete implementations for FAISS, Qdrant, Weaviate, Milvus, and Pinecone backends.semantica/embedding_provider.py: Specifies the pluggable interface for text-to-vector encoding across different LLM providers.
Practical Usage Examples
The following examples demonstrate how to invoke precedent search at different abstraction levels.
Graph-Only Precedent Lookup
For scenarios requiring explicit precedent chains without semantic similarity:
from semantica import AgentContext
ctx = AgentContext(...)
# Returns Precedent objects linked via PRECEDENT_FOR edges
precedents = ctx.find_precedents(
scenario="Approve mortgage refinance",
category="mortgage",
limit=5
)
for p in precedents:
print(p.decision_id, p.content[:80], f"{p.score:.2f}")
Hybrid Search (Vector + Graph)
For comprehensive retrieval combining semantic similarity with explicit relationships:
from semantica import DecisionQuery
dq = DecisionQuery(...)
# Retrieves similar past decisions plus their explicit precedents
hybrid_precedents = dq.find_precedents_hybrid(
scenario="Supply-chain procurement optimisation",
category="supply_chain",
limit=10
)
for p in hybrid_precedents:
print(p.decision_id, p.score, p.metadata["source"])
Integration via CrewAI Tool
For agent workflows using the CrewAI integration:
from semantica.integrations.crewai import DecisionTool
tool = DecisionTool(context=ctx)
result_json = tool.run(action="find_precedents", scenario="New loan application")
print(result_json) # JSON list of precedent objects
Summary
- Hybrid Architecture: Semantica combines vector similarity (semantic matching) with graph traversal (explicit
PRECEDENT_FORrelationships) to surface relevant decisions from the ContextGraph. - Entry Points: Use
AgentContext.find_precedents()insemantica/context_graph.pyfor graph-only queries, orDecisionQuery.find_precedents_hybrid()for unified search. - Modular Design: The system injects
VectorStoreFacade,KnowledgeGraph, andEmbeddingProviderdependencies, allowing swaps between FAISS/Qdrant/Weaviate or Neo4j/ArangoDB without API changes. - Result Model: All methods return standardized
Precedentobjects containing decision IDs, content snippets, reconciled scores, and provenance metadata.
Frequently Asked Questions
How does Semantica handle conflicts when the same decision appears in both vector and graph results?
When DecisionQuery.find_precedents_hybrid() detects duplicate decisions across vector and graph result sets, it reconciles the scores by selecting the higher similarity value. The merged list maintains unique entries sorted by descending relevance score, ensuring the strongest signal—whether from semantic similarity or explicit graph relationships—determines the final ranking.
Can I use Semantica's precedent search without a vector database?
Yes. The AgentContext.find_precedents() method in semantica/context_graph.py performs graph-only retrieval using PRECEDENT_FOR edge traversal without invoking the embedding provider or vector store. This mode requires only the KnowledgeGraph backend (Neo4j, ArangoDB, or in-memory) and skips semantic similarity entirely, relying solely on explicit precedent links recorded during decision storage.
What embedding providers does Semantica support for precedent search?
Semantica's EmbeddingProvider interface supports pluggable implementations including OpenAI, Cohere, HuggingFace, and custom providers. The provider is injected into DecisionQuery or ContextRetriever at construction time, allowing you to switch between cloud-hosted and local embedding models without modifying the precedent search logic in decision_query.py or retriever.py.
How does the graph traversal respect category filters?
The find_precedents() method accepts optional category and limit parameters that constrain the PRECEDENT_FOR edge traversal in semantica/context_graph.py. When specified, the graph query filters nodes by category labels before traversing precedent relationships, ensuring only decisions within the specified domain context are considered relevant for the search results.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →