# How Semantica's Precedent Search Works: Hybrid Vector and Graph Retrieval Explained

> Discover how Semantica's hybrid precedent search merges vector similarity and graph traversal for powerful retrieval. Understand the PRECENDENT_FOR edge and ranked Precedent objects.

- Repository: [Semantica /semantica](https://github.com/semantica-agi/semantica)
- Tags: deep-dive
- Published: 2026-09-08

---

**Semantica performs precedent search using a hybrid approach that combines vector similarity search with graph-based relationship traversal through `PRECEDENT_FOR` edges, merging results from both sources to return ranked `Precedent` objects.**

Semantica is an open-source AGI framework that stores decisions as structured nodes in a **ContextGraph** and retrieves them using a sophisticated dual-path retrieval system. Understanding how Semantica's precedent search works requires examining both the vector similarity layer and the explicit relationship graph that connects related decisions.

## The Hybrid Retrieval Pipeline

Semantica's precedent search operates through a multi-stage pipeline that unifies semantic similarity with explicit relational links recorded at storage time.

### Step 1: Scenario Embedding and Vector Lookup

When a user submits a scenario, the system first encodes the text into embeddings using the configured **EmbeddingProvider** (OpenAI, Cohere, HuggingFace, etc.). These embeddings query the **VectorStoreFacade**, which abstracts concrete backends like FAISS, Qdrant, Weaviate, or Pinecone. The vector store returns the *k* most similar decision records, each carrying a similarity score representing semantic proximity.

### Step 2: Graph Traversal via PRECEDENT_FOR Edges

Simultaneously, the system queries the **KnowledgeGraph** via `AgentContext.find_precedents` in [`semantica/context_graph.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/context_graph.py). This method walks **`PRECEDENT_FOR`** edges to locate precedents that share explicit links with candidate decisions. The traversal respects optional filters such as `category` and `limit` parameters, and can be persisted in Neo4j, ArangoDB, or in-memory storage.

### Step 3: Score Reconciliation and Result Normalization

The `DecisionQuery.find_precedents_hybrid` method in [`semantica/decision_query.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/decision_query.py) orchestrates the merging phase. When the same decision appears in both vector and graph results, scores are reconciled—typically favoring higher similarity values—and sorted by descending relevance. The final list is wrapped in lightweight `Precedent` models exposing `decision_id`, `content`, `score`, and provenance metadata.

## Core Components and Architecture

The precedent search relies on three injected abstractions that enable completely offline or remote deployment configurations without code changes.

### AgentContext and Explicit Graph Queries

The `AgentContext` class provides the primary interface for graph-only precedent retrieval through `find_precedents()`. This method performs targeted lookups against the **ContextGraph** without vector similarity, useful when you need strictly explicit precedent chains rather than semantic matches.

### DecisionQuery for Hybrid Orchestration

`DecisionQuery` serves as the high-level orchestrator implementing `find_precedents_hybrid()`. This class coordinates between embedding generation, vector store queries, and graph traversals, implementing the business logic for merging heterogeneous result sets into unified `Precedent` objects.

### ContextRetriever and Integration Facade

For higher-level APIs and framework integrations (such as CrewAI), `ContextRetriever` in [`semantica/retriever.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/retriever.py) provides `find_precedents_hybrid()` as a thin wrapper. This facade allows third-party tools to invoke the full hybrid search without managing underlying component dependencies directly.

## Key Source Files and Implementation Details

The precedent search implementation spans five critical modules in the Semantica codebase:

- **[`semantica/context_graph.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/context_graph.py)**: Implements the `AgentContext` class and `find_precedents()` method that traverses `PRECEDENT_FOR` edges in the knowledge graph layer.
- **[`semantica/decision_query.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/decision_query.py)**: Contains `DecisionQuery.find_precedents_hybrid()`, the main orchestration logic for combining vector and graph results.
- **[`semantica/retriever.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/retriever.py)**: Houses `ContextRetriever.find_precedents_hybrid()`, a facade used by integrations and external APIs.
- **[`semantica/vector_store.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/vector_store.py)**: Defines `VectorStoreFacade` and concrete implementations for FAISS, Qdrant, Weaviate, Milvus, and Pinecone backends.
- **[`semantica/embedding_provider.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/embedding_provider.py)**: Specifies the pluggable interface for text-to-vector encoding across different LLM providers.

## Practical Usage Examples

The following examples demonstrate how to invoke precedent search at different abstraction levels.

### Graph-Only Precedent Lookup

For scenarios requiring explicit precedent chains without semantic similarity:

```python
from semantica import AgentContext

ctx = AgentContext(...)

# Returns Precedent objects linked via PRECEDENT_FOR edges

precedents = ctx.find_precedents(
    scenario="Approve mortgage refinance",
    category="mortgage",
    limit=5
)

for p in precedents:
    print(p.decision_id, p.content[:80], f"{p.score:.2f}")

```

### Hybrid Search (Vector + Graph)

For comprehensive retrieval combining semantic similarity with explicit relationships:

```python
from semantica import DecisionQuery

dq = DecisionQuery(...)

# Retrieves similar past decisions plus their explicit precedents

hybrid_precedents = dq.find_precedents_hybrid(
    scenario="Supply-chain procurement optimisation",
    category="supply_chain",
    limit=10
)

for p in hybrid_precedents:
    print(p.decision_id, p.score, p.metadata["source"])

```

### Integration via CrewAI Tool

For agent workflows using the CrewAI integration:

```python
from semantica.integrations.crewai import DecisionTool

tool = DecisionTool(context=ctx)
result_json = tool.run(action="find_precedents", scenario="New loan application")
print(result_json)  # JSON list of precedent objects

```

## Summary

- **Hybrid Architecture**: Semantica combines **vector similarity** (semantic matching) with **graph traversal** (explicit `PRECEDENT_FOR` relationships) to surface relevant decisions from the ContextGraph.
- **Entry Points**: Use `AgentContext.find_precedents()` in [`semantica/context_graph.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/context_graph.py) for graph-only queries, or `DecisionQuery.find_precedents_hybrid()` for unified search.
- **Modular Design**: The system injects `VectorStoreFacade`, `KnowledgeGraph`, and `EmbeddingProvider` dependencies, allowing swaps between FAISS/Qdrant/Weaviate or Neo4j/ArangoDB without API changes.
- **Result Model**: All methods return standardized `Precedent` objects containing decision IDs, content snippets, reconciled scores, and provenance metadata.

## Frequently Asked Questions

### How does Semantica handle conflicts when the same decision appears in both vector and graph results?

When `DecisionQuery.find_precedents_hybrid()` detects duplicate decisions across vector and graph result sets, it reconciles the scores by selecting the higher similarity value. The merged list maintains unique entries sorted by descending relevance score, ensuring the strongest signal—whether from semantic similarity or explicit graph relationships—determines the final ranking.

### Can I use Semantica's precedent search without a vector database?

Yes. The `AgentContext.find_precedents()` method in [`semantica/context_graph.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/context_graph.py) performs graph-only retrieval using `PRECEDENT_FOR` edge traversal without invoking the embedding provider or vector store. This mode requires only the `KnowledgeGraph` backend (Neo4j, ArangoDB, or in-memory) and skips semantic similarity entirely, relying solely on explicit precedent links recorded during decision storage.

### What embedding providers does Semantica support for precedent search?

Semantica's `EmbeddingProvider` interface supports pluggable implementations including **OpenAI**, **Cohere**, **HuggingFace**, and custom providers. The provider is injected into `DecisionQuery` or `ContextRetriever` at construction time, allowing you to switch between cloud-hosted and local embedding models without modifying the precedent search logic in [`decision_query.py`](https://github.com/semantica-agi/semantica/blob/main/decision_query.py) or [`retriever.py`](https://github.com/semantica-agi/semantica/blob/main/retriever.py).

### How does the graph traversal respect category filters?

The `find_precedents()` method accepts optional `category` and `limit` parameters that constrain the `PRECEDENT_FOR` edge traversal in [`semantica/context_graph.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/context_graph.py). When specified, the graph query filters nodes by category labels before traversing precedent relationships, ensuring only decisions within the specified domain context are considered relevant for the search results.