# Implementing Hybrid Search: Dense Vectors and Sparse Retrieval in Semantica

> Master hybrid search by combining dense vectors and sparse retrieval in Semantica. Achieve precise, context-aware results balancing semantic relevance and filtering.

- Repository: [Semantica /semantica](https://github.com/semantica-agi/semantica)
- Tags: how-to-guide
- Published: 2026-09-11

---

**Semantica's hybrid search engine fuses dense vector similarity with composable metadata predicates in [`semantica/vector_store/hybrid_search.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/vector_store/hybrid_search.py) to deliver precise, context-aware retrieval that balances semantic relevance with deterministic filtering.**

Semantica implements a production-grade hybrid search system that combines dense vector embeddings with sparse metadata filtering to improve recall and precision when querying knowledge-graph elements. The implementation centers on three core classes—`MetadataFilter`, `HybridSearch`, and `SearchRanker`—which orchestrate dual-stage retrieval and result fusion. This architecture enables developers to enforce strict business rules through metadata constraints while preserving the flexible semantic matching that dense embeddings provide.

## Core Architecture of Hybrid Search

The hybrid search implementation in [`semantica/vector_store/hybrid_search.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/vector_store/hybrid_search.py) separates concerns across three specialized components, each handling a distinct aspect of the retrieval pipeline.

### MetadataFilter for Sparse Retrieval

The `MetadataFilter` class (defined beginning at line 47) constructs composable predicates for evaluating document metadata through chainable methods including `eq`, `gt`, `lt`, and `contains`. These condition builders (lines 47‑99) construct filter trees that the `matches` method (lines 101‑136) evaluates against candidate metadata dictionaries, discarding any vectors that violate the specified constraints.

### HybridSearch Orchestration Layer

The `HybridSearch` class (line 17) serves as the primary entry point for hybrid queries, coordinating the dual-stage retrieval process. It first executes dense vector similarity searches against the underlying backend—whether Qdrant, Weaviate, or PGVector—then applies `MetadataFilter` instances to the candidate set. The class also supports `multi_source_search` (illustrated at lines 31‑33), which aggregates results from multiple vector sources before fusion.

### SearchRanker for Result Fusion

After filtering, the `SearchRanker` merges ranked lists from different sources using configurable fusion strategies. The default **Reciprocal Rank Fusion (RRF)** implementation (lines 48‑70) calculates scores as `1 / (k + rank)` for each candidate across all result sets, summing these values to produce a final ranking. For weighted scoring scenarios, the class provides a weighted-average fusion method beginning at line 88.

## How Hybrid Search Works

The retrieval pipeline executes in four distinct stages, combining neural semantic search with deterministic rule evaluation.

1. **Vector Retrieval** – The system submits the query vector to the configured vector store backend, retrieving the top-k nearest neighbors with their similarity scores.

2. **Metadata Filtering** – The `MetadataFilter.matches` method iterates through the retrieved candidates, checking each document's metadata dictionary against the constructed predicate tree.

3. **Result Fusion** – Surviving candidates enter the `SearchRanker`, which applies RRF or weighted-average strategies to combine rankings from multiple sources or modalities.

4. **Multi-Source Aggregation** – When querying across heterogeneous backends, `multi_source_search` normalizes results from each source before fusion, ensuring consistent ranking despite varying vector spaces.

## Implementation Examples

The following examples demonstrate practical usage of the hybrid search API, from basic filtered queries to multi-backend fusion.

### Basic Hybrid Search with Metadata Filtering

```python
from semantica.vector_store import HybridSearch, MetadataFilter

# Initialize the hybrid engine

search = HybridSearch()

# Build a composable metadata filter

meta_filter = (
    MetadataFilter()
    .eq("category", "research")
    .gt("published_year", 2018)
)

# Execute hybrid search (vectors retrieved from configured backend)

results = search.search(
    query_vector=my_query_vector,
    filter=meta_filter,
    k=15,
)

```

### Fusing Results from Multiple Backends

```python
from semantica.vector_store import SearchRanker

# Initialize ranker with Reciprocal Rank Fusion

ranker = SearchRanker(strategy="reciprocal_rank_fusion")

# Combine results from Qdrant and PGVector

fused = ranker.rank(
    [results_from_qdrant, results_from_pgvector], 
    k=20
)

```

### Multi-Source Hybrid Search

```python

# Define heterogeneous sources with distinct vector sets

sources = [
    {"vectors": vecs_a, "metadata": meta_a, "ids": ids_a},
    {"vectors": vecs_b, "metadata": meta_b, "ids": ids_b},
]

# Query across all sources with unified filtering

multi_results = search.multi_source_search(
    query_vector=my_query_vector,
    sources=sources,
    filter=meta_filter,
    k=10,
)

```

## Key Supporting Files

The hybrid search system relies on several additional modules within the `semantica/vector_store/` package:

- **[`hybrid_similarity.py`](https://github.com/semantica-agi/semantica/blob/main/hybrid_similarity.py)** – Contains `HybridSimilarityCalculator`, which blends semantic and structural similarity scores for internal ranking adjustments.
- **[`vector_store.py`](https://github.com/semantica-agi/semantica/blob/main/vector_store.py)** – Defines the abstract base class for vector store backends that `HybridSearch` utilizes for dense retrieval.
- **[`metadata_store.py`](https://github.com/semantica-agi/semantica/blob/main/metadata_store.py)** – Persists metadata dictionaries accessed by the `MetadataFilter.matches` method during candidate evaluation.
- **[`registry.py`](https://github.com/semantica-agi/semantica/blob/main/registry.py)** – Registers available backend implementations, enabling runtime selection of Qdrant, PGVector, or Weaviate connectors.

## Summary

- **HybridSearch** in [`semantica/vector_store/hybrid_search.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/vector_store/hybrid_search.py) orchestrates dual-stage retrieval, combining dense vector similarity with sparse metadata filtering.
- **MetadataFilter** provides chainable methods (`eq`, `gt`, `lt`, `contains`) for constructing complex predicates evaluated at lines 101‑136.
- **SearchRanker** implements Reciprocal Rank Fusion (lines 48‑70) and weighted-average strategies to merge results from multiple sources or backends.
- The architecture supports multi-source queries through `multi_source_search`, enabling consistent hybrid retrieval across heterogeneous vector stores.
- Result fusion balances semantic relevance from embeddings with deterministic business rules from metadata constraints.

## Frequently Asked Questions

### What is the difference between dense and sparse retrieval in Semantica?

Dense retrieval utilizes vector embeddings to find semantically similar content based on neural network representations, while sparse retrieval refers to the metadata filtering system that applies exact-match predicates, ranges, and set inclusions through the `MetadataFilter` class. The hybrid approach combines these modalities, using dense vectors for semantic recall and sparse filters for precision control.

### How does Reciprocal Rank Fusion work in the SearchRanker?

The RRF implementation in lines 48‑70 of [`hybrid_search.py`](https://github.com/semantica-agi/semantica/blob/main/hybrid_search.py) assigns each candidate a score of `1 / (k + rank)` where `k` is a constant (typically 60) and `rank` is the position in a given result list. The algorithm sums these scores across all sources, producing a fused ranking that reduces bias toward any single retrieval method while preserving high-ranked items from any source.

### Can I use multiple vector backends simultaneously with hybrid search?

Yes, the `multi_source_search` method (demonstrated at lines 31‑33) accepts a list of sources containing distinct vectors, metadata, and IDs, allowing queries to span Qdrant, PGVector, Weaviate, or custom backends simultaneously. The system normalizes results from each source before applying the shared `MetadataFilter` and fusing rankings through the configured `SearchRanker` strategy.

### How do I construct complex metadata filters with multiple conditions?

Chain the condition builder methods on a `MetadataFilter` instance to create conjunctive predicates; for example, `MetadataFilter().eq("status", "active").gt("priority", 5).contains("tags", "urgent")` creates a filter requiring all three conditions. The `matches` method evaluates these against document metadata using the logic defined at lines 101‑136, supporting nested boolean logic for sophisticated filtering scenarios.