# How the CRG Query Engine Works: A Deep Dive into Hybrid Code Search in Code‑Review‑Graph

> Discover how the CRG query engine uses hybrid code search, combining FTS5 BM25 and vector embeddings with Reciprocal Rank Fusion, for advanced context-aware code navigation in Code-Review-Graph.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: deep-dive
- Published: 2026-08-10

---

**The CRG query engine combines FTS5 BM25 full‑text search with vector embeddings, fuses results via Reciprocal Rank Fusion, and applies multiple heuristic boosts to deliver context‑aware code navigation.**

The **Code‑Review‑Graph (CRG)** query engine powers intelligent code search across repositories parsed into a SQLite-based graph. Implemented in [`code_review_graph/search.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/search.py), it bridges classic information retrieval with modern semantic search. This article breaks down the engine's five‑phase pipeline, its hybrid architecture, and how you can leverage its API for precise code discovery.

---

## Query Preprocessing: Extracting Intent from Natural Language

Before any database query executes, the engine analyzes the raw search string to extract structural cues. Two functions drive this phase in [`code_review_graph/search.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/search.py):

- **`extract_query_identifiers`** (lines 81–100) parses **dotted paths** (`Context.Next`), **snake_case** (`get_user`), and **PascalCase** (`MyClass`) identifiers from the query.
- **`detect_query_kind_boost`** (lines 101–144) applies heuristics to infer which **node kinds**—`Class`, `Function`, `Method`, `Variable`, etc.—deserve relevance boosting.

This preprocessing ensures that a query like `"Session manager authenticate"` automatically elevates `Class` and `Function` nodes containing those terms, without requiring explicit filters.

---

## Three Parallel Search Back‑Ends

The engine executes three search strategies, falling back gracefully when dependencies are unavailable:

| Back‑end | Implementation | When It Runs |
|----------|---------------|--------------|
| **FTS5 BM25** | `_fts_search` (lines 182–206) | Always if the `nodes_fts` virtual table exists; ranks by lexical match |
| **Embedding search** | `_embedding_search` (lines 213–250) | When an `EmbeddingStore` is configured; ranks by cosine similarity |
| **Keyword LIKE** | `_keyword_search` (lines 256–301) | Fallback if FTS5 is unavailable or embeddings are disabled |

The **FTS5** virtual table provides millisecond‑scale full‑text retrieval over node names, docstrings, and source snippets. The **embedding** path (via [`code_review_graph/embeddings.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/embeddings.py)) enables semantic matching—finding "authentication" when the query says "login". If neither is viable, the **LIKE** fallback ensures the CLI remains functional.

---

## Result Fusion with Reciprocal Rank Fusion (RRF)

Two ranked lists (FTS5 and embeddings) require merging without fragile weight tuning. The engine uses **Reciprocal Rank Fusion**:

```python

# From search.py lines 151-176

def rrf_merge(fts_results, embedding_results, k=60):
    scores = {}
    for rank, item in enumerate(fts_results):
        scores[item["id"]] = scores.get(item["id"], 0) + 1 / (k + rank + 1)
    for rank, item in enumerate(embedding_results):
        scores[item["id"]] = scores.get(item["id"], 0) + 1 / (k + rank + 1)
    # Return merged, deduplicated list sorted by fused score

```

RRF's **rank‑based discounting** (the `k=60` constant dampens top‑rank dominance) means a strong lexical hit and a strong semantic hit both surface effectively, even if they appear at different positions in their respective lists.

---

## Batch Fetching and Multi‑Factor Score Boosting

After fusion, the engine fetches all candidate nodes in a **single SQL batch**—critical for performance on large codebases. Four boost multipliers then refine the RRF scores (lines 384–438):

1. **Kind‑specific boost** (lines 384–403): `Class` nodes may receive `2.0×`, `Function` nodes `1.5×`, based on `detect_query_kind_boost` heuristics.

2. **Qualified‑name match boost** (`_qualified`, lines 424–430): Exact matches against the full dotted path (`auth.session.SessionManager`) receive elevated scores.

3. **Identifier‑match boost** (`_qualified_identifiers`, lines 430–435): Matches against extracted identifiers from the preprocessing phase.

4. **Context‑file boost** (lines 436–438): Nodes inside files passed via `context_files` receive a **1.5× multiplier**—ideal for IDE integrations where open files indicate developer focus.

These boosts are **multiplicative and composable**, allowing fine‑grained relevance tuning without hard filtering.

---

## Filtering and Final Result Assembly

The `hybrid_search` function (lines 309–470) performs final orchestration:

- Applies optional `kind` filtering (e.g., `kind="Class"` discards non‑class nodes).
- Sorts by boosted score descending.
- Returns the top `limit` results as dictionaries with:
  - `name`: Short identifier
  - `qualified_name`: Full dotted path
  - `kind`: Node type
  - `location`: File path and line number
  - `language`: Detected programming language
  - `score`: Final computed relevance

---

## Practical Usage: Calling the Search API

```python
from code_review_graph.search import hybrid_search
from code_review_graph.graph import GraphStore

# Initialize connection to a parsed repository

store = GraphStore("my_repo.graphdb")

# Default hybrid search (FTS + embeddings if available)

results = hybrid_search(store, "authenticate")
print(results[0]["qualified_name"], results[0]["score"])

# Constrained search with file context boosting

results = hybrid_search(
    store,
    "Session",
    kind="Class",
    context_files=["src/auth/session.py"],
    limit=5,
)
for hit in results:
    print(f"{hit['kind']}: {hit['qualified_name']} (score={hit['score']:.3f})")

```

The second call demonstrates **contextual awareness**: nodes in [`src/auth/session.py`](https://github.com/tirth8205/code-review-graph/blob/main/src/auth/session.py) receive preferential scoring, mimicking IDE "files of interest" behavior.

---

## Architectural Strengths of the CRG Query Engine

| Capability | Implementation Benefit |
|-----------|------------------------|
| **Hybrid retrieval** | Combines exact identifier matching (FTS5) with conceptual similarity (embeddings) |
| **Parameter‑less fusion** | RRF eliminates need to tune α/β weights for lexical vs. semantic balance |
| **Query‑driven boosting** | Heuristic detection of code patterns (PascalCase, dotted paths) improves recall |
| **Context sensitivity** | `context_files` parameter enables workspace‑aware results |
| **Resilient operation** | LIKE fallback ensures functionality without FTS5 or embedding services |

---

## Summary

- The **CRG query engine** in [`code_review_graph/search.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/search.py) implements a **five‑phase pipeline**: preprocessing, parallel search back‑ends, RRF fusion, batch fetching with multi‑factor boosting, and final filtering.
- **FTS5 BM25**, **embedding search**, and **LIKE fallback** provide speed, semantics, and resilience.
- **Reciprocal Rank Fusion** merges heterogeneous rankings without fragile hyperparameters.
- **Heuristic boosts** for node kinds, qualified names, identifiers, and context files tailor results to developer intent.
- The **`hybrid_search`** API exposes all functionality through a single call with optional `kind` and `context_files` parameters.

---

## Frequently Asked Questions

### What makes the CRG query engine "hybrid"?

The engine runs **FTS5 BM25** and **vector embedding search** simultaneously, then fuses their rankings. This hybrid approach captures both exact identifier matches (critical for code) and semantic paraphrases (useful for natural‑language queries like "how to log in").

### How does the engine handle queries without exact keyword matches?

When embeddings are configured via `EmbeddingStore`, `_embedding_search` computes cosine similarity between the query vector and pre‑computed node embeddings. This surfaces conceptually related code even when lexical overlap is minimal.

### Why does the engine use Reciprocal Rank Fusion instead of score‑weighted averaging?

RRF depends only on **rank positions**, not raw scores. This avoids calibration problems: BM25 scores (unbounded, document‑length normalized) and cosine similarities (bounded [−1, 1]) have incompatible scales. RRF's `1/(k+rank)` formula robustly interleaves both lists.

### What happens if SQLite FTS5 is unavailable or embeddings aren't installed?

The engine gracefully degrades to `_keyword_search` (lines 256–301), which executes `LIKE '%term%'` queries against the nodes table. Performance is slower and ranking coarser, but the search remains functional—ensuring CLI tools never fail completely.