How the CRG Query Engine Works: A Deep Dive into Hybrid Code Search in Code‑Review‑Graph
The CRG query engine combines FTS5 BM25 full‑text search with vector embeddings, fuses results via Reciprocal Rank Fusion, and applies multiple heuristic boosts to deliver context‑aware code navigation.
The Code‑Review‑Graph (CRG) query engine powers intelligent code search across repositories parsed into a SQLite-based graph. Implemented in code_review_graph/search.py, it bridges classic information retrieval with modern semantic search. This article breaks down the engine's five‑phase pipeline, its hybrid architecture, and how you can leverage its API for precise code discovery.
Query Preprocessing: Extracting Intent from Natural Language
Before any database query executes, the engine analyzes the raw search string to extract structural cues. Two functions drive this phase in code_review_graph/search.py:
extract_query_identifiers(lines 81–100) parses dotted paths (Context.Next), snake_case (get_user), and PascalCase (MyClass) identifiers from the query.detect_query_kind_boost(lines 101–144) applies heuristics to infer which node kinds—Class,Function,Method,Variable, etc.—deserve relevance boosting.
This preprocessing ensures that a query like "Session manager authenticate" automatically elevates Class and Function nodes containing those terms, without requiring explicit filters.
Three Parallel Search Back‑Ends
The engine executes three search strategies, falling back gracefully when dependencies are unavailable:
| Back‑end | Implementation | When It Runs |
|---|---|---|
| FTS5 BM25 | _fts_search (lines 182–206) |
Always if the nodes_fts virtual table exists; ranks by lexical match |
| Embedding search | _embedding_search (lines 213–250) |
When an EmbeddingStore is configured; ranks by cosine similarity |
| Keyword LIKE | _keyword_search (lines 256–301) |
Fallback if FTS5 is unavailable or embeddings are disabled |
The FTS5 virtual table provides millisecond‑scale full‑text retrieval over node names, docstrings, and source snippets. The embedding path (via code_review_graph/embeddings.py) enables semantic matching—finding "authentication" when the query says "login". If neither is viable, the LIKE fallback ensures the CLI remains functional.
Result Fusion with Reciprocal Rank Fusion (RRF)
Two ranked lists (FTS5 and embeddings) require merging without fragile weight tuning. The engine uses Reciprocal Rank Fusion:
# From search.py lines 151-176
def rrf_merge(fts_results, embedding_results, k=60):
scores = {}
for rank, item in enumerate(fts_results):
scores[item["id"]] = scores.get(item["id"], 0) + 1 / (k + rank + 1)
for rank, item in enumerate(embedding_results):
scores[item["id"]] = scores.get(item["id"], 0) + 1 / (k + rank + 1)
# Return merged, deduplicated list sorted by fused score
RRF's rank‑based discounting (the k=60 constant dampens top‑rank dominance) means a strong lexical hit and a strong semantic hit both surface effectively, even if they appear at different positions in their respective lists.
Batch Fetching and Multi‑Factor Score Boosting
After fusion, the engine fetches all candidate nodes in a single SQL batch—critical for performance on large codebases. Four boost multipliers then refine the RRF scores (lines 384–438):
-
Kind‑specific boost (lines 384–403):
Classnodes may receive2.0×,Functionnodes1.5×, based ondetect_query_kind_boostheuristics. -
Qualified‑name match boost (
_qualified, lines 424–430): Exact matches against the full dotted path (auth.session.SessionManager) receive elevated scores. -
Identifier‑match boost (
_qualified_identifiers, lines 430–435): Matches against extracted identifiers from the preprocessing phase. -
Context‑file boost (lines 436–438): Nodes inside files passed via
context_filesreceive a 1.5× multiplier—ideal for IDE integrations where open files indicate developer focus.
These boosts are multiplicative and composable, allowing fine‑grained relevance tuning without hard filtering.
Filtering and Final Result Assembly
The hybrid_search function (lines 309–470) performs final orchestration:
- Applies optional
kindfiltering (e.g.,kind="Class"discards non‑class nodes). - Sorts by boosted score descending.
- Returns the top
limitresults as dictionaries with:name: Short identifierqualified_name: Full dotted pathkind: Node typelocation: File path and line numberlanguage: Detected programming languagescore: Final computed relevance
Practical Usage: Calling the Search API
from code_review_graph.search import hybrid_search
from code_review_graph.graph import GraphStore
# Initialize connection to a parsed repository
store = GraphStore("my_repo.graphdb")
# Default hybrid search (FTS + embeddings if available)
results = hybrid_search(store, "authenticate")
print(results[0]["qualified_name"], results[0]["score"])
# Constrained search with file context boosting
results = hybrid_search(
store,
"Session",
kind="Class",
context_files=["src/auth/session.py"],
limit=5,
)
for hit in results:
print(f"{hit['kind']}: {hit['qualified_name']} (score={hit['score']:.3f})")
The second call demonstrates contextual awareness: nodes in src/auth/session.py receive preferential scoring, mimicking IDE "files of interest" behavior.
Architectural Strengths of the CRG Query Engine
| Capability | Implementation Benefit |
|---|---|
| Hybrid retrieval | Combines exact identifier matching (FTS5) with conceptual similarity (embeddings) |
| Parameter‑less fusion | RRF eliminates need to tune α/β weights for lexical vs. semantic balance |
| Query‑driven boosting | Heuristic detection of code patterns (PascalCase, dotted paths) improves recall |
| Context sensitivity | context_files parameter enables workspace‑aware results |
| Resilient operation | LIKE fallback ensures functionality without FTS5 or embedding services |
Summary
- The CRG query engine in
code_review_graph/search.pyimplements a five‑phase pipeline: preprocessing, parallel search back‑ends, RRF fusion, batch fetching with multi‑factor boosting, and final filtering. - FTS5 BM25, embedding search, and LIKE fallback provide speed, semantics, and resilience.
- Reciprocal Rank Fusion merges heterogeneous rankings without fragile hyperparameters.
- Heuristic boosts for node kinds, qualified names, identifiers, and context files tailor results to developer intent.
- The
hybrid_searchAPI exposes all functionality through a single call with optionalkindandcontext_filesparameters.
Frequently Asked Questions
What makes the CRG query engine "hybrid"?
The engine runs FTS5 BM25 and vector embedding search simultaneously, then fuses their rankings. This hybrid approach captures both exact identifier matches (critical for code) and semantic paraphrases (useful for natural‑language queries like "how to log in").
How does the engine handle queries without exact keyword matches?
When embeddings are configured via EmbeddingStore, _embedding_search computes cosine similarity between the query vector and pre‑computed node embeddings. This surfaces conceptually related code even when lexical overlap is minimal.
Why does the engine use Reciprocal Rank Fusion instead of score‑weighted averaging?
RRF depends only on rank positions, not raw scores. This avoids calibration problems: BM25 scores (unbounded, document‑length normalized) and cosine similarities (bounded [−1, 1]) have incompatible scales. RRF's 1/(k+rank) formula robustly interleaves both lists.
What happens if SQLite FTS5 is unavailable or embeddings aren't installed?
The engine gracefully degrades to _keyword_search (lines 256–301), which executes LIKE '%term%' queries against the nodes table. Performance is slower and ranking coarser, but the search remains functional—ensuring CLI tools never fail completely.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →