How Semantic Search Works with Hybrid FTS5 and Vector Similarity in code-review-graph

Semantic search in code-review-graph combines SQLite FTS5 full-text search with vector similarity scoring to deliver fast, context-aware code discovery that works whether or not embeddings have been generated.

The code-review-graph repository implements a sophisticated hybrid search architecture that bridges traditional keyword matching and modern embedding-based retrieval. Understanding this dual-mode system helps developers optimize query performance and result relevance across code review knowledge graphs.

Overview of the Hybrid Search Architecture

The search implementation spans two core modules. The entry point lives in [code_review_graph/tools/query.py](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/tools/query.py) with the semantic_search_nodes function, while the vector operations reside in [code_review_graph/embeddings.py](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/embeddings.py) through the semantic_search function.

This design reflects a pragmatic approach to code intelligence: FTS5 provides deterministic, fast keyword matches while vector similarity captures semantic relationships between concepts that share no surface-level text overlap.

Step 1: Query Preprocessing and Embedding Generation

When semantic_search_nodes receives a query string, it first normalizes the input and checks whether the embeddings subsystem is initialized.


# From code_review_graph/tools/query.py

# The semantic_search_nodes function prepares the query and branches based on

# embedding availability

If embeddings are available, the system generates a query vector using the model configuration established during graph initialization. This vector becomes the basis for similarity comparisons against stored node embeddings in the SQLite semantic_embeddings table.

Step 2: Vector-Based Retrieval

The semantic_search function in [embeddings.py](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/embeddings.py) executes the core similarity computation:

  • Loads the query vector into the SQLite vector extension
  • Performs cosine similarity against all stored node embeddings
  • Returns ranked node IDs with their similarity scores
  • Applies configurable thresholds via min_score and max_results parameters

The vector table is created and populated by the embed_graph_tool, which runs lazily—meaning the hybrid system automatically adapts when embeddings become available without requiring explicit configuration changes.

Step 3: FTS5 Full-Text Fallback and Parallel Execution

Simultaneously, the original query string routes through SQLite's FTS5 virtual table. This parallel execution guarantees results even when:

  • Embeddings have not yet been generated for the repository
  • The vector index is incomplete due to partial processing
  • The query contains exact terms that benefit from precise matching

The FTS5 component leverages SQLite's built-in ranking algorithms, providing deterministic performance characteristics that complement the probabilistic nature of vector similarity.

Step 4: Score Fusion and Result Merging

The critical innovation occurs in the merging logic within semantic_search_nodes. The system combines three possible result categories:

  • Nodes present in both vectors and FTS5: Scores are fused using configurable weighting
  • Vector-only matches: Retain cosine similarity scores
  • FTS5-only matches: Retain BM25-derived relevance scores

This fusion strategy ensures that semantic matches augment rather than replace keyword matches, preventing the "hallucination" problem where purely vector-based systems return conceptually related but irrelevant code.

Step 5: Provenance Enrichment and Response Formatting

The final processing stage attaches metadata essential for downstream consumption:

  • source: Origin repository or file path
  • node_type: Function, class, module, or documentation
  • match_type: Whether the result came from vector, FTS5, or both

Results respect limit and offset parameters for pagination, returning data in the standardized CRG tool contract format that integrates seamlessly with the broader tool ecosystem.

CLI and Python API Usage


# Enable semantic search alongside standard graph querying

code-review-graph serve --tools query_graph_tool,semantic_search_nodes_tool

Basic Python API Call

from code_review_graph.tools.query import semantic_search_nodes

# Hybrid search (default): uses vectors if available, FTS5 otherwise

results = semantic_search_nodes("authentication middleware", limit=10)

Explicit Search Mode Control


# Force pure vector similarity (requires pre-built embeddings)

vector_results = semantic_search_nodes(
    "user session management",
    limit=10,
    search_mode="vector"
)

# Force pure FTS5 keyword search

fts_results = semantic_search_nodes(
    "user session management",
    limit=10,
    search_mode="fts"
)

# Explicit hybrid with custom weights (implementation-specific)

hybrid_results = semantic_search_nodes(
    "user session management",
    limit=10,
    search_mode="hybrid"
)

Handling Fallback Scenarios Programmatically

results = semantic_search_nodes("payment gateway integration", limit=10)

if not results["nodes"]:
    # No vector matches—explicitly pivot to FTS5

    results = semantic_search_nodes(
        "payment gateway integration",
        search_mode="fts",
        limit=20  # Expand limit for keyword-only results

    )

Key Design Characteristics

Feature Implementation Benefit
Lazy embedding building Vector index constructed on-demand via embed_graph_tool No blocking operations during initial import
Thread-safe caching Per-process embedding matrix cache Eliminates redundant model loading
Configurable thresholds min_score, max_results parameters Precision-recall tradeoff control
Deterministic fallback FTS5 always executes Guaranteed result availability

File Reference Guide

Path Purpose
[code_review_graph/tools/query.py](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/tools/query.py) semantic_search_nodes hybrid orchestration
[code_review_graph/embeddings.py](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/embeddings.py) semantic_search vector similarity, embedding storage
[docs/COMMANDS.md](https://github.com/tirth8205/code-review-graph/blob/main/docs/COMMANDS.md) CLI documentation for semantic_search_nodes_tool
[docs/USAGE.md](https://github.com/tirth8205/code-review-graph/blob/main/docs/USAGE.md) Automatic vector search behavior description

Summary

  • Hybrid search in code-review-graph combines FTS5 keyword matching with vector cosine similarity for robust code discovery.
  • The semantic_search_nodes function in query.py orchestrates parallel execution and result fusion.
  • Vector operations in embeddings.py leverage SQLite's vector extension with lazy index building and thread-safe caching.
  • Three search modes—"vector", "fts", and "hybrid"—provide explicit control over retrieval behavior.
  • The architecture automatically degrades gracefully when embeddings are unavailable, ensuring consistent API behavior.

Frequently Asked Questions

How do I know if vector search is actually being used?

Check whether embeddings exist by attempting a search_mode="vector" query. If empty results return, the vector index has not been built. Run embed_graph_tool to generate embeddings, after which hybrid search will automatically incorporate similarity scores.

What embedding model does code-review-graph use?

The model configuration is set during graph initialization in embeddings.py. The default uses a sentence-transformer compatible with SQLite's vector extension, though specific model details depend on the initialization parameters passed to the embedding subsystem.

Can I adjust the weighting between vector and FTS5 scores?

Yes. The semantic_search_nodes function accepts parameters that control score fusion. While the exact parameter names depend on the implementation version, the merging logic in query.py supports configurable weighting when both vector and FTS5 scores exist for a node.

Why would I ever use search_mode="fts" explicitly?

Use pure FTS5 when you need deterministic, reproducible results based on exact keyword presence—particularly useful for compliance queries, precise identifier searches, or when debugging why certain results appear in hybrid mode.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →