How FTSS Full-Text Search Combines Keyword and Vector Similarity
FTSS Hybrid Search Architecture: Blend SQLite FTS5 keyword filtering (BM25 scoring) with cosine vector similarity using a configurable weighted sum (α * bm25Score + (1 - α) * cosineScore) for semantic + lexical ranking.
FTSS (Full-Text Search System) in the tirth8205/code-review-graph repository powers intelligent code graph exploration by merging traditional keyword matching with modern embedding-based similarity. This hybrid approach ensures developers find syntactically relevant results while capturing semantic intent—even when exact keywords don't match.
FTSS Architecture Overview
The search pipeline operates in four distinct phases implemented across the backend modules. Understanding this flow reveals how lexical and semantic signals complement each other.
Phase 1: FTS5 Keyword Filtering
The system first leverages SQLite's native FTS5 extension for rapid candidate selection. In src/backend/sqlite.ts, the MATCH query executes against the nodes_fts virtual table, returning nodes containing search terms with BM25 relevance scores.
// src/backend/sqlite.ts – FTS5 candidate retrieval
const candidates = await db.prepare(`
SELECT rowid, bm25(fts) AS bm25Score, embedding
FROM nodes_fts
WHERE nodes_fts MATCH ?
`).all(searchTerms);
This coarse filter eliminates irrelevant nodes before expensive vector computations begin. BM25 scoring rewards term frequency and field length normalization, providing a strong lexical baseline.
Phase 2: Embedding Generation
FTSS generates dense vector representations for both stored nodes and the incoming query. The query string passes through the embedding model (typically OpenAI text-embedding-ada-002) to produce a query vector comparable against candidate embeddings.
// Typical embedding flow in search pipeline
const queryVec = await embed(queryText); // 1536-dimensional vector
Node embeddings are pre-computed during indexing and stored as BLOB columns in the SQLite database, enabling retrieval without repeated model calls.
Phase 3: Cosine Similarity Computation
For each FTS5 candidate, FTSS calculates semantic closeness using cosine similarity between the query vector and stored node embedding.
// src/backend/sqlite.ts – similarity calculation
const cosineScore = cosineSimilarity(queryVec, c.embedding);
Cosine similarity measures directional alignment regardless of vector magnitude, normalizing to [-1, 1] and typically scaled to [0, 1] for score fusion.
Phase 4: Weighted Score Fusion
The decisive operation combines both signals through a configurable blend factor α (default 0.7), producing the final ranking score:
// Core fusion formula implemented in search orchestration
const finalScore = α * bm25Score + (1 - α) * cosineScore;
Nodes ranking high in both keyword relevance and semantic similarity surface first. Pure keyword matches without semantic alignment—or semantically related nodes lacking keyword overlap—receive appropriately dampened scores.
Implementing FTSS Hybrid Search
Basic API Usage
The high-level ftssSearch function in src/features/search.ts encapsulates the entire pipeline:
import { ftssSearch } from '@/features/search';
// Hybrid search: "auth" keywords + semantic "user login flow"
const results = await ftssSearch({
query: 'auth',
embedQuery: true, // Enable vector similarity
blendAlpha: 0.6, // 60% keyword, 40% vector weight
limit: 20,
});
results.forEach(r => {
console.log(`${r.nodeId} – fused score: ${r.score.toFixed(3)}`);
});
Adjusting Semantic Emphasis
Reduce blendAlpha to prioritize meaning over exact term matching:
// Emphasize semantic similarity for conceptual exploration
await ftssSearch({
query: 'database schema',
embedQuery: true,
blendAlpha: 0.3, // 30% keyword, 70% vector weight
});
Lower alpha values excel for exploratory queries where vocabulary varies but intent aligns—finding "data model" when searching "table structure," for example.
Low-Level Pipeline Access
Direct database and embedding module access enables custom fusion strategies:
// Manual orchestration bypassing the high-level API
import { db } from '@/backend/sqlite';
import { embed, cosineSimilarity } from '@/backend/embedding';
const customSearch = async (query: string, candidates: number = 50) => {
const rows = await db.prepare(`
SELECT rowid, bm25(fts, 1.2, 0.75) AS bm25Score, embedding
FROM nodes_fts
WHERE nodes_fts MATCH ?
ORDER BY bm25Score DESC
LIMIT ?
`).all(query, candidates);
const qVec = await embed(query);
return rows.map(r => ({
...r,
cosine: cosineSimilarity(qVec, r.embedding),
score: 0.5 * r.bm25Score + 0.5 * r.cosine // Custom 50/50 blend
})).sort((a, b) => b.score - a.score);
};
Core Implementation Files
| File | Purpose | Key Exports |
|---|---|---|
src/features/search.ts |
Public API orchestrating the complete FTSS pipeline | ftssSearch(), SearchOptions interface |
src/backend/sqlite.ts |
SQLite/FTS5 database layer, BM25 retrieval, score computation | Database connection, bm25() scoring |
src/backend/embedding.ts |
Embedding model client and similarity utilities | embed(), cosineSimilarity() |
src/views/graphWebview.ts |
VS Code webview rendering search results | Result visualization, node highlighting |
All paths reference the code-review-graph-vscode directory within the repository as implemented in tirth8205/code-review-graph.
Tuning FTSS for Your Dataset
Keyword-Heavy Corpora — Increase blendAlpha (0.8–0.9) when precise terminology dominates, such as API reference searches.
Conceptually Varied Content — Decrease blendAlpha (0.2–0.4) for design discussions, architecture reviews, or natural language documentation where paraphrasing is common.
Performance Constraints — The FTS5 filter dramatically reduces vector computation scope. For latency-critical paths, lower the candidate LIMIT in the initial query before embedding comparison.
Summary
- FTSS combines signals through SQLite FTS5 BM25 scoring and embedding cosine similarity in a weighted fusion model
- Configurable blending via
blendAlphaparameter tunes lexical vs. semantic priority per query context - Pipeline efficiency relies on FTS5 pre-filtering to minimize expensive vector operations
- Implementation spans
src/features/search.ts(orchestration),src/backend/sqlite.ts(storage/similarity), andsrc/backend/embedding.ts(vector generation)
Frequently Asked Questions
How does FTSS handle queries with no keyword matches?
FTSS returns empty result sets when FTS5 MATCH yields zero candidates. Unlike pure vector systems, the hybrid architecture requires at least partial lexical overlap. To enable pure semantic search, bypass ftssSearch and query the embeddings table directly using cosineSimilarity against all nodes—trading recall for computational cost.
What embedding model does FTSS use by default?
The default configuration targets OpenAI's text-embedding-ada-002 (1536 dimensions) as implemented in src/backend/embedding.ts. The modular design permits swapping to local models (sentence-transformers, Ollama) by reimplementing the embed() function signature without changing downstream consumers.
Can FTSS search across multiple node types simultaneously?
Yes. The FTS5 virtual table nodes_fts indexes heterogeneous node contents (functions, classes, comments) into a unified searchable field. The rowid joins back to type-specific tables post-ranking. Filter by node type after scoring via WHERE type = 'function' clauses, or include type metadata in the FTS5 indexed columns for query-time filtering.
Why cosine similarity over dot product or Euclidean distance?
Cosine similarity ignores vector magnitude, preventing longer documents from dominating purely due to embedding scale variations. This normalization proves essential when mixing code snippets (short) with documentation blocks (long) in the same search space. Dot products require explicit L2-normalization; Euclidean distance penalizes magnitude differences that may carry meaningful semantic information.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →