How the FTS5 Knowledge Base in context-mode Combines Porter Stemming and Trigram Search
The context-mode FTS5 knowledge base stores every text chunk in two separate virtual tables—one using Porter stemming for morphological matching and another using trigrams for substring search—then merges results with Reciprocal Rank Fusion (RRF) to ensure high recall across exact and stemmed queries.
The mksglu/context-mode repository implements a sophisticated full-text search system using SQLite FTS5. Its knowledge base architecture solves the classic search trade-off between stemming accuracy and substring matching by maintaining dual indexes and executing a three-layer fallback pipeline that automatically selects the best results from either approach.
Dual-Index Architecture in src/store.ts
The system creates two distinct FTS5 virtual tables during initialization in src/store.ts (lines 1007–1021).
The chunks Table: Porter Stemming
The primary index uses the porter unicode61 tokenizer. This implementation applies Porter stemming to normalize words to their root forms, ensuring that queries like "running" match documents containing "run" or "runs".
The chunks_trigram Table: Substring Matching
A secondary index uses the trigram tokenizer. This table indexes every three-character sequence in the text, enabling substring search capabilities that Porter stemming cannot provide, such as matching camelCase identifiers ("responseBody") or partial word segments.
Both tables are populated simultaneously during the indexing phase via the #insertChunks method (lines 783–787), ensuring every chunk exists in both representations without duplication of the source text.
The Three-Layer Search Pipeline
When searchWithFallback (lines 639–674) executes, it orchestrates queries across both indexes and fuses the results.
Layer 1: Porter Stemmed Search
The pipeline first calls search (lines 462–475), which sanitizes the query and executes it against the chunks table. This layer excels at linguistic matches but fails on camelCase substrings or exact partial matches where stemming breaks the token.
Layer 2: Trigram Fallback Search
If Layer 1 returns insufficient results, the system invokes searchTrigram (lines 505–517) against the chunks_trigram table. This captures matches that the Porter tokenizer misses, such as code identifiers or technical terms where the user remembers only a substring.
Layer 3: RRF Fusion and Proximity Reranking
The #rrfSearch method (lines 549–590) implements Reciprocal Rank Fusion (RRF) to merge the two result sets. It assigns scores based on reciprocal rank positions, then applies a proximity model that boosts matches in titles and tightly clustered term occurrences. The final list represents the optimal combination of stemmed linguistic matches and exact substring hits.
Practical Code Examples
Indexing Content with Dual Tokenizers
import { ContentStore } from "./src/store.js";
const store = new ContentStore();
store.index({ path: "docs/guide.md", source: "guide" });
The index method delegates to #insertChunks, which inserts each chunk into both chunks and chunks_trigram tables.
Executing a Fallback Search
// Matches "responseBody" even if Porter stems "response" separately
const results = store.searchWithFallback("responseBody", 5);
console.log(results.map(r => `${r.title}: ${r.highlighted}`));
This executes #rrfSearch, combining Porter and trigram results through reciprocal rank fusion.
Debugging Individual Layers
// Porter-only results (linguistic matching)
const porterHits = store.search("caching", 5);
// Trigram-only results (substring matching)
const trigramHits = store.searchTrigram("responseBody", 5);
Direct access to search (lines 462–475) and searchTrigram (lines 505–517) enables performance analysis and debugging.
Summary
- Dual indexing strategy: Every chunk is stored in both
chunks(Porter) andchunks_trigram(trigram) tables during indexing. - Automatic fallback: The
searchWithFallbackmethod attempts Porter matching first, then supplements with trigram results when linguistic stemming fails. - RRF fusion: Results from both layers are merged using Reciprocal Rank Fusion in
#rrfSearch, then reranked by proximity and title relevance. - CamelCase support: The trigram layer ensures code identifiers and technical terms match even when Porter stemming breaks apart compound words.
Frequently Asked Questions
Why does context-mode use both Porter and trigram tokenizers?
Porter stemming excels at matching word variations (searching for "run" finds "running" and "runs"), but it fails on camelCase identifiers like "responseBody" where stemming breaks the token. The trigram tokenizer indexes every three-character sequence, enabling substring matches that capture these technical terms. Using both ensures high recall across natural language and code-focused content.
How does Reciprocal Rank Fusion (RRF) combine results from both indexes?
RRF assigns a score to each document based on its rank in each result list using the formula 1 / (k + rank), where k is a constant (typically 60). Documents appearing in both the Porter and trigram results receive higher combined scores than those appearing in only one list. The #rrfSearch method implements this fusion, then applies additional proximity-based reranking to prioritize title matches and tightly clustered terms.
When does the system specifically fall back to trigram search?
The fallback triggers when the Porter stemmed search returns no results or insufficient matches for queries containing camelCase, technical prefixes, or partial words. For example, a query for "responseBodyParser" will fail the Porter layer because "response" stems separately from the camelCase suffix, but succeed in the trigram layer which matches the exact character sequences "res", "esp", "spo", etc.
Can I query only the trigram index for specific use cases?
Yes. While searchWithFallback automatically handles the fusion, you can call searchTrigram directly (lines 505–517) to force a substring-only search. This is useful for debugging, exact pattern matching, or when you specifically need to find code fragments where you know the exact character sequence but not the full tokenization.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →