# How the FTS5 Knowledge Base in context-mode Combines Porter Stemming and Trigram Search

> Discover how FTS5 in context-mode merges Porter stemming and trigram search with Reciprocal Rank Fusion for comprehensive recall on exact and stemmed queries within mksglu/context-mode.

- Repository: [Mert Köseoğlu/context-mode](https://github.com/mksglu/context-mode)
- Tags: internals
- Published: 2026-04-24

---

**The context-mode FTS5 knowledge base stores every text chunk in two separate virtual tables—one using Porter stemming for morphological matching and another using trigrams for substring search—then merges results with Reciprocal Rank Fusion (RRF) to ensure high recall across exact and stemmed queries.**

The `mksglu/context-mode` repository implements a sophisticated full-text search system using SQLite FTS5. Its knowledge base architecture solves the classic search trade-off between stemming accuracy and substring matching by maintaining dual indexes and executing a three-layer fallback pipeline that automatically selects the best results from either approach.

## Dual-Index Architecture in src/store.ts

The system creates two distinct FTS5 virtual tables during initialization in [`src/store.ts`](https://github.com/mksglu/context-mode/blob/main/src/store.ts) (lines 1007–1021).

### The chunks Table: Porter Stemming

The primary index uses the `porter unicode61` tokenizer. This implementation applies **Porter stemming** to normalize words to their root forms, ensuring that queries like "running" match documents containing "run" or "runs".

### The chunks_trigram Table: Substring Matching

A secondary index uses the `trigram` tokenizer. This table indexes every three-character sequence in the text, enabling **substring search** capabilities that Porter stemming cannot provide, such as matching camelCase identifiers ("responseBody") or partial word segments.

Both tables are populated simultaneously during the indexing phase via the `#insertChunks` method (lines 783–787), ensuring every chunk exists in both representations without duplication of the source text.

## The Three-Layer Search Pipeline

When `searchWithFallback` (lines 639–674) executes, it orchestrates queries across both indexes and fuses the results.

### Layer 1: Porter Stemmed Search

The pipeline first calls `search` (lines 462–475), which sanitizes the query and executes it against the `chunks` table. This layer excels at linguistic matches but fails on camelCase substrings or exact partial matches where stemming breaks the token.

### Layer 2: Trigram Fallback Search

If Layer 1 returns insufficient results, the system invokes `searchTrigram` (lines 505–517) against the `chunks_trigram` table. This captures matches that the Porter tokenizer misses, such as code identifiers or technical terms where the user remembers only a substring.

### Layer 3: RRF Fusion and Proximity Reranking

The `#rrfSearch` method (lines 549–590) implements **Reciprocal Rank Fusion (RRF)** to merge the two result sets. It assigns scores based on reciprocal rank positions, then applies a proximity model that boosts matches in titles and tightly clustered term occurrences. The final list represents the optimal combination of stemmed linguistic matches and exact substring hits.

## Practical Code Examples

### Indexing Content with Dual Tokenizers

```typescript
import { ContentStore } from "./src/store.js";

const store = new ContentStore();
store.index({ path: "docs/guide.md", source: "guide" });

```

*The `index` method delegates to `#insertChunks`, which inserts each chunk into both `chunks` and `chunks_trigram` tables.*

### Executing a Fallback Search

```typescript
// Matches "responseBody" even if Porter stems "response" separately
const results = store.searchWithFallback("responseBody", 5);
console.log(results.map(r => `${r.title}: ${r.highlighted}`));

```

*This executes `#rrfSearch`, combining Porter and trigram results through reciprocal rank fusion.*

### Debugging Individual Layers

```typescript
// Porter-only results (linguistic matching)
const porterHits = store.search("caching", 5);

// Trigram-only results (substring matching)
const trigramHits = store.searchTrigram("responseBody", 5);

```

*Direct access to `search` (lines 462–475) and `searchTrigram` (lines 505–517) enables performance analysis and debugging.*

## Summary

- **Dual indexing strategy**: Every chunk is stored in both `chunks` (Porter) and `chunks_trigram` (trigram) tables during indexing.
- **Automatic fallback**: The `searchWithFallback` method attempts Porter matching first, then supplements with trigram results when linguistic stemming fails.
- **RRF fusion**: Results from both layers are merged using Reciprocal Rank Fusion in `#rrfSearch`, then reranked by proximity and title relevance.
- **CamelCase support**: The trigram layer ensures code identifiers and technical terms match even when Porter stemming breaks apart compound words.

## Frequently Asked Questions

### Why does context-mode use both Porter and trigram tokenizers?

Porter stemming excels at matching word variations (searching for "run" finds "running" and "runs"), but it fails on camelCase identifiers like "responseBody" where stemming breaks the token. The trigram tokenizer indexes every three-character sequence, enabling substring matches that capture these technical terms. Using both ensures high recall across natural language and code-focused content.

### How does Reciprocal Rank Fusion (RRF) combine results from both indexes?

RRF assigns a score to each document based on its rank in each result list using the formula `1 / (k + rank)`, where `k` is a constant (typically 60). Documents appearing in both the Porter and trigram results receive higher combined scores than those appearing in only one list. The `#rrfSearch` method implements this fusion, then applies additional proximity-based reranking to prioritize title matches and tightly clustered terms.

### When does the system specifically fall back to trigram search?

The fallback triggers when the Porter stemmed search returns no results or insufficient matches for queries containing camelCase, technical prefixes, or partial words. For example, a query for "responseBodyParser" will fail the Porter layer because "response" stems separately from the camelCase suffix, but succeed in the trigram layer which matches the exact character sequences "res", "esp", "spo", etc.

### Can I query only the trigram index for specific use cases?

Yes. While `searchWithFallback` automatically handles the fusion, you can call `searchTrigram` directly (lines 505–517) to force a substring-only search. This is useful for debugging, exact pattern matching, or when you specifically need to find code fragments where you know the exact character sequence but not the full tokenization.