How Context-Mode Handles Fuzzy Correction for Misspelled Search Queries
Context-mode corrects misspelled query terms using a Levenshtein distance algorithm with dynamic edit-distance ceilings, SQLite-backed vocabulary lookups, and an in-memory LRU cache to optimize repeated corrections.
The mksglu/context-mode repository implements an intelligent typo-tolerance system that automatically fixes misspelled search queries before executing the search. Understanding how this fuzzy correction for misspelled search queries works reveals a sophisticated multi-stage pipeline that balances accuracy with performance through dynamic distance calculations and strategic caching.
The Three-Stage Fuzzy Correction Pipeline
The core correction logic resides in src/store.ts and operates through three distinct computational stages.
Stage 1: Dynamic Edit-Distance Ceiling Calculation
Before comparing strings, the system calculates a maximum allowable edit distance based on query word length. The maxEditDistance(wordLength) function at src/store.ts:L139-L143 returns 1, 2, or 3 depending on the input length, ensuring short words receive tighter tolerance limits while longer words allow more flexibility for typos.
Stage 2: SQLite Vocabulary Candidate Lookup
Once the ceiling is established, the system queries the SQLite database using the prepared statement #stmtFuzzyVocab. This operation, defined at src/store.ts:L333-L337, retrieves all vocabulary entries whose length falls within word.length ± maxDist, narrowing the candidate pool to plausible matches before expensive distance calculations begin.
Stage 3: Levenshtein Distance Scoring and Selection
For each candidate retrieved, the classic Levenshtein distance algorithm computes the exact edit distance between the misspelled term and the candidate vocabulary word. The implementation at src/store.ts:L902-L946 selects the candidate with the smallest distance that remains within the calculated ceiling. If an exact match exists, the function returns null to indicate no correction is necessary.
Caching Strategy for Performance Optimization
To prevent redundant database queries and repeated distance calculations, context-mode employs an LRU (Least Recently Used) cache strategy.
In-Process LRU Cache Implementation
The private field #fuzzyCache stores correction results as string | null mappings between misspelled inputs and their corrections. When fuzzyCorrect() executes, it first checks this cache at src/store.ts:L906-L915. Cache hits are promoted to the tail of the LRU structure, keeping frequently accessed corrections in fast memory. Once the cache exceeds ContentStore.FUZZY_CACHE_SIZE, the least recently used entries are evicted automatically at src/store.ts:L939-L945.
End-to-End Query Processing and Fallback Strategy
Individual word corrections integrate into the broader search architecture through the searchWithFallback() method.
Multi-Word Query Correction
When processing a search query, searchWithFallback() at src/store.ts:L653-L670 splits the input into individual tokens. It skips stop-words and words shorter than three characters, then invokes fuzzyCorrect() on each remaining term. The system rebuilds the query string with corrected terms, and if any modifications occurred, re-executes the search through the RRF-fusion pipeline.
Result Layer Tagging
Corrected queries are tagged with matchLayer: "rrf-fuzzy" to distinguish them from direct hits. This metadata allows downstream consumers to identify when results originate from fuzzy-corrected searches rather than exact matches, providing transparency into the search process.
Implementation Examples
The following examples demonstrate the fuzzy correction API in practice.
Basic Fuzzy-Corrected Search
import { ContentStore } from "./src/store.js";
const store = new ContentStore();
// Index content containing the correct term
store.index({ content: "# Authentication\n\nThe system uses **authentication** tokens.", source: "doc:auth" });
// Search with a misspelled query
const results = store.searchWithFallback("autentication", 5);
console.log(results[0].matchLayer); // → "rrf-fuzzy"
console.log(results[0].title); // → "Authentication"
This example triggers the complete correction pipeline described in src/store.ts:L653-L670.
Direct Correction Inspection
const correction = store.fuzzyCorrect("orchestraton"); // typo for "orchestration"
console.log(correction); // → "orchestration"
The fuzzyCorrect method at src/store.ts:L902-L946 returns the corrected string or null if no suitable candidate exists within the maximum edit distance.
Cache Behavior Verification
// First lookup populates the cache
store.fuzzyCorrect("authentiction");
// Second lookup hits cache, avoiding database work
store.fuzzyCorrect("authentiction"); // Retrieved from #fuzzyCache
The LRU eviction policy activates when the cache size exceeds ContentStore.FUZZY_CACHE_SIZE, as implemented at src/store.ts:L939-L945.
Summary
- Context-mode calculates dynamic edit-distance ceilings via
maxEditDistance(), allowing tighter tolerance for short words and more flexibility for long words. - The system queries SQLite vocabulary using
#stmtFuzzyVocaband filters candidates by length proximity before computing exact Levenshtein distances. - An LRU cache (
#fuzzyCache) eliminates redundant correction work for repeated misspellings, with automatic eviction when size limits are exceeded. - The
searchWithFallback()method orchestrates multi-word corrections, skipping stop-words and short tokens while tagging results withmatchLayer: "rrf-fuzzy"for transparency.
Frequently Asked Questions
What algorithm does context-mode use for fuzzy correction?
Context-mode uses the classic Levenshtein distance algorithm to calculate the edit distance between misspelled query terms and vocabulary candidates. The implementation dynamically adjusts the maximum allowed distance based on word length, ensuring short words like "cat" don't match distant alternatives while longer technical terms have more tolerance for typos.
How does the fuzzy correction cache work?
The system maintains an in-process LRU (Least Recently Used) cache via the private #fuzzyCache field. When fuzzyCorrect() encounters a previously seen misspelling, it returns the cached result immediately without querying the SQLite database. Frequently accessed corrections migrate to the tail of the cache structure, while entries exceeding ContentStore.FUZZY_CACHE_SIZE are evicted from the front.
Which words get skipped during fuzzy correction?
The searchWithFallback() method intentionally skips two categories of tokens: standard stop-words (common words like "the" or "and") and any word shorter than three characters. This optimization prevents the system from attempting to correct short functional words where corrections would likely introduce false positives rather than helpful fixes.
How can I identify if search results came from a corrected query?
Results originating from fuzzy-corrected queries carry the metadata matchLayer: "rrf-fuzzy" rather than standard layer identifiers. This tagging occurs in searchWithFallback() at src/store.ts:L653-L670, allowing your application to indicate to users when results match a corrected version of their original input.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →