How Wigolo's Local Cache Handles Keyword and Semantic Search

Wigolo's cache system routes queries through src/tools/cache.ts to either SQLite FTS5 for pure keyword search or a hybrid pipeline that fuses keyword and vector rankings using Reciprocal Rank Fusion (RRF), automatically falling back to keyword-only mode when embedding services are unavailable.

The KnockOutEZ/wigolo repository implements a robust local caching mechanism that enables fast content retrieval through both lexical and semantic matching. Understanding how wigolo's local cache function with keyword and semantic search works requires examining its dual-path architecture that combines traditional full-text search with modern vector similarity, all backed by SQLite storage and intelligent rank fusion algorithms.

Request Dispatch and Cache Entry Point

The handleCache function in src/tools/cache.ts serves as the primary router that interprets incoming requests based on operation flags. It distinguishes between cache maintenance tasks and search operations:

  • check_changes – Runs a change-detector on cached rows to identify stale content
  • stats – Returns cache metadata via getCacheStats()
  • clear – Deletes rows matching supplied time-based or URL filters
  • mode: "hybrid" – Activates the combined keyword and semantic search pipeline
  • Default (no mode) – Executes plain keyword search via FTS5

This dispatch logic ensures that maintenance operations bypass search infrastructure entirely while routing queries to the appropriate retrieval engine.

Keyword Search with SQLite FTS5

For standard keyword queries, the system invokes ftsSearchRanked(query, limit) implemented in src/cache/store.ts. This function executes SQLite FTS5 (Full-Text Search version 5) queries against the local cache database, scoring rows by term frequency and BM25 relevance to return an ordered list of URLs.

The keyword path operates entirely within the SQLite layer without invoking embedding models or vector stores. This design guarantees minimal latency and zero dependency on external AI services, making it suitable for environments without vector infrastructure.

Automatic Fallback Behavior

If the vector store is empty, the embed provider fails initialization, or the query parameter is omitted, handleCache automatically degrades to this pure FTS5 path. This fallback mechanism ensures that cache queries always return usable results regardless of semantic service availability.

Semantic Vector Search Pipeline

When mode: "hybrid" is specified, the system activates the semantic search layer through a multi-step process defined in src/tools/cache.ts:

  1. Query Embedding – Retrieves an embed provider via getEmbedProvider() and converts the query text into a dense vector using embedProvider.embed([query])
  2. Vector Retrieval – Accesses the vector store through getVectorStore() and executes store.search(queryVector, candidateLimit) to find nearest neighbors based on vector similarity
  3. Candidate Selection – Returns the top candidates up to the specified limit for downstream fusion

This pipeline depends on abstraction layers in src/providers/embed-provider.ts for text vectorization and src/providers/vector-store.ts for persistent vector storage, allowing pluggable backend implementations.

Hybrid Fusion with Reciprocal Rank Fusion (RRF)

The hybrid mode merges results from both the FTS5 keyword search and the vector semantic search using the Reciprocal Rank Fusion algorithm implemented in src/search/rrf.ts.

Rank Combination Logic

The system constructs rank maps via buildRankMap() for both the keyword result list and the vector result list, then applies reciprocalRankFusion() with a constant k = 60. Each document receives a fusion score calculated as:


score = Σ 1 / (60 + rank_i)

Where rank_i is the document's position in each source list. The sortByRRFScore() function then reorders the final output. This mathematical approach favors documents that appear highly ranked in both lexical and semantic contexts, typically surfacing more relevant content than either method in isolation.

Result Hydration and Token Budget Enforcement

After rank fusion, the system performs two final processing stages before returning results:

Content Hydration – For each URL in the fused ranking (up to the requested limit), the system calls getCachedContentByNormalizedUrl() to retrieve the full cached row from src/cache/store.ts. Missing or expired rows are silently skipped without breaking the result set.

Token Trimming – Before returning, applyAggregateMarkdownBudget() in src/search/evidence.ts processes the aggregated markdown bodies to enforce the max_tokens_out parameter. This operation truncates content to fit within specified token limits while preserving the complete list of result metadata (URLs, titles, and timestamps).

The final output conforms to the CacheResultItem type, returning an array of {url, title, markdown, fetched_at} objects, or an {error} object if unrecoverable exceptions occur.

Practical Implementation Examples

// Pure keyword search - retrieve 5 latest matches for "typescript"
const kwResult = await handleCache({
  query: 'typescript',
  limit: 5,
});
// Returns: { results: [{url, title, markdown, fetched_at}, ...] }
// Hybrid semantic + keyword search with token budget constraints
const hybridResult = await handleCache({
  query: 'async event handling',
  mode: 'hybrid',
  limit: 8,
  max_tokens_out: 400,
});
// Returns: fused, re-ranked results combining FTS5 and vector similarity
// Cache maintenance - clear entries older than 30 days
await handleCache({
  clear: true,
  since: new Date(Date.now() - 30 * 24 * 60 * 60 * 1000),
});

Summary

  • Triple-mode architecture – Wigolo supports pure keyword (FTS5), pure semantic (vector), and hybrid search modes through src/tools/cache.ts
  • SQLite FTS5 implementation – Keyword search uses ftsSearchRanked() in src/cache/store.ts without external dependencies
  • Vector pipeline – Semantic search relies on embedProvider.embed() and vectorStore.search() from the providers abstraction layer
  • RRF fusion – Hybrid mode merges rankings using Reciprocal Rank Fusion with k = 60 as defined in src/search/rrf.ts
  • Graceful degradation – Automatic fallback to keyword search when vector services fail or return empty results
  • Token management – applyAggregateMarkdownBudget() in src/search/evidence.ts enforces max_tokens_out limits on returned content bodies

Frequently Asked Questions

What happens if the vector embedding service is unavailable?

If the embed provider fails initialization or the vector store contains no entries, handleCache in src/tools/cache.ts automatically routes the query to the pure SQLite FTS5 path via ftsSearchRanked(). This fallback behavior ensures the cache remains fully operational for keyword retrieval even when semantic infrastructure is offline.

How does Reciprocal Rank Fusion improve search relevance?

RRF, implemented in src/search/rrf.ts, combines rankings from both keyword and semantic engines by calculating 1 / (60 + rank) for each document in each list and summing the scores. Documents appearing highly in both lexical and semantic rankings receive disproportionately higher fusion scores, typically yielding more relevant results than either search method alone.

What is the purpose of the max_tokens_out parameter?

The max_tokens_out parameter triggers applyAggregateMarkdownBudget() in src/search/evidence.ts, which trims the markdown content of cached pages to fit within the specified token budget. This prevents context window overflow in downstream LLM processing while preserving the full list of result URLs, titles, and metadata.

Yes. The handleCache function supports maintenance operations including stats (returns cache statistics via getCacheStats()), clear (deletes filtered rows by time or URL), and check_changes (runs change detection on cached content). These operations execute independently of both keyword and semantic search logic.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →