How Claude-Mem Combines SQLite FTS5 and Chroma Vector Embeddings for Hybrid Search

Claude-Mem implements a strategy-orchestrator pattern that routes queries between SQLite FTS5 for deterministic metadata filtering and Chroma vector embeddings for semantic ranking, using a four-step hybrid algorithm that intersects both result sets to deliver precise yet contextually relevant results.

Claude-Mem is an open-source memory layer that unifies structured metadata storage with semantic search capabilities. By combining SQLite FTS5 full-text search with Chroma vector embeddings, the system achieves hybrid search that respects strict metadata constraints while preserving semantic relevance. This architecture is implemented across several key services in the thedotmack/claude-mem repository.

The Three Search Strategies

Claude-Mem’s SearchOrchestrator coordinates three distinct search backends, selecting the appropriate strategy based on query parameters and service availability:

Strategy Implementation Use Case
SQLite FTS5 SQLiteSearchStrategy Filter-only queries with no free-text query parameter, or as fallback when Chroma is unavailable. Uses virtual tables (observations_fts, session_summaries_fts) for fast full-text matching and relevance ranking.
Chroma ChromaSearchStrategy Pure semantic search when a free-text query is present and the Chroma server is running. Performs vector-based similarity search on document embeddings.
Hybrid HybridSearchStrategy Metadata-constrained semantic search for find_by_concept, find_by_type, and find_by_file operations. Combines SQLite filtering with Chroma ranking.

How SearchOrchestrator Routes Queries

The routing logic resides in src/services/worker/search/SearchOrchestrator.ts. The executeWithFallback method implements the decision tree:

When no free-text query is provided, the orchestrator immediately delegates to the SQLite strategy for filter-only retrieval【/tmp/instagit_r1if4bdc/src/services/worker/search/SearchOrchestrator.ts#L84-L89】.

When a query is present and Chroma is available, it attempts the Chroma strategy first; if that fails, it falls back to SQLite【/tmp/instagit_r1if4bdc/src/services/worker/search/SearchOrchestrator.ts#L90-L107】.

For specialized lookups, the orchestrator explicitly invokes the Hybrid strategy. For example, findByConcept routes through HybridSearchStrategy【/tmp/instagit_r1if4bdc/src/services/worker/search/SearchOrchestrator.ts#L24-L31】.

The Hybrid Search Algorithm

The HybridSearchStrategy in src/services/worker/search/strategies/HybridSearchStrategy.ts implements a four-step pipeline that combines SQLite FTS5 metadata filtering with Chroma vector ranking【/tmp/instagit_r1if4bdc/src/services/worker/search/strategies/HybridSearchStrategy.ts#L1-L8】.

Step 1: Metadata Filtering with SQLite FTS5

The strategy first applies strict metadata constraints using SQLite. It retrieves IDs that satisfy the requested concepts, types, files, project, or date ranges:

const metadataResults = this.sessionSearch.findByConcept(concept, filterOptions);

The same pattern applies for findByType and findByFile【/tmp/instagit_r1if4bdc/src/services/worker/search/strategies/HybridSearchStrategy.ts#L74-L81】. This step leverages the FTS5 virtual tables for fast text relevance when filtering.

Step 2: Semantic Ranking with Chroma

Next, the strategy submits the candidate IDs to Chroma for semantic ranking against the original query:

const chromaResults = await this.chromaSync.queryChroma(
    concept,
    Math.min(ids.length, SEARCH_CONSTANTS.CHROMA_BATCH_SIZE)
);

This retrieves vector-based similarity scores for the filtered subset【/tmp/instagit_r1if4bdc/src/services/worker/search/strategies/HybridSearchStrategy.ts#L84-L90】.

Step 3: Intersection and Reordering

The strategy intersects the SQLite metadata results with the Chroma semantic ranking, preserving the Chroma order:

const rankedIds = this.intersectWithRanking(ids, chromaResults.ids);

The intersectWithRanking helper walks the Chroma-ordered list and filters against a Set of SQLite IDs【/tmp/instagit_r1if4bdc/src/services/worker/search/strategies/HybridSearchStrategy.ts#L58-L66】.

Step 4: Result Hydration

Finally, the strategy fetches full observation or session rows from SQLite using the ranked IDs, then reorders them to match the semantic ranking:

const observations = this.sessionStore.getObservationsByIds(rankedIds, { limit });
observations.sort((a, b) => rankedIds.indexOf(a.id) - rankedIds.indexOf(b.id));

If any step fails, the strategy falls back to pure SQLite metadata results【/tmp/instagit_r1if4bdc/src/services/worker/search/strategies/HybridSearchStrategy.ts#L13-L23】.

SQLite FTS5 Implementation Details

The FTS5 virtual tables are created during database migration in src/services/sqlite/migrations.ts. The migration creates observations_fts and session_summaries_fts virtual tables that index the content for full-text search【/tmp/instagit_r1if4bdc/src/services/sqlite/migrations.ts#L379-L420】.

When executing filter-only queries, SessionSearch in src/services/sqlite/SessionSearch.ts can order results by FTS5 relevance. The buildOrderClause method checks for hasFTS === true and appends an ORDER BY …rank clause to leverage the virtual table's built-in relevance scoring【/tmp/instagit_r1if4bdc/src/services/sqlite/SessionSearch.ts#L27-L31】.

Code Examples

import { SearchOrchestrator } from './services/worker/search/SearchOrchestrator.js';

const orchestrator = new SearchOrchestrator(sessionSearch, sessionStore, chromaSync);

const result = await orchestrator.findByConcept('authentication', {
  project: 'my-app',
  dateRange: { start: '2024-01-01', end: '2024-12-31' },
  limit: 20,
  orderBy: 'date_desc'
});

console.log(result.results.observations);

This executes the four-step hybrid pipeline: SQLite metadata filtering, Chroma semantic ranking, intersection, and hydration.

Filter-Only Search Using SQLite FTS5

const result = await orchestrator.search({
  obs_type: 'decision',
  concepts: 'security',
  project: 'my-app',
  orderBy: 'relevance'
});

With no free-text query parameter, the orchestrator routes to SQLiteSearchStrategy, which uses the FTS5 virtual table's rank column for relevance ordering.

const result = await orchestrator.search({
  query: 'how to encrypt data in transit',
  limit: 10
});

When a query is present and Chroma is available, ChromaSearchStrategy performs pure vector similarity search without metadata constraints.

Summary

  • Claude-Mem implements a strategy-orchestrator pattern that dynamically selects between SQLite FTS5, Chroma vector search, or a hybrid combination based on query parameters and service availability.
  • Hybrid search executes a four-step pipeline: SQLite metadata filtering, Chroma semantic ranking, intersection with preserved ranking, and SQLite hydration.
  • SQLite FTS5 provides fast full-text matching and relevance scoring through virtual tables (observations_fts, session_summaries_fts) created in migrations.ts.
  • Fallback mechanisms ensure resilience: if Chroma fails, the system falls back to SQLite; if hybrid steps fail, it returns pure metadata results.

Frequently Asked Questions

How does Claude-Mem decide which search strategy to use?

The SearchOrchestrator in src/services/worker/search/SearchOrchestrator.ts examines the query parameters. If no free-text query is provided, it uses the SQLite strategy. If a query exists and Chroma is available, it attempts the Chroma strategy first, falling back to SQLite if Chroma fails. For specialized lookups like findByConcept, it explicitly invokes the Hybrid strategy.

What happens if the Chroma vector database is unavailable?

When Chroma is unreachable or returns an error, the SearchOrchestrator catches the failure and falls back to the SQLiteSearchStrategy. This ensures that metadata filtering and FTS5 full-text search remain available even when the vector service is down, maintaining system resilience.

Currently, the ChromaSearchStrategy does not apply metadata constraints when running pure vector search. To combine metadata filtering with semantic relevance, use the hybrid approach via findByConcept, findByType, or findByFile methods, which first filter by metadata in SQLite then rank by semantic similarity in Chroma.

Relevance in hybrid search comes from Chroma's vector similarity scores. The SQLite step only filters by metadata constraints and retrieves candidate IDs. Chroma then ranks these IDs by semantic similarity to the query. The final results preserve Chroma's ranking order but only include IDs that passed the SQLite metadata filters, combining precise filtering with semantic relevance.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →