Migrating from Another Search System to QMD: A Complete Step-by-Step Guide
Migrate any existing search infrastructure to QMD by installing the binary, defining YAML collections that point to your current document directories, running qmd update to populate the SQLite index, and optionally generating embeddings with qmd embed for semantic search capabilities.
Migrating from another search system to QMD moves your searchable corpus into a self‑contained, on‑device engine that requires no external services or credentials. Whether you are transitioning from Elasticsearch, MeiliSearch, or a custom SQLite schema, the process centers on mapping your existing directories to QMD collections and rebuilding the index locally.
Install QMD and Verify Prerequisites
Begin by installing the QMD binary from the tobi/qmd repository. The package bundles compiled TypeScript in dist/ and downloads required GGUF models on first run.
# Using npm (Node ≥22) or Bun (≥1.0)
npm install -g @tobilu/qmd
# or
bun install -g @tobilu/qmd
Verify the installation by checking the version and ensuring the default cache directory exists at ~/.cache/qmd/.
Define Collections Mapping to Existing Directories
QMD stores collection definitions in a human‑readable YAML file located at ~/.config/qmd/index.yml. Each collection maps a name to an absolute path and an optional glob pattern.
# ~/.config/qmd/index.yml
collections:
wiki:
path: /Users/alice/wiki
pattern: "**/*.md"
tickets:
path: /Users/alice/support/tickets
pattern: "**/*.txt"
global_context: "My personal knowledge base"
Alternatively, use the CLI to generate the same configuration:
qmd collection add /Users/alice/wiki --name wiki
qmd collection add /Users/alice/support/tickets --name tickets --mask "**/*.txt"
qmd context add / "My personal knowledge base"
The addCollection and setGlobalContext functions in src/collections.ts manage this YAML file, ensuring that your existing directory structure is preserved without moving files.
Index Documents into SQLite
Once collections are defined, run the indexing command. QMD walks the filesystem, extracts titles (first H1 or filename), hashes content, and populates three core tables in ~/.cache/qmd/index.sqlite:
content– Content‑addressable storage mappinghash → doc.documents– Mapping of collection name + relative path → metadata (title, hash, timestamps).documents_fts– SQLite FTS5 virtual table for BM25 full‑text search.
qmd update # Walk all collections and (re)index files
qmd status # Verify document counts and embedding gaps
During indexing, the initializeDatabase function in src/store.ts creates the schema and triggers, including documents_ai, which automatically updates the FTS index when new rows are inserted into documents.
Generate Vector Embeddings for Semantic Search
To enable semantic search alongside BM25, generate embeddings for all indexed chunks. This step is optional but required if you want vector similarity capabilities.
qmd embed # Chunk documents → 900‑token chunks → embed with embeddinggemma
The embedding pipeline in src/store.ts performs the following:
- Chunking –
chunkDocumentrespects markdown headings, code fences, and a 200‑token window (CHUNK_WINDOW_TOKENS). - Formatting –
formatDocForEmbeddingrenders each chunk as"title: … | text: …". - Embedding –
node‑llama‑cpploads the default GGUF modelembeddinggemma-300M-Q8_0. - Storage – Vectors are stored in
content_vectorsand indexed byvectors_vecusingsqlite‑vec.
Verify the Migration
Validate that your data migrated correctly by running search queries and health checks:
qmd search "my old Elasticsearch query" -c wiki -n 5
qmd query "semantic search test"
qmd status
The getIndexHealth and getStatus functions in src/store.ts provide the statistics displayed by these commands, ensuring document counts match your source system and no embedding gaps remain.
Handle Legacy Schema Upgrades
If you are upgrading from a QMD version < 0.5 that stored collections in a separate SQLite table (collections) and used a foreign‑key collection_id in documents, you must run the legacy migration script to convert to the new name‑based schema.
./migrate-schema.ts # Run with Bun: bun migrate-schema.ts
The script performs these steps:
- Adds a
collection TEXTcolumn todocuments. - Populates it by joining the old
collectionstable (mapping IDs to names). - Re‑creates
documentswithout the foreign key, copies data, and drops the old table. - Re‑installs FTS triggers to use the new
collectioncolumn.
After migration, run qmd update once more to rebuild the FTS index. The complete logic resides in migrate-schema.ts at the repository root.
Programmatic Migration for Non-Filesystem Sources
For documents stored in databases or cloud buckets rather than local files, use the store API directly to import data:
import { createStore } from "./src/store.js";
async function importFromExternalSystem(docs: Array<{content: string, path: string, collection: string}>) {
const store = createStore(); // Uses default DB path ~/.cache/qmd/index.sqlite
for (const {content, path, collection} of docs) {
const hash = await store.hashContent(content);
const title = store.extractTitle(content, path);
const now = new Date().toISOString();
store.insertContent(hash, content, now);
store.insertDocument(collection, path, title, hash, now, now);
}
console.log("✅ Migration complete – run `qmd embed` to enable semantic search");
}
Key functions in src/store.ts include hashContent, extractTitle, insertContent, and insertDocument, which handle the content‑addressable storage and document metadata mapping.
Integrate with AI Agents via MCP
If your previous search system exposed an HTTP endpoint, replace it with QMD’s MCP server, which provides a JSON‑RPC interface compatible with Claude, GPT‑4, and other agents:
# Start a long‑lived HTTP server (default port 8181)
qmd mcp --http &
The MCP server exposes these tools: qmd_search, qmd_vector_search, qmd_deep_search, qmd_get, qmd_multi_get, and qmd_status. Implementation resides in src/mcp.ts, with response formatting handled by src/formatter.ts.
Post-Migration Maintenance
Keep the index healthy after migration:
qmd cleanup # Delete LLM cache, orphaned content/vectors, vacuum DB
qmd status # Verify no missing embeddings
The cleanupOrphanedContent and vacuumDatabase functions in src/store.ts handle database optimization and removal of stale data.
Summary
- Install QMD via npm or Bun to get the binary and bundled TypeScript core.
- Define collections in
~/.config/qmd/index.ymlto map existing directories without moving files. - Run
qmd updateto populate the SQLite tables (documents,content,documents_fts) with BM25‑ready data. - Run
qmd embedto generate 900‑token chunks and store vectors incontent_vectorsfor semantic search. - Verify with
qmd statusand search commands to ensure document counts match your source system. - Upgrade legacy installs (< 0.5) using
migrate-schema.tsto convert foreign‑key schemas to name‑based collections. - Maintain with
qmd cleanupto vacuum the database and remove orphaned content.
Frequently Asked Questions
How do I migrate documents that are not stored on the local filesystem?
Use the programmatic store API in src/store.ts. Import the createStore function, then iterate over your external data source and call hashContent, extractTitle, insertContent, and insertDocument for each document. This bypasses the filesystem walker and imports directly into the SQLite schema.
What is the difference between qmd update and qmd embed?
qmd update performs the full‑text indexing phase: it walks collections, hashes content, populates the documents and content tables, and updates the FTS5 virtual table documents_fts for BM25 search. qmd embed is the semantic indexing phase: it chunks documents into 900‑token segments, generates vectors using embeddinggemma, and stores them in content_vectors with an index in vectors_vec.
Do I need to run the legacy schema migration?
Only if you are upgrading from QMD version < 0.5. Earlier versions stored collections in a separate SQLite table and used integer foreign keys (collection_id) in the documents table. Version 0.5+ uses a YAML‑based configuration and stores the collection name directly as a text column. Run bun migrate-schema.ts to convert old databases to the new format.
How can I verify that my migration succeeded?
Run qmd status to check that document counts match your source system and that no "needs embedding" gaps remain. Then execute test searches using qmd search for BM25 queries and qmd query for semantic vector search. The getIndexHealth function in src/store.ts provides the underlying statistics for these checks.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →