QMD vs ripgrep: Semantic Search vs Lexical Pattern Matching
QMD is a semantic-aware search engine that uses SQLite FTS5 and vector embeddings for ranked, context-aware results, while ripgrep performs fast, stateless regex scans without indexing or ranking.
When comparing local search tools, developers often pit QMD (Query-Markup-Documents) against ripgrep. While both run from the command line, they serve fundamentally different purposes. According to the tobi/qmd source code, QMD builds persistent SQLite indexes with full-text and vector capabilities, whereas ripgrep scans files on-the-fly using optimized regex.
Architecture: Indexed vs On-the-Fly Search
QMD's SQLite and FTS5 Foundation
In src/store.ts, QMD implements searchFTS (lines 1999-2008) using SQLite's FTS5 extension for full-text search. The system also maintains vector embeddings via searchVec (lines 5550-5580), leveraging the sqlite-vec extension. This dual-index approach requires an initial indexing phase via qmd collection add and qmd embed, but enables sub-millisecond queries thereafter.
ripgrep's Regex-First Approach
ripgrep does not maintain a persistent index. It traverses the file system at runtime, applying optimized regex patterns directly to file contents. This stateless approach eliminates setup overhead but requires scanning files repeatedly for each query.
Search Capabilities: Semantic vs Lexical
Vector Embeddings and BM25 Ranking in QMD
QMD combines lexical and semantic search through hybridQuery in src/store.ts (lines 2806-2848). The implementation uses:
- BM25 scoring from FTS5, normalized as
|bm25|/(1+|bm25|)insearchFTS - Cosine similarity for vector comparisons in
searchVec - Reciprocal Rank Fusion (RRF) to combine lexical and semantic results
The system uses LLM-based embeddings defined by DEFAULT_EMBED_MODEL_URI and optional reranking via DEFAULT_RERANK_MODEL_URI in src/llm.ts.
Pattern Matching in ripgrep
ripgrep performs pure lexical matching using regex. It has no notion of semantic similarity, relevance ranking, or document embeddings. Results appear in file traversal order rather than by relevance.
Performance and Workflow Trade-offs
QMD requires upfront investment: running qmd collection add . --name notes followed by qmd embed to compute vector embeddings. However, subsequent searches via qmd search, qmd vsearch, or qmd query execute as sub-millisecond SQLite lookups.
ripgrep offers immediate execution with rg <pattern> and excels for ad-hoc searches across small to medium codebases where indexing overhead would outweigh benefits.
When to Choose QMD vs ripgrep
Choose QMD when you need:
- Semantic understanding (e.g., "find design docs about OAuth" matching conceptual rather than literal text)
- Ranked results using BM25 and vector similarity
- Persistent collections with virtual paths (
qmd://collection/path.mdas implemented inparseVirtualPathinsrc/store.tslines 1313-1327) - Context-aware output using folder-level contexts from YAML collection configs in
src/collections.ts
Choose ripgrep when you need:
- Fast, stateless regex searches without setup
- Minimal installation footprint (single ~2MB binary vs. Bun + SQLite + model files ~200MB)
- Searches across constantly changing file trees where index maintenance would be costly
Summary
- QMD builds persistent SQLite indexes with FTS5 and vector embeddings, enabling semantic search and BM25 ranking.
- ripgrep performs stateless regex scans without indexing, offering immediate results for literal pattern matching.
- QMD excels at conceptual queries and ranked retrieval via
hybridQueryandsearchVec. - ripgrep excels at fast, ad-hoc searches with minimal overhead.
- Choose QMD for document collections requiring semantic understanding; choose ripgrep for codebase grepping and regex workflows.
Frequently Asked Questions
Can QMD replace ripgrep entirely?
No. While QMD offers superior semantic search capabilities, it requires upfront indexing and heavier dependencies (Bun, SQLite with extensions, and LLM models). ripgrep remains the optimal choice for quick, stateless regex searches across codebases where semantic understanding is unnecessary.
How does QMD's hybrid search work?
QMD's hybridQuery function in src/store.ts (lines 2806-2848) combines BM25 scores from the FTS5 full-text index with cosine similarity scores from vector embeddings. It uses Reciprocal Rank Fusion (RRF) to merge these signals, producing a unified relevance ranking that captures both literal term matches and conceptual similarity.
What are the storage requirements for QMD compared to ripgrep?
ripgrep has virtually zero storage overhead beyond the ~2MB binary. QMD requires SQLite database files storing inverted indices and vector embeddings, plus optional LLM model files (~200MB). The qmd collection add and qmd embed commands populate these persistent stores, trading disk space for query performance.
Does ripgrep support semantic or vector search?
No. ripgrep performs pure lexical matching using optimized regex engines. It has no functionality for computing text embeddings, measuring semantic similarity, or ranking results by conceptual relevance. For semantic capabilities, tools like QMD that integrate LLM-based embeddings and vector databases are required.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →