# qmd | Tobias Lütke | Knowledge Base | Instagit

mini cli search engine for your docs, knowledge bases, meeting notes, whatever. Tracking current sota approaches while being all local

GitHub Stars: 8.6k

Repository: https://github.com/tobi/qmd

---

## Articles

### [How the LLM Cache in QMD Delivers 10-30× Performance Improvements](/tobi/qmd/how-llm-cache-qmd-improves-performance)

Discover how the QMD LLM cache achieves 10-30x performance gains by storing responses locally, slashing latency from seconds to milliseconds and eliminating redundant network calls.

- Tags: performance
- Published: 2026-02-16

### [Which GGUF Models Does QMD Support and How to Configure Them](/tobi/qmd/gguf-models-supported-qmd-configuration)

Discover which GGUF models QMD supports for embeddings, reranking, and query expansion. Learn to configure custom HuggingFace GGUF models with LlamaCppConfig.

- Tags: how-to-guide
- Published: 2026-02-16

### [Migrating from Another Search System to QMD: A Complete Step-by-Step Guide](/tobi/qmd/migrate-from-other-search-system-to-qmd)

Seamlessly migrate your search system to QMD. Follow our step-by-step guide to install, configure collections, update the index, and enable semantic search with embeddings.

- Tags: migration-guide
- Published: 2026-02-16

### [QMD CPU vs GPU Performance: Benchmarks and Device Selection Guide](/tobi/qmd/qmd-performance-cpu-vs-gpu)

Discover QMD CPU vs GPU performance differences. See benchmarks showing 10x faster inference on GPU, reducing latency from 300ms to 30ms. Choose the right device for your needs.

- Tags: performance
- Published: 2026-02-16

### [QMD vs ripgrep: Semantic Search vs Lexical Pattern Matching](/tobi/qmd/qmd-vs-ripgrep-local-search-tools)

Discover how QMD semantic search with vector embeddings outperforms ripgrep's regex scans for context-aware, ranked local search results. Compare QMD and ripgrep.

- Tags: comparison
- Published: 2026-02-16

### [Chunk Size and Overlap Settings for QMD Embeddings: A Technical Deep Dive](/tobi/qmd/factors-determining-chunk-size-overlap-qmd-embeddings)

Discover how Tobi/qmd determines chunk size and overlap for QMD embeddings. Learn about the 900-token limit, 15% overlap ratio, and intelligent boundary detection for optimal context.

- Tags: deep-dive
- Published: 2026-02-16

### [How the Top-Rank Bonus in QMD Fusion Preserves Exact Matches](/tobi/qmd/how-top-rank-bonus-qmd-fusion-preserves-exact-matches)

Discover how the QMD Fusion top-rank bonus preserves exact matches by adding score increments, preventing high-precision results from being lost.

- Tags: internals
- Published: 2026-02-16

### [How to Use QMD to Manage Collections on Network Drives: A Complete Guide](/tobi/qmd/can-qmd-manage-collections-network-drives)

Learn how QMD manages collections on network drives with this comprehensive guide. QMD unifies filesystem paths for seamless network storage management.

- Tags: how-to-guide
- Published: 2026-02-16

### [Database Schema for a QMD Index: Complete SQLite Structure Explained](/tobi/qmd/database-schema-qmd-index)

Explore the complete SQLite database schema for a QMD index. Understand the seven core tables storing metadata, content, vectors, and FTS indexes for efficient document management in tobi/qmd.

- Tags: api-reference
- Published: 2026-02-16

### [How to Debug Irrelevant Search Results in QMD: 8 Common Causes and Fixes](/tobi/qmd/debug-irrelevant-search-results-qmd)

Debug irrelevant search results in QMD by understanding 8 common causes like BM25-only queries and stale indexes. Learn quick fixes to improve search accuracy.

- Tags: how-to-guide
- Published: 2026-02-16

### [How QMD Handles Large Files During Multi-Get Operations](/tobi/qmd/how-qmd-handles-large-files-multi-get-operations)

Learn how QMD handles large files during multi-get operations. Discover default limits, automatic skipping, and how to override settings for efficient data transfer.

- Tags: internals
- Published: 2026-02-16

### [How QMD Uses Position-Aware Blending for Intelligent Reranking](/tobi/qmd/describe-position-aware-blending-qmd-reranking)

Discover how QMD employs position-aware blending to intelligently rerank search results by fusing retrieval and neural scores. Learn about its tiered weighting system.

- Tags: internals
- Published: 2026-02-16

### [How to Run the QMD MCP Server in HTTP Mode for Shared Access](/tobi/qmd/run-qmd-mcp-server-http-mode-shared-access)

Learn how to run the QMD MCP server in HTTP mode. This allows multiple clients to share a single LLM instance, keeping models loaded in VRAM for faster access.

- Tags: how-to-guide
- Published: 2026-02-16

### [How Context Metadata Enhances Search Relevance in QMD](/tobi/qmd/how-context-metadata-enhances-search-relevance-qmd)

Discover how QMD context metadata boosts search relevance. Learn how hierarchical descriptions empower LLM rerankers for semantic matching beyond keywords.

- Tags: deep-dive
- Published: 2026-02-16

### [How to Use Docids to Reference Specific Documents in QMD Scripts](/tobi/qmd/how-to-use-docids-reference-documents-qmd-scripts)

Learn how to use docids for path-independent document retrieval in QMD scripts. This guide explains QMDs stable docid system for efficient CLI workflows.

- Tags: how-to-guide
- Published: 2026-02-16

### [How QMD Implements Query Expansion to Improve Search Results](/tobi/qmd/how-qmd-implements-query-expansion-improve-search-results)

Discover how QMD implements query expansion using a grammar guided LLM to generate typed query variations for enhanced search results. Explore lexical, vector, and hypothetical document expansions.

- Tags: internals
- Published: 2026-02-16

### [Why QMD Requires Separate Embedding Generation for Vector Search](/tobi/qmd/why-separate-embedding-generation-qmd-vector-search)

Discover why QMD needs separate embedding generation for vector search. Understand the cost-effective approach to processing expensive embeddings once and performing lightweight SQL lookups repeatedly.

- Tags: internals
- Published: 2026-02-16

### [How QMD Uses the RRF Fusion Algorithm to Combine Search Results](/tobi/qmd/explain-rrf-fusion-algorithm-qmd-combine-search-results)

Learn how QMD uses Reciprocal Rank Fusion RRF to combine BM25 and vector search results. Discover weighted fusion with a top-rank bonus for superior ranking accuracy.

- Tags: deep-dive
- Published: 2026-02-16

### [How QMD Detects Document Changes and Handles Updates: A Complete Technical Guide](/tobi/qmd/how-qmd-detects-document-changes-handles-updates)

Learn how QMD detects document changes using SHA-256 hashes and a SQLite database. Understand its four-state logic for handling updates efficiently.

- Tags: deep-dive
- Published: 2026-02-16

### [QMD Search Commands Explained: Differences Between search, vsearch, and query](/tobi/qmd/qmd-search-vsearch-query-command-differences)

Understand QMD's search commands. Learn the differences between search for BM25, vsearch for vector semantic search, and query for hybrid retrieval including LLM expansion.

- Tags: deep-dive
- Published: 2026-02-16

### [How to Configure Multiple QMD Collections with Different Glob Patterns](/tobi/qmd/configure-multiple-qmd-collections-glob-patterns)

Learn to configure multiple QMD collections with unique glob patterns using the qmd collection add command and the --mask flag. Manage your files efficiently.

- Tags: how-to-guide
- Published: 2026-02-16

### [How QMD Combines BM25, Vector Search, and LLM Reranking for Hybrid Search](/tobi/qmd/how-qmd-combines-bm25-vector-search-llm-reranking-hybrid-search)

Discover how QMD's eight-stage hybrid search pipeline uniquely merges BM25, vector search, and LLM reranking. Achieve superior relevance with advanced fusion and LLM techniques. Learn more today

- Tags: how-to-guide
- Published: 2026-02-16

### [How QMD's Smart Chunking Preserves Markdown Structure: A Technical Deep Dive](/tobi/qmd/how-qmd-smart-chunking-preserves-markdown-structure)

Discover how QMD's smart chunking preserves markdown structure. Learn the technical details of boundary scoring, fence protection, and decay algorithms for optimal code splitting that respects your formatting.

- Tags: deep-dive
- Published: 2026-02-16

### [How QMD's MCP Server Enables Claude Desktop and AI Agents to Interact with QMD](/tobi/qmd/qmd-mcp-server-functionality-claude-desktop-ai-agents)

Discover how QMD's MCP server bridges Claude Desktop and AI agents. Enable real-time document search and queries via stdio or HTTP with structured JSON messages.

- Tags: internals
- Published: 2026-02-16

### [How to Use the QMD --index Option to Manage Separate Knowledge Bases with Named Indexes](/tobi/qmd/qmd-use-named-indexes-separate-knowledge-bases)

Learn to manage separate knowledge bases with named indexes using the QMD --index option. Isolate documents, embeddings, and collections into independent SQLite databases and YAML files. Build multiple knowledge bases easily.

- Tags: how-to-guide
- Published: 2026-02-16

### [How QMD's Docid System Generates 6-Character Content Hashes](/tobi/qmd/qmd-docid-system-content-hash-generation)

Learn how QMD's docid system generates 6-character content hashes using SHA-256 for content-addressed storage without filenames or timestamps.

- Tags: internals
- Published: 2026-02-16

### [How QMD's MCP HTTP Daemon Manages Model Loading for Maximum Efficiency](/tobi/qmd/qmd-mcp-http-daemon-model-loading)

Discover how QMD's MCP HTTP daemon efficiently loads GGUF models, keeping them in memory across requests by default and only disposing lightweight contexts to maximize performance.

- Tags: internals
- Published: 2026-02-16

### [How to Configure Custom GGUF Models for Embedding, Reranking, and Query Expansion in QMD](/tobi/qmd/qmd-configure-custom-gguf-models)

Configure custom GGUF models in QMD for embedding reranking and query expansion using the LlamaCpp class. Override defaults for powerful text analysis.

- Tags: how-to-guide
- Published: 2026-02-16

### [How QMD's Smart Chunking Algorithm Identifies Markdown Boundaries for Optimal Tokenization](/tobi/qmd/qmd-smart-chunking-markdown-boundaries)

Discover how QMD's smart chunking algorithm finds markdown boundaries for efficient tokenization. Learn about its structural break point scanning, semantic scoring, and distance-decay function.

- Tags: internals
- Published: 2026-02-15

