How to Debug Irrelevant Search Results in QMD: 8 Common Causes and Fixes
Irrelevant search results in QMD usually stem from BM25-only queries without vector search, skipped query expansion due to strong signals, missing collection context, or stale indexes.
The tobi/qmd search engine combines BM25 keyword matching, vector semantic search, and LLM reranking to surface relevant markdown documents. When irrelevant search results in QMD appear, the issue typically lies in one of eight specific pipeline stages that can be diagnosed using built-in CLI flags and environment variables.
BM25-Only Queries Without Vector Search or Expansion
QMD’s hybrid pipeline starts with a BM25 keyword search via searchFTS() in src/store.ts (lines 1999–2035). If the initial probe reports a strong signal—a very high top score with a large gap to the second result—the engine skips LLM query expansion and vector search (lines 2829–2832).
When this happens, the source field in every result shows "fts", indicating no semantic search contributed to the ranking.
To verify and fix:
# Check which backend produced each result
qmd search "your query" --json | jq '.[] | {file:.filepath, source:.source}'
# If all sources are "fts", force expansion by disabling strong-signal detection
export QMD_STRONG_SIGNAL_MIN_SCORE=0
export QMD_STRONG_SIGNAL_MIN_GAP=0
qmd search "your query"
Disabled Vector Search or Embedding Model Failures
The vector-semantic search back-end relies on searchVec() (lines 560–595 in src/store.ts). If the vectors_vec table does not exist, or if getEmbedding() returns null due to an unsupported model, the function returns an empty list immediately (lines 560–562).
When vector search fails, the hybrid orchestrator falls back to BM25-only results, often producing generic matches.
To diagnose:
# Test vector search directly
qmd vsearch "machine learning" --json
# If you see "No vector embeddings. Only search (BM25) is available."
# (as defined in src/mcp.ts lines 116-118), verify the extension:
sqlite3 ~/.cache/qmd/index.sqlite "SELECT load_extension('sqlite-vec');"
To resolve, re-index with embeddings:
qmd embed --model nomic-embed-text
Query Expansion Skipped Due to Strong BM25 Signals
QMD’s hybridQuery() (lines 2946–3000) uses Reciprocal Rank Fusion to combine BM25 and vector candidates. However, if the initial BM25 probe detects a strong signal, the pipeline skips query expansion entirely (lines 2829–2832).
This optimization assumes the top BM25 hit is highly relevant. If that hit is a false positive—such as a generic README—the engine never explores alternative query formulations.
Debug by watching for the onStrongSignal hook output in verbose mode:
qmd search "ambiguous term" --verbose
If you see “Strong BM25 signal, skipping expansion,” override the thresholds:
export QMD_STRONG_SIGNAL_MIN_SCORE=0.0
export QMD_STRONG_SIGNAL_MIN_GAP=0.0
qmd search "ambiguous term"
Missing Collection Context Causing Generic Matches
QMD augments results with context snippets via getContextForFile(). When a collection lacks context, getCollectionsWithoutContext() (lines 1901–1923 in src/store.ts) reports the deficiency. Without these hints, the reranker cannot boost relevant sections, leading to generic matches.
Check for missing context:
qmd context check
Add context to improve relevance:
# Add context to the collection root
qmd context add . "Documentation for the QMD search engine"
# Add context to specific subdirectories
qmd context add src "Core TypeScript source files"
Stale Indexes After Document Changes
QMD relies on SQLite FTS5 and vector tables built during qmd embed or qmd collection add. If files change without re-indexing, searchFTS() and searchVec() return matches from old content or miss new files entirely.
Verify index freshness:
qmd status # shows document count per collection
qmd collection list # verify paths
Refresh the index:
qmd update --pull # re-index all collections
Incorrect Collection Names or Path Filters
Both searchFTS() and searchVec() accept an optional collectionName parameter. Supplying the wrong name restricts the search to an empty or tiny subset, forcing the engine to return the only available matches regardless of relevance.
Verify available collections:
qmd collection list
Search the correct collection:
qmd search -c mynotes "query term"
Low Minimum Scores and Aggressive Result Limits
The hybrid pipeline in hybridQuery() (lines 2946–3000) filters candidates using options.minScore (default 0). Combining a low threshold with a high --limit allows low-relevance hits to surface.
Tune the scoring:
qmd search "term" --limit 10 --min-score 0.4
Reranking Model Failures
After fusing candidates, QMD reranks the best chunk of each document using store.rerank. If the LLM backend (node-llama-cpp) is unavailable or returns NaN scores, the final order becomes unpredictable.
Check reranker health:
# Run the benchmark harness
npx ts-node src/bench-rerank.ts "test query"
Or test manually:
import { Store } from "./src/store";
const store = new Store(db);
const res = await store.rerank("test query", [
{file:"qmd://demo/doc.md", text:"sample chunk"}
]);
console.log(res);
Summary
- Check the source: Use
--jsonto verify if results come fromfts,vec, orhybridto identify missing vector contributions. - Force expansion: Set
QMD_STRONG_SIGNAL_MIN_SCORE=0andQMD_STRONG_SIGNAL_MIN_GAP=0to disable strong-signal optimization when top hits are false positives. - Verify context: Run
qmd context checkand add context to collections to improve reranking precision. - Refresh indexes: Execute
qmd update --pullafter file changes to prevent stale results. - Tune thresholds: Use
--min-scoreand--limitto filter out low-relevance candidates from the hybrid fusion pipeline.
Frequently Asked Questions
Why do my QMD searches only return results from the FTS backend?
When every result shows "source": "fts" in JSON output, the hybrid pipeline detected a strong BM25 signal and skipped vector search and query expansion. This happens in src/store.ts (lines 2829–2832) when the top BM25 score exceeds the strong-signal threshold. Disable this optimization by setting QMD_STRONG_SIGNAL_MIN_SCORE=0 and QMD_STRONG_SIGNAL_MIN_GAP=0 before running your search.
How can I tell if my vector embeddings are working correctly?
Run qmd vsearch "your query" --json. If you receive the message “No vector embeddings. Only search (BM25) is available.” (defined in src/mcp.ts lines 116–118), the vectors_vec table is missing or searchVec() (lines 560–562 in src/store.ts) returned empty results. Resolve this by running qmd embed to regenerate embeddings with your configured model.
What causes generic, seemingly random results at the top of my search?
Generic matches often indicate missing collection context. When getContextForFile() returns null for a collection, the reranker lacks semantic hints to boost relevant sections. Use qmd context check (backed by getCollectionsWithoutContext() in src/store.ts lines 1901–1923) to identify collections without context, then add descriptive context using qmd context add <path> "<description>".
Why do I get outdated results after editing my markdown files?
QMD maintains separate FTS5 and vector tables that are only updated during indexing operations. If you modify files without re-indexing, searchFTS() and searchVec() reference stale data. Run qmd status to verify document counts, then execute qmd update --pull to refresh the SQLite indexes and align them with your current filesystem.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →