How Vector Search Is Configured and Used in LLM Wiki: A Complete LanceDB Implementation Guide
LLM Wiki implements vector search using a local LanceDB instance managed by a Rust backend, where markdown chunks are embedded via configurable providers (OpenAI, Anthropic, Gemini, or Ollama) and retrieved through Tauri commands vector_upsert_chunks and vector_search_chunks, supporting keyword, vector, and hybrid search modes with BM25 fallback and Reciprocal Rank Fusion.
LLM Wiki is an open-source knowledge management application that combines markdown editing with semantic search capabilities. The repository nashsu/llm_wiki stores and retrieves semantic embeddings of wiki pages using LanceDB as the embedded vector store, orchestrated through Tauri commands for cross-platform native performance. This article examines the configuration, indexing workflow, and search APIs that enable semantic retrieval within the application.
Vector Database Configuration
LanceDB Local Storage Architecture
LLM Wiki uses LanceDB as its vector database, storing data locally within each project's .llm_wiki directory. Unlike external vector databases that require network connections, the LanceDB files live on disk and are accessed directly by the Rust backend. This eliminates the need for external service URLs and ensures offline functionality. The database is initialized and managed through the Tauri command layer in src-tauri/src/commands/page_embedding.rs.
Embedding Provider Configuration
The system supports multiple embedding providers selected at runtime via the embeddingProvider field in the project configuration (defined in src/lib/config.ts). Supported providers include:
- OpenAI (
text-embedding-ada-002and related models) - Anthropic (Claude embedding models)
- Google Gemini (Gemini embedding API)
- Ollama (local models via OpenAI-compatible batch API)
When a page is indexed or a query is performed, the frontend calls embedBatch() or embedQuery() from src/lib/embedding.ts, which routes the request to the configured provider before passing the resulting float vectors to the LanceDB backend.
Indexing Workflow: From Markdown to Vectors
Text Chunking Strategy
Before embedding, markdown content is split into appropriately sized chunks to balance semantic coherence with vector granularity. The src/lib/text-chunker.ts module handles this segmentation, breaking pages into fragments that preserve context while staying within token limits.
Vector Upsert Process
The indexing pipeline follows three steps:
- Chunk Generation: The markdown is parsed and split into chunks with metadata.
- Embedding Generation: Each chunk text is sent to the configured provider via
embedBatch()insrc/lib/embedding.ts. - Database Insertion: Vectors are persisted via the
vector_upsert_chunksTauri command.
The frontend wrapper vectorUpsertChunks() in src/lib/embedding.ts serializes the data and invokes the backend:
import { vectorUpsertChunks } from "./lib/embedding";
async function indexPage(projectPath: string, pageId: string, markdown: string) {
// 1. Split markdown into chunks
const chunks = await chunkMarkdown(markdown);
// 2. Generate embeddings for all chunks
const texts = chunks.map(c => c.text);
const embeddings = await embedBatch(texts);
// 3. Upsert to LanceDB via Tauri command
await vectorUpsertChunks(projectPath, pageId, chunks.map((chunk, i) => ({
chunk_index: i,
chunk_text: chunk.text,
embedding: embeddings[i],
})));
}
On the backend, src-tauri/src/commands/page_embedding.rs implements vector_upsert_chunks, which writes the vectors and metadata to the LanceDB table using SQL syntax: INSERT INTO chunks ... or equivalent upsert operations.
Search Implementation and Query Modes
Pure Vector Search
When a user submits a query with mode: "vector", the system embeds the query text using the same provider configured for indexing, then executes a similarity search. The vectorSearchChunks() function in src/lib/embedding.ts invokes vector_search_chunks on the backend:
import { search } from "./lib/search";
async function semanticSearch(projectPath: string, query: string) {
const results = await search({
projectPath,
query,
mode: "vector",
topK: 10,
});
// Results contain vectorScore indicating L2 distance
return results.map(r => ({
title: r.title,
snippet: r.snippet,
score: r.vectorScore,
}));
}
The backend command in src-tauri/src/commands/search.rs constructs a LanceDB query using L2 distance: SELECT * FROM chunks ORDER BY l2_distance(?, embedding) LIMIT ?. This returns chunk IDs and similarity scores that the frontend maps back to page metadata.
Hybrid Search with Reciprocal Rank Fusion
For mode: "hybrid", LLM Wiki executes both vector and keyword (BM25) searches concurrently, then merges results using Reciprocal Rank Fusion (RRF). The search() function in src/lib/search.ts coordinates this:
- Keyword Search: Executes BM25 full-text search against the page content.
- Vector Search: Runs the L2 distance query against chunk embeddings.
- RRF Merging:
src/lib/search-rrf.tscombines the two ranked lists using the standard RRF formula:score = Σ(1 / (k + rank))wherekis typically 60.
If the vector search returns zero hits (e.g., when embeddings are missing), the system automatically falls back to keyword results, ensuring users always receive results regardless of indexing state.
const hybridResults = await search({
projectPath,
query: "retrieval augmented generation",
mode: "hybrid",
topK: 20,
});
Vector Maintenance Commands
Beyond search and insertion, the backend exposes additional vector operations in src-tauri/src/commands/page_embedding.rs:
vector_delete_page: Removes all chunks associated with a specific page ID.vector_optimize_chunks: Compacts and optimizes the LanceDB table for performance.vector_clear_chunks: Truncates the entire vector table for re-indexing.
Key Source Files
| File | Responsibility | Key Exports/Commands |
|---|---|---|
src/lib/embedding.ts |
Frontend embedding interface | embedQuery(), embedBatch(), vectorUpsertChunks(), vectorSearchChunks() |
src/lib/search.ts |
Public search API | search(), mode routing, fallback logic |
src/lib/search-rrf.ts |
Rank fusion algorithm | RRF implementation for hybrid results |
src/lib/text-chunker.ts |
Content segmentation | chunkMarkdown() |
src-tauri/src/commands/page_embedding.rs |
Backend vector persistence | vector_upsert_chunks, vector_delete_page, vector_optimize_chunks |
src-tauri/src/commands/search.rs |
Backend query execution | vector_search_chunks, distance calculations |
Summary
- LLM Wiki uses LanceDB as a local vector store, eliminating external dependencies and enabling offline semantic search.
- The embedding pipeline in
src/lib/embedding.tsabstracts multiple providers (OpenAI, Anthropic, Gemini, Ollama) and exposes them through Tauri commands. - Two primary commands handle vector operations:
vector_upsert_chunksfor indexing andvector_search_chunksfor retrieval. - Three search modes are available:
keyword(BM25),vector(L2 similarity), andhybrid(RRF fusion), with automatic fallback to keyword if vector results are empty. - All vector data resides in the project's
.llm_wikidirectory, managed by Rust commands insrc-tauri/src/commands/.
Frequently Asked Questions
What vector database does LLM Wiki use?
LLM Wiki uses LanceDB, an embedded vector database that stores data locally in the .llm_wiki folder of each project. This allows the application to perform semantic search without requiring external database servers or network connectivity, as the Rust backend directly manages the LanceDB files through SQL-like queries executed via Tauri commands.
How does LLM Wiki generate embeddings for wiki pages?
The application breaks markdown content into chunks using src/lib/text-chunker.ts, then sends these chunks to a configurable embedding provider (OpenAI, Anthropic, Gemini, or Ollama) through src/lib/embedding.ts. The resulting float vectors are passed to the vector_upsert_chunks Tauri command, which writes them to LanceDB along with metadata linking vectors to specific page chunks.
What happens if vector search returns no results?
If vector_search_chunks returns zero hits (for example, if a page hasn't been indexed yet), the search logic in src/lib/search.ts automatically falls back to BM25 keyword search. This ensures users receive lexical matches even when semantic embeddings are unavailable, providing a robust search experience across all content states.
Can I use local models for embeddings instead of cloud APIs?
Yes. LLM Wiki supports Ollama and other local providers through the OpenAI-compatible batch API. By setting the embeddingProvider configuration to a local endpoint, src/lib/embedding.ts routes embedding requests to your local model via invoke("embedding_batch", ...), allowing completely offline operation without sending data to external services.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →