How Vector Search Is Configured and Used in LLM Wiki: A Complete LanceDB Implementation Guide

LLM Wiki implements vector search using a local LanceDB instance managed by a Rust backend, where markdown chunks are embedded via configurable providers (OpenAI, Anthropic, Gemini, or Ollama) and retrieved through Tauri commands vector_upsert_chunks and vector_search_chunks, supporting keyword, vector, and hybrid search modes with BM25 fallback and Reciprocal Rank Fusion.

LLM Wiki is an open-source knowledge management application that combines markdown editing with semantic search capabilities. The repository nashsu/llm_wiki stores and retrieves semantic embeddings of wiki pages using LanceDB as the embedded vector store, orchestrated through Tauri commands for cross-platform native performance. This article examines the configuration, indexing workflow, and search APIs that enable semantic retrieval within the application.

Vector Database Configuration

LanceDB Local Storage Architecture

LLM Wiki uses LanceDB as its vector database, storing data locally within each project's .llm_wiki directory. Unlike external vector databases that require network connections, the LanceDB files live on disk and are accessed directly by the Rust backend. This eliminates the need for external service URLs and ensures offline functionality. The database is initialized and managed through the Tauri command layer in src-tauri/src/commands/page_embedding.rs.

Embedding Provider Configuration

The system supports multiple embedding providers selected at runtime via the embeddingProvider field in the project configuration (defined in src/lib/config.ts). Supported providers include:

  • OpenAI (text-embedding-ada-002 and related models)
  • Anthropic (Claude embedding models)
  • Google Gemini (Gemini embedding API)
  • Ollama (local models via OpenAI-compatible batch API)

When a page is indexed or a query is performed, the frontend calls embedBatch() or embedQuery() from src/lib/embedding.ts, which routes the request to the configured provider before passing the resulting float vectors to the LanceDB backend.

Indexing Workflow: From Markdown to Vectors

Text Chunking Strategy

Before embedding, markdown content is split into appropriately sized chunks to balance semantic coherence with vector granularity. The src/lib/text-chunker.ts module handles this segmentation, breaking pages into fragments that preserve context while staying within token limits.

Vector Upsert Process

The indexing pipeline follows three steps:

  1. Chunk Generation: The markdown is parsed and split into chunks with metadata.
  2. Embedding Generation: Each chunk text is sent to the configured provider via embedBatch() in src/lib/embedding.ts.
  3. Database Insertion: Vectors are persisted via the vector_upsert_chunks Tauri command.

The frontend wrapper vectorUpsertChunks() in src/lib/embedding.ts serializes the data and invokes the backend:

import { vectorUpsertChunks } from "./lib/embedding";

async function indexPage(projectPath: string, pageId: string, markdown: string) {
  // 1. Split markdown into chunks
  const chunks = await chunkMarkdown(markdown);
  
  // 2. Generate embeddings for all chunks
  const texts = chunks.map(c => c.text);
  const embeddings = await embedBatch(texts);
  
  // 3. Upsert to LanceDB via Tauri command
  await vectorUpsertChunks(projectPath, pageId, chunks.map((chunk, i) => ({
    chunk_index: i,
    chunk_text: chunk.text,
    embedding: embeddings[i],
  })));
}

On the backend, src-tauri/src/commands/page_embedding.rs implements vector_upsert_chunks, which writes the vectors and metadata to the LanceDB table using SQL syntax: INSERT INTO chunks ... or equivalent upsert operations.

Search Implementation and Query Modes

When a user submits a query with mode: "vector", the system embeds the query text using the same provider configured for indexing, then executes a similarity search. The vectorSearchChunks() function in src/lib/embedding.ts invokes vector_search_chunks on the backend:

import { search } from "./lib/search";

async function semanticSearch(projectPath: string, query: string) {
  const results = await search({
    projectPath,
    query,
    mode: "vector",
    topK: 10,
  });
  
  // Results contain vectorScore indicating L2 distance
  return results.map(r => ({
    title: r.title,
    snippet: r.snippet,
    score: r.vectorScore,
  }));
}

The backend command in src-tauri/src/commands/search.rs constructs a LanceDB query using L2 distance: SELECT * FROM chunks ORDER BY l2_distance(?, embedding) LIMIT ?. This returns chunk IDs and similarity scores that the frontend maps back to page metadata.

Hybrid Search with Reciprocal Rank Fusion

For mode: "hybrid", LLM Wiki executes both vector and keyword (BM25) searches concurrently, then merges results using Reciprocal Rank Fusion (RRF). The search() function in src/lib/search.ts coordinates this:

  • Keyword Search: Executes BM25 full-text search against the page content.
  • Vector Search: Runs the L2 distance query against chunk embeddings.
  • RRF Merging: src/lib/search-rrf.ts combines the two ranked lists using the standard RRF formula: score = Σ(1 / (k + rank)) where k is typically 60.

If the vector search returns zero hits (e.g., when embeddings are missing), the system automatically falls back to keyword results, ensuring users always receive results regardless of indexing state.

const hybridResults = await search({
  projectPath,
  query: "retrieval augmented generation",
  mode: "hybrid",
  topK: 20,
});

Vector Maintenance Commands

Beyond search and insertion, the backend exposes additional vector operations in src-tauri/src/commands/page_embedding.rs:

  • vector_delete_page: Removes all chunks associated with a specific page ID.
  • vector_optimize_chunks: Compacts and optimizes the LanceDB table for performance.
  • vector_clear_chunks: Truncates the entire vector table for re-indexing.

Key Source Files

File Responsibility Key Exports/Commands
src/lib/embedding.ts Frontend embedding interface embedQuery(), embedBatch(), vectorUpsertChunks(), vectorSearchChunks()
src/lib/search.ts Public search API search(), mode routing, fallback logic
src/lib/search-rrf.ts Rank fusion algorithm RRF implementation for hybrid results
src/lib/text-chunker.ts Content segmentation chunkMarkdown()
src-tauri/src/commands/page_embedding.rs Backend vector persistence vector_upsert_chunks, vector_delete_page, vector_optimize_chunks
src-tauri/src/commands/search.rs Backend query execution vector_search_chunks, distance calculations

Summary

  • LLM Wiki uses LanceDB as a local vector store, eliminating external dependencies and enabling offline semantic search.
  • The embedding pipeline in src/lib/embedding.ts abstracts multiple providers (OpenAI, Anthropic, Gemini, Ollama) and exposes them through Tauri commands.
  • Two primary commands handle vector operations: vector_upsert_chunks for indexing and vector_search_chunks for retrieval.
  • Three search modes are available: keyword (BM25), vector (L2 similarity), and hybrid (RRF fusion), with automatic fallback to keyword if vector results are empty.
  • All vector data resides in the project's .llm_wiki directory, managed by Rust commands in src-tauri/src/commands/.

Frequently Asked Questions

What vector database does LLM Wiki use?

LLM Wiki uses LanceDB, an embedded vector database that stores data locally in the .llm_wiki folder of each project. This allows the application to perform semantic search without requiring external database servers or network connectivity, as the Rust backend directly manages the LanceDB files through SQL-like queries executed via Tauri commands.

How does LLM Wiki generate embeddings for wiki pages?

The application breaks markdown content into chunks using src/lib/text-chunker.ts, then sends these chunks to a configurable embedding provider (OpenAI, Anthropic, Gemini, or Ollama) through src/lib/embedding.ts. The resulting float vectors are passed to the vector_upsert_chunks Tauri command, which writes them to LanceDB along with metadata linking vectors to specific page chunks.

What happens if vector search returns no results?

If vector_search_chunks returns zero hits (for example, if a page hasn't been indexed yet), the search logic in src/lib/search.ts automatically falls back to BM25 keyword search. This ensures users receive lexical matches even when semantic embeddings are unavailable, providing a robust search experience across all content states.

Can I use local models for embeddings instead of cloud APIs?

Yes. LLM Wiki supports Ollama and other local providers through the OpenAI-compatible batch API. By setting the embeddingProvider configuration to a local endpoint, src/lib/embedding.ts routes embedding requests to your local model via invoke("embedding_batch", ...), allowing completely offline operation without sending data to external services.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →