How to Use the LLM Wiki API for Hybrid Search: A Complete Guide

LLM Wiki exposes a local HTTP API at http://127.0.0.1:19828 that performs hybrid search by combining BM25 keyword scoring, vector similarity search via LanceDB, and optional graph-based link expansion through a single POST endpoint.

The open-source LLM Wiki repository (nashsu/llm_wiki) provides a self-hosted knowledge base with a built-in hybrid search engine. This architecture enables external scripts, AI agents, and frontend applications to execute complex retrieval-augmented generation (RAG) workflows without managing separate search infrastructure.

Understanding the Hybrid Search Architecture

When you send a request to the LLM Wiki API, the query is processed through a three-stage pipeline implemented in the Rust backend. The entry point is the search_project command located in [src-tauri/src/commands/search.rs](https://github.com/nashsu/llm_wiki/blob/main/src-tauri/src/commands/search.rs#L107-L127), which is exposed via the Tauri bridge.

The system merges results from three distinct retrieval strategies:

  • Lexical matching using BM25-style heuristics with title and filename bonuses
  • Semantic similarity using vector embeddings stored in LanceDB
  • Graph expansion adding one-hop Wikipedia-style link context

This multi-modal approach ensures high recall for exact keyword matches while maintaining semantic understanding for conceptually related content.

Making API Requests to the LLM Wiki Hybrid Search Endpoint

The primary endpoint is POST /api/v1/projects/{id}/search. You must replace {id} with your project identifier. The server expects a JSON payload containing your query string and optional configuration parameters.

cURL Example

Use this command to test the endpoint from any terminal:

curl -X POST http://127.0.0.1:19828/api/v1/projects/12345/search \
  -H "Content-Type: application/json" \
  -d '{
        "query": "how to reset my password",
        "topK": 10,
        "includeContent": false,
        "queryEmbedding": null,
        "embeddingConfig": null
      }'

Python Implementation

For production Python applications, use the requests library to parse the ProjectSearchResponse:

import requests

url = "http://127.0.0.1:19828/api/v1/projects/12345/search"
payload = {
    "query": "how to reset my password",
    "topK": 10,
    "includeContent": False,
    "queryEmbedding": None,
    "embeddingConfig": None      # omit or supply an embedding config to enable vector search

}
resp = requests.post(url, json=payload)
data = resp.json()

print("Search mode:", data["mode"])
for hit in data["results"]:
    print(f"- {hit['title']} ({hit['path']}) → score {hit['score']:.2f}")

TypeScript Frontend Integration

If you are building within the LLM Wiki ecosystem, use the provided helper in [src/lib/search.ts](https://github.com/nashsu/llm_wiki/blob/main/src/lib/search.ts):

import { searchWiki } from "@/lib/search";

const results = await searchWiki("/path/to/project", "how to reset my password");
console.log(results);

How Hybrid Ranking Works Internally

The LLM Wiki API implements a sophisticated ranking pipeline that processes your query through sequential retrieval stages before returning fused results.

Stage 1: Keyword Retrieval with BM25

The query string is first tokenized using the tokenize_query function in [src/lib/search.ts](https://github.com/nashsu/llm_wiki/blob/main/src/lib/search.ts#L38-L60). Each markdown file in the project's wiki/ folder receives a BM25-style relevance score that includes:

  • Title-match bonus: Higher weight for terms appearing in document titles
  • Filename exact bonus: Exact matches against file paths
  • Phrase-in-content weighting: Positional scoring for query terms within body text

Results are sorted via token_rank and passed to the next stage.

Stage 2: Vector Retrieval and RRF Fusion

If an embedding configuration is provided (or server-side defaults are enabled), the system performs a parallel vector search. The embedding_fetch endpoint generates query embeddings, and search_by_embedding (lines 331-350 in search.rs) queries the LanceDB vector store for nearest neighbors.

The system merges keyword and vector results using Reciprocal Rank Fusion (apply_rrf_scores), which balances the ranked positions from both retrieval methods without requiring score normalization.

After fusion, the blend_graph_results function performs a lightweight graph pass that adds one-hop Wikipedia-style links to the result set. This respects the configured MIN_GRAPH_RESULT_RATIO and MAX_GRAPH_RESULT_RATIO constants, ensuring graph expansion supplements rather than overwhelms the primary search results.

Interpreting the Search Response

The API returns a ProjectSearchResponse object containing:

  • mode: Indicates which retrieval methods activated—"hybrid" when both token and vector sources contribute, otherwise "keyword" or "vector" (see search_mode)
  • token_hits, vector_hits, graph_hits: Integer counts for each retrieval type
  • results: An array of hit objects containing:
    • path: File path relative to the wiki directory
    • title: Document title
    • snippet: Text excerpt around matching terms
    • score: Final fused relevance score
    • vectorScore: Raw similarity score when applicable
    • images: Array of embedded image references

This structured response allows clients to filter or rerank results based on confidence thresholds or specific hit type requirements.

Summary

  • LLM Wiki provides a local HTTP API at http://127.0.0.1:19828 for hybrid search operations
  • The /api/v1/projects/{id}/search endpoint accepts POST requests and routes through the search_project Rust command
  • Search combines BM25 keyword retrieval, LanceDB vector similarity, and graph link expansion through Reciprocal Rank Fusion
  • Configuration is controlled via the embeddingConfig parameter and server-side constants like MIN_GRAPH_RESULT_RATIO
  • Response objects include detailed metadata about which retrieval modes contributed to each result

Frequently Asked Questions

What is the default port for the LLM Wiki API?

The local HTTP server binds to port 19828 by default, making the base URL http://127.0.0.1:19828. This is defined in the API server implementation within src-tauri/src/api_server.rs.

How does LLM Wiki combine keyword and vector search results?

The system uses Reciprocal Rank Fusion (apply_rrf_scores) to merge ranked lists from the BM25 keyword stage and the LanceDB vector stage. This algorithm weights results by their inverse rank positions, ensuring high-ranked items from either method surface in the final output without requiring score calibration between the different retrieval modalities.

Can I disable vector search and use only keyword matching?

Yes. Set the queryEmbedding field to null and omit the embeddingConfig parameter in your request payload. When no embedding configuration is present, the search_project command skips the search_by_embedding stage and returns results flagged with mode "keyword" rather than "hybrid".

LLM Wiki uses LanceDB as its vector store, implemented in src-tauri/src/commands/vectorstore.rs. The hybrid search system queries this database during the optional vector retrieval stage to perform nearest-neighbor searches against pre-computed document embeddings.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →