How to Use the LLM Wiki API for Hybrid Search: A Complete Guide
LLM Wiki exposes a local HTTP API at http://127.0.0.1:19828 that performs hybrid search by combining BM25 keyword scoring, vector similarity search via LanceDB, and optional graph-based link expansion through a single POST endpoint.
The open-source LLM Wiki repository (nashsu/llm_wiki) provides a self-hosted knowledge base with a built-in hybrid search engine. This architecture enables external scripts, AI agents, and frontend applications to execute complex retrieval-augmented generation (RAG) workflows without managing separate search infrastructure.
Understanding the Hybrid Search Architecture
When you send a request to the LLM Wiki API, the query is processed through a three-stage pipeline implemented in the Rust backend. The entry point is the search_project command located in [src-tauri/src/commands/search.rs](https://github.com/nashsu/llm_wiki/blob/main/src-tauri/src/commands/search.rs#L107-L127), which is exposed via the Tauri bridge.
The system merges results from three distinct retrieval strategies:
- Lexical matching using BM25-style heuristics with title and filename bonuses
- Semantic similarity using vector embeddings stored in LanceDB
- Graph expansion adding one-hop Wikipedia-style link context
This multi-modal approach ensures high recall for exact keyword matches while maintaining semantic understanding for conceptually related content.
Making API Requests to the LLM Wiki Hybrid Search Endpoint
The primary endpoint is POST /api/v1/projects/{id}/search. You must replace {id} with your project identifier. The server expects a JSON payload containing your query string and optional configuration parameters.
cURL Example
Use this command to test the endpoint from any terminal:
curl -X POST http://127.0.0.1:19828/api/v1/projects/12345/search \
-H "Content-Type: application/json" \
-d '{
"query": "how to reset my password",
"topK": 10,
"includeContent": false,
"queryEmbedding": null,
"embeddingConfig": null
}'
Python Implementation
For production Python applications, use the requests library to parse the ProjectSearchResponse:
import requests
url = "http://127.0.0.1:19828/api/v1/projects/12345/search"
payload = {
"query": "how to reset my password",
"topK": 10,
"includeContent": False,
"queryEmbedding": None,
"embeddingConfig": None # omit or supply an embedding config to enable vector search
}
resp = requests.post(url, json=payload)
data = resp.json()
print("Search mode:", data["mode"])
for hit in data["results"]:
print(f"- {hit['title']} ({hit['path']}) → score {hit['score']:.2f}")
TypeScript Frontend Integration
If you are building within the LLM Wiki ecosystem, use the provided helper in [src/lib/search.ts](https://github.com/nashsu/llm_wiki/blob/main/src/lib/search.ts):
import { searchWiki } from "@/lib/search";
const results = await searchWiki("/path/to/project", "how to reset my password");
console.log(results);
How Hybrid Ranking Works Internally
The LLM Wiki API implements a sophisticated ranking pipeline that processes your query through sequential retrieval stages before returning fused results.
Stage 1: Keyword Retrieval with BM25
The query string is first tokenized using the tokenize_query function in [src/lib/search.ts](https://github.com/nashsu/llm_wiki/blob/main/src/lib/search.ts#L38-L60). Each markdown file in the project's wiki/ folder receives a BM25-style relevance score that includes:
- Title-match bonus: Higher weight for terms appearing in document titles
- Filename exact bonus: Exact matches against file paths
- Phrase-in-content weighting: Positional scoring for query terms within body text
Results are sorted via token_rank and passed to the next stage.
Stage 2: Vector Retrieval and RRF Fusion
If an embedding configuration is provided (or server-side defaults are enabled), the system performs a parallel vector search. The embedding_fetch endpoint generates query embeddings, and search_by_embedding (lines 331-350 in search.rs) queries the LanceDB vector store for nearest neighbors.
The system merges keyword and vector results using Reciprocal Rank Fusion (apply_rrf_scores), which balances the ranked positions from both retrieval methods without requiring score normalization.
Stage 3: Graph Link Expansion
After fusion, the blend_graph_results function performs a lightweight graph pass that adds one-hop Wikipedia-style links to the result set. This respects the configured MIN_GRAPH_RESULT_RATIO and MAX_GRAPH_RESULT_RATIO constants, ensuring graph expansion supplements rather than overwhelms the primary search results.
Interpreting the Search Response
The API returns a ProjectSearchResponse object containing:
mode: Indicates which retrieval methods activated—"hybrid"when both token and vector sources contribute, otherwise"keyword"or"vector"(seesearch_mode)token_hits,vector_hits,graph_hits: Integer counts for each retrieval typeresults: An array of hit objects containing:path: File path relative to the wiki directorytitle: Document titlesnippet: Text excerpt around matching termsscore: Final fused relevance scorevectorScore: Raw similarity score when applicableimages: Array of embedded image references
This structured response allows clients to filter or rerank results based on confidence thresholds or specific hit type requirements.
Summary
- LLM Wiki provides a local HTTP API at
http://127.0.0.1:19828for hybrid search operations - The
/api/v1/projects/{id}/searchendpoint accepts POST requests and routes through thesearch_projectRust command - Search combines BM25 keyword retrieval, LanceDB vector similarity, and graph link expansion through Reciprocal Rank Fusion
- Configuration is controlled via the
embeddingConfigparameter and server-side constants likeMIN_GRAPH_RESULT_RATIO - Response objects include detailed metadata about which retrieval modes contributed to each result
Frequently Asked Questions
What is the default port for the LLM Wiki API?
The local HTTP server binds to port 19828 by default, making the base URL http://127.0.0.1:19828. This is defined in the API server implementation within src-tauri/src/api_server.rs.
How does LLM Wiki combine keyword and vector search results?
The system uses Reciprocal Rank Fusion (apply_rrf_scores) to merge ranked lists from the BM25 keyword stage and the LanceDB vector stage. This algorithm weights results by their inverse rank positions, ensuring high-ranked items from either method surface in the final output without requiring score calibration between the different retrieval modalities.
Can I disable vector search and use only keyword matching?
Yes. Set the queryEmbedding field to null and omit the embeddingConfig parameter in your request payload. When no embedding configuration is present, the search_project command skips the search_by_embedding stage and returns results flagged with mode "keyword" rather than "hybrid".
What embedding database does LLM Wiki use for vector search?
LLM Wiki uses LanceDB as its vector store, implemented in src-tauri/src/commands/vectorstore.rs. The hybrid search system queries this database during the optional vector retrieval stage to perform nearest-neighbor searches against pre-computed document embeddings.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →