How LLM Wiki Performs Tokenized Search Across Wiki and Source Directories
LLM Wiki performs tokenized search by normalizing queries into deduplicated tokens in the TypeScript frontend, then passing them to a Rust backend that scans both markdown wiki files and plaintext source files, scoring matches using configurable weights before returning ranked results to the UI.
LLM Wiki (nashsu/llm_wiki) is an open-source knowledge management application built with Tauri, combining a Rust backend with a TypeScript frontend. The application implements a unified search pipeline that applies the same tokenization logic to query both markdown documentation in the wiki directory and plaintext source files in the src tree, ensuring consistent relevance scoring across heterogeneous content types.
Frontend Query Tokenization
The search journey begins in src/lib/search.ts, where the searchWiki function prepares the user query for backend processing. This module handles all client-side normalization before invoking the Tauri command layer.
The tokenization pipeline performs the following transformations:
- Lowercasing and splitting – The raw query string is converted to lowercase and segmented on whitespace, punctuation, and a comprehensive list of stop-words defined in the
STOP_WORDSconstant. - CJK bigram generation – For Chinese, Japanese, and Korean text, the tokenizer generates bigram tokens (adjacent character pairs) alongside individual characters and the full phrase, ensuring partial matches capture relevant context.
- Deduplication – Duplicate tokens are removed using
[...new Set(tokens)]to optimize network payload and backend processing time.
Once normalized, the frontend calls the Tauri command search_project, passing the original query string along with optional embedding vectors for hybrid search scenarios.
Backend Search Execution
The Rust backend entry point resides in src-tauri/src/commands/search.rs. The search_project async command receives the query parameters and delegates to search_project_inner, which orchestrates the actual file system traversal and scoring.
To guarantee deterministic behavior, the backend re-runs tokenization via the tokenize_query function, mirroring the frontend logic exactly. This ensures that stop-word removal and CJK handling remain consistent regardless of which layer initiates the search.
Scanning Markdown Wiki Files
For wiki content, the backend constructs the path project_path/wiki and walks the directory using walkdir. Every file matching *.md is processed sequentially:
- Title extraction – The
extract_titlefunction parses markdown front-matter or the first heading to derive a human-readable title. - Content tokenization – The full markdown text is tokenized using the same rules applied to the query.
- Relevance scoring – The
score_filefunction calculates a composite score based on:- FILENAME_EXACT_BONUS – Exact matches between query tokens and the filename.
- PHRASE_IN_TITLE_BONUS – Full phrase occurrences in the extracted title.
- PHRASE_IN_CONTENT_PER_OCC – Per-occurrence weight for phrase matches in body text.
- TITLE_TOKEN_WEIGHT and CONTENT_TOKEN_WEIGHT – Individual token match weights for titles versus body content.
Indexing Source Directories with Any‑Txt
Beyond markdown wikis, LLM Wiki supports "any‑txt" search for plaintext source files (code, notes, configuration) residing in the src directory. The frontend logic in src/lib/anytxt-search.ts—referenced via hasConfiguredAnyTxt in src/lib/web-search.ts—indexes every file under the source tree.
This module reuses the identical token set generated for wiki queries and applies the same score_file logic, ensuring that a search for "error handling" returns consistently ranked results whether the term appears in a README.md or a src/error.rs source file.
Ranking Algorithm and Result Fusion
After scanning both directories, the backend aggregates results using a multi-stage ranking process:
- Keyword scoring – Raw token matches generate a
token_rankscore based on the weighted bonuses defined in the scoring routine. - Reciprocal Rank Fusion – If vector embeddings are provided, keyword results are merged with embedding-based hits using
apply_rrf_scores, which balances the two signals without requiring score normalization. - Graph expansion – Optional graph-based enrichment via
blend_graph_resultscan augment the result set, though tokenized matching remains the primary relevance signal.
The final output is a ProjectSearchResponse struct containing the search mode, ranked results, and hit counts, which the frontend maps to UI components displaying titles, snippets, and extracted image references.
Implementation Examples
Trigger a cross-directory search from the frontend:
import { searchWiki } from "@/lib/search";
const projectPath = "/Users/alice/knowledge-base";
const query = "self-attention mechanisms";
searchWiki(projectPath, query).then(results => {
console.log("Found", results.length, "pages:");
results.forEach(r => console.log(`${r.title} – ${r.snippet}`));
});
The Tauri command signature handling the request:
#[tauri::command]
pub async fn search_project(
project_path: String,
query: String,
top_k: Option<usize>,
include_content: Option<bool>,
query_embedding: Option<Vec<f32>>,
embedding_config: Option<SearchEmbeddingConfig>,
) -> Result<ProjectSearchResponse, String> {
// Tokenise query, walk wiki & source files, compute scores …
// (implementation lives in `search_project_inner`)
}
Example of CJK-aware tokenization:
import { tokenizeQuery } from "@/lib/search";
const tokens = tokenizeQuery("注意力机制");
// → ["注意力", "注意", "意力", "力机", "机制", "注意", "力"]
console.log(tokens);
Summary
- Unified tokenization – Both frontend (
src/lib/search.ts) and backend (src-tauri/src/commands/search.rs) use identicaltokenize_querylogic to handle stop-words, punctuation, and CJK bigrams. - Dual-directory scanning – The backend walks
project_path/wikifor markdown files and thesrctree for plaintext source files, applying the same scoring algorithm to both. - Configurable scoring – Relevance is calculated using weighted bonuses for filename matches, title phrases, and content tokens via the
score_fileroutine. - Hybrid fusion – Keyword results are merged with optional vector search hits using Reciprocal Rank Fusion (
apply_rrf_scores) before returning to the UI.
Frequently Asked Questions
How does LLM Wiki handle CJK characters in search queries?
LLM Wiki generates bigram tokens for CJK text by creating adjacent character pairs, individual characters, and the full phrase. This approach ensures that queries like "注意力机制" produce tokens covering both partial overlaps ("注意", "意力") and complete phrases, improving recall for ideographic languages without requiring language-specific dictionaries.
What distinguishes wiki search from source file search?
While both use the same tokenization and scoring logic, wiki search targets *.md files in the wiki directory and extracts structured titles via extract_title, whereas source file search (any‑txt) scans the src directory for all plaintext files. The any‑txt integration in src/lib/web-search.ts enables searching code comments and documentation alongside markdown articles using identical query tokens.
How are keyword and embedding results combined?
The backend uses Reciprocal Rank Fusion via apply_rrf_scores to merge token-based rankings (token_rank) with vector similarity scores. This method ranks results by their reciprocal positions in each list, eliminating the need to normalize scores between the two different search modalities while preserving the relevance signals from both.
Where is the tokenization logic defined to ensure frontend-backend parity?
The tokenization algorithm is implemented in both src/lib/search.ts for the frontend and src-tauri/src/commands/search.rs (within the tokenize_query function) for the backend. Both implementations share identical logic for lowercasing, stop-word filtering, punctuation splitting, and CJK bigram generation, ensuring that a query tokenized in the browser will match exactly against tokens generated from files on disk.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →