# How LLM Wiki Performs Tokenized Search Across Wiki and Source Directories

> Discover how LLM Wiki executes tokenized search across wiki and source directories. Learn about its TypeScript frontend, Rust backend, and ranked result scoring for efficient data retrieval.

- Repository: [nash_su/llm_wiki](https://github.com/nashsu/llm_wiki)
- Tags: deep-dive
- Published: 2026-09-12

---

**LLM Wiki performs tokenized search by normalizing queries into deduplicated tokens in the TypeScript frontend, then passing them to a Rust backend that scans both markdown wiki files and plaintext source files, scoring matches using configurable weights before returning ranked results to the UI.**

LLM Wiki (nashsu/llm_wiki) is an open-source knowledge management application built with Tauri, combining a Rust backend with a TypeScript frontend. The application implements a unified search pipeline that applies the same tokenization logic to query both markdown documentation in the `wiki` directory and plaintext source files in the `src` tree, ensuring consistent relevance scoring across heterogeneous content types.

## Frontend Query Tokenization

The search journey begins in [`src/lib/search.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/search.ts), where the `searchWiki` function prepares the user query for backend processing. This module handles all client-side normalization before invoking the Tauri command layer.

The tokenization pipeline performs the following transformations:

- **Lowercasing and splitting** – The raw query string is converted to lowercase and segmented on whitespace, punctuation, and a comprehensive list of stop-words defined in the `STOP_WORDS` constant.
- **CJK bigram generation** – For Chinese, Japanese, and Korean text, the tokenizer generates bigram tokens (adjacent character pairs) alongside individual characters and the full phrase, ensuring partial matches capture relevant context.
- **Deduplication** – Duplicate tokens are removed using `[...new Set(tokens)]` to optimize network payload and backend processing time.

Once normalized, the frontend calls the Tauri command `search_project`, passing the original query string along with optional embedding vectors for hybrid search scenarios.

## Backend Search Execution

The Rust backend entry point resides in [`src-tauri/src/commands/search.rs`](https://github.com/nashsu/llm_wiki/blob/main/src-tauri/src/commands/search.rs). The `search_project` async command receives the query parameters and delegates to `search_project_inner`, which orchestrates the actual file system traversal and scoring.

To guarantee deterministic behavior, the backend re-runs tokenization via the `tokenize_query` function, mirroring the frontend logic exactly. This ensures that stop-word removal and CJK handling remain consistent regardless of which layer initiates the search.

### Scanning Markdown Wiki Files

For wiki content, the backend constructs the path `project_path/wiki` and walks the directory using `walkdir`. Every file matching `*.md` is processed sequentially:

1. **Title extraction** – The `extract_title` function parses markdown front-matter or the first heading to derive a human-readable title.
2. **Content tokenization** – The full markdown text is tokenized using the same rules applied to the query.
3. **Relevance scoring** – The `score_file` function calculates a composite score based on:
   - **FILENAME_EXACT_BONUS** – Exact matches between query tokens and the filename.
   - **PHRASE_IN_TITLE_BONUS** – Full phrase occurrences in the extracted title.
   - **PHRASE_IN_CONTENT_PER_OCC** – Per-occurrence weight for phrase matches in body text.
   - **TITLE_TOKEN_WEIGHT** and **CONTENT_TOKEN_WEIGHT** – Individual token match weights for titles versus body content.

### Indexing Source Directories with Any‑Txt

Beyond markdown wikis, LLM Wiki supports "any‑txt" search for plaintext source files (code, notes, configuration) residing in the `src` directory. The frontend logic in [`src/lib/anytxt-search.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/anytxt-search.ts)—referenced via `hasConfiguredAnyTxt` in [`src/lib/web-search.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/web-search.ts)—indexes every file under the source tree.

This module reuses the identical token set generated for wiki queries and applies the same `score_file` logic, ensuring that a search for "error handling" returns consistently ranked results whether the term appears in a [`README.md`](https://github.com/nashsu/llm_wiki/blob/main/README.md) or a [`src/error.rs`](https://github.com/nashsu/llm_wiki/blob/main/src/error.rs) source file.

## Ranking Algorithm and Result Fusion

After scanning both directories, the backend aggregates results using a multi-stage ranking process:

- **Keyword scoring** – Raw token matches generate a `token_rank` score based on the weighted bonuses defined in the scoring routine.
- **Reciprocal Rank Fusion** – If vector embeddings are provided, keyword results are merged with embedding-based hits using `apply_rrf_scores`, which balances the two signals without requiring score normalization.
- **Graph expansion** – Optional graph-based enrichment via `blend_graph_results` can augment the result set, though tokenized matching remains the primary relevance signal.

The final output is a `ProjectSearchResponse` struct containing the search mode, ranked results, and hit counts, which the frontend maps to UI components displaying titles, snippets, and extracted image references.

## Implementation Examples

Trigger a cross-directory search from the frontend:

```typescript
import { searchWiki } from "@/lib/search";

const projectPath = "/Users/alice/knowledge-base";
const query = "self-attention mechanisms";

searchWiki(projectPath, query).then(results => {
  console.log("Found", results.length, "pages:");
  results.forEach(r => console.log(`${r.title} – ${r.snippet}`));
});

```

The Tauri command signature handling the request:

```rust
#[tauri::command]
pub async fn search_project(
    project_path: String,
    query: String,
    top_k: Option<usize>,
    include_content: Option<bool>,
    query_embedding: Option<Vec<f32>>,
    embedding_config: Option<SearchEmbeddingConfig>,
) -> Result<ProjectSearchResponse, String> {
    // Tokenise query, walk wiki & source files, compute scores …
    // (implementation lives in `search_project_inner`)
}

```

Example of CJK-aware tokenization:

```typescript
import { tokenizeQuery } from "@/lib/search";

const tokens = tokenizeQuery("注意力机制");
// → ["注意力", "注意", "意力", "力机", "机制", "注意", "力"]
console.log(tokens);

```

## Summary

- **Unified tokenization** – Both frontend ([`src/lib/search.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/search.ts)) and backend ([`src-tauri/src/commands/search.rs`](https://github.com/nashsu/llm_wiki/blob/main/src-tauri/src/commands/search.rs)) use identical `tokenize_query` logic to handle stop-words, punctuation, and CJK bigrams.
- **Dual-directory scanning** – The backend walks `project_path/wiki` for markdown files and the `src` tree for plaintext source files, applying the same scoring algorithm to both.
- **Configurable scoring** – Relevance is calculated using weighted bonuses for filename matches, title phrases, and content tokens via the `score_file` routine.
- **Hybrid fusion** – Keyword results are merged with optional vector search hits using Reciprocal Rank Fusion (`apply_rrf_scores`) before returning to the UI.

## Frequently Asked Questions

### How does LLM Wiki handle CJK characters in search queries?

LLM Wiki generates bigram tokens for CJK text by creating adjacent character pairs, individual characters, and the full phrase. This approach ensures that queries like "注意力机制" produce tokens covering both partial overlaps ("注意", "意力") and complete phrases, improving recall for ideographic languages without requiring language-specific dictionaries.

### What distinguishes wiki search from source file search?

While both use the same tokenization and scoring logic, wiki search targets `*.md` files in the `wiki` directory and extracts structured titles via `extract_title`, whereas source file search (any‑txt) scans the `src` directory for all plaintext files. The any‑txt integration in [`src/lib/web-search.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/web-search.ts) enables searching code comments and documentation alongside markdown articles using identical query tokens.

### How are keyword and embedding results combined?

The backend uses Reciprocal Rank Fusion via `apply_rrf_scores` to merge token-based rankings (`token_rank`) with vector similarity scores. This method ranks results by their reciprocal positions in each list, eliminating the need to normalize scores between the two different search modalities while preserving the relevance signals from both.

### Where is the tokenization logic defined to ensure frontend-backend parity?

The tokenization algorithm is implemented in both [`src/lib/search.ts`](https://github.com/nashsu/llm_wiki/blob/main/src/lib/search.ts) for the frontend and [`src-tauri/src/commands/search.rs`](https://github.com/nashsu/llm_wiki/blob/main/src-tauri/src/commands/search.rs) (within the `tokenize_query` function) for the backend. Both implementations share identical logic for lowercasing, stop-word filtering, punctuation splitting, and CJK bigram generation, ensuring that a query tokenized in the browser will match exactly against tokens generated from files on disk.