# How the Search Results Reranker Service Operates in MiniSearch

> Discover how MiniSearch reranks search results using a local LLM for enhanced relevance. Learn about its Jina model integration and statistical filtering.

- Repository: [Victor Nogueira/minisearch](https://github.com/felladrin/minisearch)
- Tags: internals
- Published: 2026-03-01

---

**MiniSearch boosts search relevance by processing raw web results through a local LLM reranker that downloads a quantized Jina model, spawns a llama-server instance, and applies statistical filtering to reorder results by semantic relevance.**

MiniSearch is an open-source search interface that improves result quality through an on-device reranking pipeline. The **search results reranker service** orchestrates a tiny LLM to score and reorder raw web-search hits before presenting them to users, operating entirely on `localhost:8012` without external API dependencies.

## Service Initialization and Model Acquisition

When the Vite development or preview server starts, the `rerankerServiceHook` in [`server/rerankerServiceHook.ts`](https://github.com/felladrin/minisearch/blob/main/server/rerankerServiceHook.ts) invokes `startRerankerService` from [`server/rerankerService.ts`](https://github.com/felladrin/minisearch/blob/main/server/rerankerService.ts). This function first ensures the model exists locally by calling `ensureModelExists`, which uses `downloadFileFromHuggingFaceRepository` to fetch the GGUF file from `Felladrin/gguf-jina-reranker-v1-tiny-en` if not already cached.

The model path is computed by `getRerankerModelPath` and resolves to `server/models/Felladrin/gguf-jina-reranker-v1-tiny-en/jina-reranker-v1-tiny-en-Q8_0.gguf`. Once confirmed, the service spawns the `llama-server` binary with flags enabling **RERANK mode**, specific context size, and batch size configurations.

## Health Monitoring and Readiness Probing

After spawning the process stored in `serverProcess`, the service enters a health-check loop that repeatedly polls `http://127.0.0.1:8012/health`. When the endpoint returns `{status: "ok"}`, a warm-up POST request to `/v1/rerank` validates that the model can actually rank documents. Only upon successful validation does the `isReady` flag become `true`, indicating the **search results reranker service** is operational.

If the `llama-server` process exits due to binary incompatibility or crashes, a 5-second restart timer automatically relaunches the service. During graceful shutdown, `rerankerServiceHook` listens for the Vite server's `close` event and invokes `stopRerankerService` to kill the child process and clear pending timers.

## Unicode Sanitization and Document Preparation

Before any reranking call, `sanitizeUnicodeSurrogates` (lines 26-58 in [`server/rerankerService.ts`](https://github.com/felladrin/minisearch/blob/main/server/rerankerService.ts)) fixes malformed UTF-16 surrogate pairs in both the query and document strings. This prevents crashes in the native server when processing web content containing invalid Unicode sequences.

The `rankSearchResults` function in [`server/rankSearchResults.ts`](https://github.com/felladrin/minisearch/blob/main/server/rankSearchResults.ts) converts raw search tuples—formatted as `[title, snippet, url]` arrays—into short Markdown snippets suitable for the reranker context window.

## The Reranking API Workflow

The core `rerank` function sends a POST request to `http://127.0.0.1:8012/v1/rerank` with a JSON body containing:

- The sanitized query string
- An array of document snippets
- `top_n` set to the document count

The server responds with an array of objects containing `index` and `relevance_score` fields. These scores drive the reordering logic in `rankSearchResults`, which reorders the original result tuples accordingly. An optional `preserveTopResults` parameter keeps the first raw hit untouched, preserving the most obvious match.

## Statistical Filtering and Thresholds

After receiving scores, `filterResultsByScore` normalizes values by shifting them by the absolute minimum to eliminate negatives. It then applies a two-stage filter:

- **Primary threshold**: Results must exceed **mean – k·σ** (where *k* = 0.3) to survive statistical outliers.
- **Fallback threshold**: If too few items remain, the system keeps any result whose normalized score exceeds **0.4 × maxScore**.

This approach removes low-confidence outliers while guaranteeing a minimum proportion of results for the user.

## Error Handling and Service Recovery

Any non-OK HTTP response from the reranker logs the full payload, throws an exception, and sets `isReady` to `false`. The failing `llama-server` process is immediately killed, triggering the automatic restart logic. This ensures the **search results reranker service** remains resilient against model loading errors or runtime crashes.

## Code Examples

```typescript
// Example: ranking search results for a user query
import { rankSearchResults } from "@/server/rankSearchResults";

// Raw search results come from the web-search service as tuples:
const rawResults: [string, string, string][] = [
  ["Cats", "Cats are small domesticated felines...", "https://en.wikipedia.org/wiki/Cat"],
  ["Caterpillars", "The larval stage of butterflies...", "https://en.wikipedia.org/wiki/Caterpillar"],
];

async function getRanked(query: string) {
  // `preserveTopResults = true` keeps the first raw hit (often the most obvious).
  const ranked = await rankSearchResults(query, rawResults, true);
  console.log(ranked);
}

getRanked("domestic felines");

```

```typescript
// Direct low-level use of the reranker (rarely needed)
import { rerank } from "@/server/rerankerService";

const query = "best Python web frameworks";
const docs = [
  "Django – a high-level Python web framework that encourages rapid development.",
  "Flask – a lightweight WSGI web application framework."
];

rerank(query, docs)
  .then(scores => console.log(scores))
  .catch(err => console.error("Reranker failure:", err));

```

## Summary

- The service downloads the GGUF model from Hugging Face on first run, caching it in `server/models/`.
- It spawns `llama-server` in RERANK mode on port 8012 with automated health checks and warm-up validation.
- `sanitizeUnicodeSurrogates` prevents crashes by fixing UTF-16 surrogate pairs in web content.
- Results are scored via the `/v1/rerank` endpoint and statistically filtered using mean-minus-sigma thresholds.
- Process crashes trigger automatic restarts with 5-second delays, while graceful shutdown hooks ensure clean termination.

## Frequently Asked Questions

### What model powers the MiniSearch reranker?

The service uses `Felladrin/gguf-jina-reranker-v1-tiny-en`, specifically the `jina-reranker-v1-tiny-en-Q8_0.gguf` quantized model file. According to the source code in [`server/rerankerService.ts`](https://github.com/felladrin/minisearch/blob/main/server/rerankerService.ts), the `downloadFileFromHuggingFaceRepository` function fetches this file from Hugging Face if it isn't already present in the local `server/models/` directory.

### How does the service recover from crashes or binary incompatibilities?

The `serverProcess` variable stores the spawned `llama-server` instance. If the process exits unexpectedly, a 5-second restart timer in [`server/rerankerService.ts`](https://github.com/felladrin/minisearch/blob/main/server/rerankerService.ts) automatically relaunches the service. Additionally, any failing HTTP request sets `isReady` to `false`, kills the process, and triggers the same recovery logic.

### What statistical method filters out low-relevance results?

The `filterResultsByScore` function applies a two-stage approach. First, it normalizes scores and applies a threshold of **mean – 0.3·σ** (standard deviations) to remove statistical outliers. If this filter removes too many items, a fallback keeps any result scoring above **40% of the maximum relevance score**, ensuring a minimum viable result set.

### Why is Unicode sanitization necessary before reranking?

The `sanitizeUnicodeSurrogates` function in [`server/rerankerService.ts`](https://github.com/felladrin/minisearch/blob/main/server/rerankerService.ts) (lines 26-58) removes malformed UTF-16 surrogate pairs from both queries and documents. This preprocessing prevents crashes in the native `llama-server` binary, which cannot handle invalid Unicode sequences that commonly appear in raw web search results.