How Wigolo Implements Multi-Engine Search with Rank Fusion and ML Reranking
Wigolo’s search pipeline executes queries across multiple engines in parallel, merges the results using Reciprocal Rank Fusion (RRF) with per-engine weights, and optionally applies a transformer-based ML reranker to reorder the final list for maximum relevance.
The KnockOutEZ/wigolo repository provides a production-grade search aggregation layer designed to combine results from disparate search providers into a single, coherent ranking. Its multi-engine search implementation and rank fusion with ML reranking architecture separates concerns into three tightly-coupled stages: parallel engine dispatch, statistical rank fusion with heuristic boosts, and optional neural reranking. This design allows the system to leverage the broad coverage of commercial APIs while refining precision through modern machine learning techniques.
Parallel Search Engine Dispatch
Search requests enter the system through handleSearch in src/tools/search.ts, which validates inputs and forwards them to the configured provider.
// src/tools/search.ts
export async function handleSearch(
input: SearchInput,
engines: SearchEngine[],
router: SmartRouter,
…,
): Promise<StageResult<SearchOutput>> {
const provider = await getSearchProvider(); // core provider or legacy SearXNG
return provider.search(input, { engines, router, … });
}
The provider delegates to runV1Search inside src/search/core/orchestrator.ts, where the system executes all enabled engines concurrently using runEnginesParallel. Each engine operates under strict timing constraints governed by constants such as ENGINE_POOL_SOFT_DEADLINE_MS and ENGINE_POOL_CHRONIC_SOFT_DEADLINE_MS, ensuring that slow providers do not block the overall result set.
Reciprocal Rank Fusion (RRF) and Score Calculation
Once raw results return from the engine pool, the orchestrator deduplicates them per-engine and applies Reciprocal Rank Fusion (RRF) inside the scoreOutcomes function. The implementation uses a fixed damping constant RRF_K = 60 and resolves per-engine weights through resolveEngineWeight in src/search/core/engine-quality.ts.
// src/search/core/orchestrator.ts – scoreOutcomes()
const base = weight / (RRF_K + rank);
fused.set(key, (fused.get(key) ?? 0) + base * recMul);
The fusion formula base = weight / (RRF_K + rank) converts each engine’s ordinal rank into a relevance score, summing contributions across all engines to produce a preliminary relevance_score for every unique URL. After fusion, the pipeline applies additional heuristic boosts including authority scoring (applyAuthorityBoost), recency multipliers (recencyMultiplier), rare-term handling, and brand-collision guards (applyBrandCollisionGuard).
ML-Based Reranking Pipeline
If the request configuration specifies a reranker other than 'none', the orchestrator invokes rerankResults from src/search/rerank.ts. This module dynamically loads a TransformersRerankProvider from src/search/reranker/transformers-rerank-provider.ts, which initializes a lightweight transformer model—typically a bi-encoder that scores query-result pairs.
// src/search/rerank.ts (excerpt)
const provider = await getRerankProvider(); // loads the ML model
const reranked = await provider.rerank(query, results);
// reranked results replace the order produced by RRF
The ML reranker reorders the fused list based on semantic similarity rather than positional rank. After reranking, the system normalizes final relevance_score values to a 0–1 scale unless the engine pool has degraded below the RANK_DEGRADED_CONFIDENCE_FLOOR of 0.05. Each result includes an evidence_score object detailing the contribution of RRF base scores, domain quality, lexical alignment, and recency factors.
Implementation Example
The following TypeScript SDK example demonstrates both the standard RRF fusion path and the optional ML reranking stage:
import { createClient } from '@wigolo/sdk';
const client = createClient({ apiKey: process.env.WIGOLO_API_KEY });
async function demo() {
// 1️⃣ Simple multi‑engine search – uses RRF fusion only
const res1 = await client.search({ query: 'typescript async await' });
console.log('RRF‑fused results:', res1.results.map(r => r.title));
// 2️⃣ Enable the ML reranker (requires the model to be warmed‑up)
const res2 = await client.search({
query: 'typescript async await',
reranker: 'onnx', // or 'transformers' depending on config
});
console.log('ML‑reranked results:', res2.results.map(r => r.title));
}
demo();
For custom integrations, you can invoke the core orchestrator directly with explicit engine selection:
import { SearchEngine } from '@wigolo/sdk/types';
import { handleSearch } from './src/tools/search.js';
import { SmartRouter } from './src/fetch/router.js';
const router = new SmartRouter(); // manages proxies & TLS tiers
const engines: SearchEngine[] = ['bing', 'duckduckgo']; // pick any subset
const out = await handleSearch(
{ query: 'open source licenses' },
engines,
router,
);
console.log(out.results.map(r => `${r.title} – ${r.url}`));
Summary
- Parallel dispatch in
src/tools/search.tsandsrc/search/core/orchestrator.tsruns multiple search engines concurrently with configurable timeouts. - RRF fusion combines per-engine rankings using the formula
weight / (60 + rank), producing a statistically robust baseline score. - Engine-specific weights are resolved via
resolveEngineWeightinsrc/search/core/engine-quality.ts, allowing quality-based calibration. - Heuristic boosts for authority, recency, and brand collision refine the fused scores before final output.
- Optional ML reranking via
TransformersRerankProviderinsrc/search/reranker/transformers-rerank-provider.tsreorders results using transformer-based semantic scoring when enabled.
Frequently Asked Questions
What is Reciprocal Rank Fusion (RRF) and why does Wigolo use it?
Reciprocal Rank Fusion is a rank aggregation algorithm that converts ordinal positions from multiple ordered lists into comparable scores using the formula 1 / (k + rank), where k is a constant (60 in Wigolo). Wigolo uses RRF because it requires no training data, handles missing results gracefully, and allows weighted contributions from engines with varying reliability through the weight parameter in src/search/core/engine-quality.ts.
How does the ML reranker integrate with the RRF phase?
The ML reranker operates as a post-processing step after RRF fusion completes. After scoreOutcomes calculates fused scores and applies heuristic boosts, the orchestrator checks for an active reranker configuration. If enabled, rerankResults in src/search/rerank.ts passes the fused list to TransformersRerankProvider, which scores each query-result pair and returns a completely reordered array that replaces the RRF-based ordering.
Can I disable ML reranking and rely solely on RRF fusion?
Yes, ML reranking is optional and controlled via the reranker configuration parameter. Setting reranker: 'none' skips the rerankResults call entirely, causing the system to return results ordered strictly by the weighted RRF scores and heuristic boosts. This mode reduces latency and eliminates model warmup requirements while maintaining high-quality aggregation through the statistical fusion layer.
What happens when search engines return conflicting or duplicate results?
Wigolo deduplicates results per-engine before fusion and uses URL-based keying during RRF scoring. In src/search/core/orchestrator.ts, the fusion logic aggregates scores under a unique key derived from each result’s URL, effectively merging duplicate links across engines. The evidence_score attached to each result traces which engines contributed to the final rank, allowing downstream consumers to audit conflicts.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →