How Wigolo's On-Device Embeddings and Reranker Avoid API Costs While Maintaining Quality
Wigolo eliminates recurring API fees by running the BGE-small-en-v1.5 embedding model and bge-reranker-v2-m3 cross-encoder locally using ONNX Runtime, downloading models once to ~/.wigolo/ and executing inference entirely on-device.
Wigolo, an open-source search tool from KnockOutEZ/wigolo, implements a zero-cost semantic search architecture by shipping native ONNX models for both embeddings and reranking. Unlike solutions that rely on OpenAI, Cohere, or other cloud APIs for every query, Wigolo's on-device embeddings and reranker perform all inference locally after a one-time model download. This approach removes per-token pricing while maintaining high-quality relevance scoring through cross-encoder reranking and intelligent post-processing heuristics.
On-Device Embeddings Architecture
FastembedEmbedProvider and fastembed-rs
At the core of Wigolo's embedding system is the FastembedEmbedProvider class in [src/embedding/fastembed-provider.ts](https://github.com/KnockOutEZ/wigolo/blob/main/src/embedding/fastembed-provider.ts). This provider wraps the fastembed-rs Rust bindings, which load the BGE-small-en-v1.5 model into memory and execute inference using the ONNX runtime. Because the model runs locally within the Node.js process via native bindings, no network requests are made to external AI services during embedding generation.
The provider handles cache-directory management automatically, storing model artifacts in ~/.wigolo/fastembed as verified in [src/cli/tui/verify.ts](https://github.com/KnockOutEZ/wigolo/blob/main/src/cli/tui/verify.ts). Subsequent application restarts reuse these cached weights, eliminating redundant downloads.
Lazy Initialization Pattern
Wigolo implements a lazy-on-first-use pattern to optimize startup performance. The embedInit function in [src/embedding/embed.ts](https://github.com/KnockOutEZ/wigolo/blob/main/src/embedding/embed.ts) initializes the embedding subsystem without loading the actual model into memory. The embedAsync method only triggers the download and model loading when the first embedding request occurs.
This architecture ensures that fresh installations incur no download cost or memory overhead until the user explicitly requests semantic search capabilities.
Provider Factory Abstraction
The getEmbedProvider factory in [src/providers/embed-provider.ts](https://github.com/KnockOutEZ/wigolo/blob/main/src/providers/embed-provider.ts) lazily imports the FastembedEmbedProvider module and returns a ready-or-lazy object. This abstraction allows the rest of the codebase to interact with a unified interface while the provider handles the complexities of native module loading and model lifecycle management.
Local Cross-Encoder Reranking
TransformersRerankProvider Implementation
For result ranking, Wigolo uses the TransformersRerankProvider located in [src/search/reranker/transformers-rerank-provider.ts](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/reranker/transformers-rerank-provider.ts). This provider loads the bge-reranker-v2-m3 cross-encoder model locally using the @huggingface/transformers library with the same ONNX runtime utilized by the embedding system.
The provider scores each (title, snippet, URL) tuple using the model's logits, producing fine-grained relevance scores that are significantly more discriminative than cosine similarity between embeddings alone.
Conditional Activation and Configuration
The rerankResults function in [src/search/rerank.ts](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/rerank.ts) checks the config.reranker flag before invoking the provider. When set to "onnx" (the default), the local reranker processes results; when set to "none", the pipeline passes results through unchanged.
This conditional logic makes it trivial to disable the reranker entirely for zero-cost operation on resource-constrained devices.
Warm-up and Failure Safety
During the CLI startup routine in [src/cli/warmup.ts](https://github.com/KnockOutEZ/wigolo/blob/main/src/cli/warmup.ts), the warmup function pre-loads both the embedding and reranker models. However, any failure during model initialization is caught and logged without aborting the process. This ensures the system never blocks users for remote service availability and continues with the best-available result ordering.
Quality Preservation Without Cloud APIs
Cross-Encoder Scoring vs. Embedding Similarity
While embedding similarity provides coarse semantic matching, the cross-encoder reranker analyzes the full interaction between query and candidate text. The TransformersRerankProvider computes logits that capture nuanced relevance signals, preserving top-rank quality comparable to cloud-based reranking APIs.
Post-Reranking Heuristics
After the reranker scores results, Wigolo applies additional lightweight heuristics entirely on-device:
- Recency boosting ([
src/search/reranker/recency.js](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/reranker/recency.js)) prioritizes newer content - Authority boosting ([
src/search/reranker/authority-boost.js](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/reranker/authority-boost.js)) weights sources by reliability metrics - Consensus scoring validates results across multiple retrieval signals
These layers improve relevance without external API calls.
Score Floor Protection
The score-floor module in [src/search/core/score-floor.ts](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/core/score-floor.ts) implements a safety threshold ensuring that even if the reranker produces very low logits for a query, the final ranking never collapses to junk results. This maintains sensible user-facing output regardless of model confidence.
Practical Implementation Examples
Generating Embeddings Locally
import { getEmbeddingService } from './embedding/embed.js';
// First call triggers lazy download if model is absent
const service = getEmbeddingService();
const vectors = await service.embedAsync([
'How does fastembed compare to sentence-transformers',
'Wigolo local search architecture'
]);
// vectors contain 384-dimensional embeddings computed locally
The embedAsync method internally calls FastembedEmbedProvider.embed, executing the ONNX model without network requests.
Executing Searches with Reranking
import { search } from './search/search.js';
// Defaults to on-device reranker (onnx) and embeddings (fastembed)
const results = await search('rust onnx runtime performance');
// Pipeline execution:
// 1. Fetch raw engine results
// 2. Compute query embedding via embedAsync
// 3. Invoke rerankResults with TransformersRerankProvider
// 4. Apply recency/authority boosts
console.log(results.map(r => ({
title: r.title,
finalScore: r.score
})));
Disabling the Reranker for Zero-Cost Operation
import { setConfig } from './config.js';
// Disable local reranker entirely
setConfig({ reranker: 'none' });
// Search executes with embedding similarity only
// Zero ONNX inference cost, minimal memory footprint
Summary
- Native ONNX Runtime: Wigolo uses
fastembed-rsand@huggingface/transformerswith ONNX to run BGE-small-en-v1.5 and bge-reranker-v2-m3 locally, eliminating per-query API costs. - Lazy Loading: Models download only on first use (
embedAsync) and cache permanently in~/.wigolo/, ensuring subsequent runs require no network access. - Conditional Execution: The
config.rerankerflag allows users to disable theTransformersRerankProviderentirely for resource-constrained environments. - Quality Maintenance: Cross-encoder logits provide fine-grained relevance scoring, supplemented by recency, authority, and score-floor heuristics to maintain result quality.
- Failure Resilience: Warm-up routines in
src/cli/warmup.tshandle model loading errors gracefully, ensuring the system remains operational even if model initialization fails.
Frequently Asked Questions
What specific models does Wigolo use for on-device embeddings and reranking?
Wigolo uses the BGE-small-en-v1.5 model for embeddings via the FastembedEmbedProvider, and the bge-reranker-v2-m3 cross-encoder for reranking via the TransformersRerankProvider. Both models run locally using ONNX Runtime through Rust and JavaScript bindings, as implemented in src/embedding/fastembed-provider.ts and src/search/reranker/transformers-rerank-provider.ts.
How does Wigolo handle the initial model download without blocking the user?
The system implements a lazy initialization pattern where embedInit prepares the service without loading models. The actual download and model loading occur only when embedAsync receives the first embedding request. Additionally, the CLI warmup command in src/cli/warmup.ts allows eager downloading during installation, while failures during warm-up are caught and logged without crashing the application.
Can I run Wigolo without the reranker to reduce resource usage?
Yes. By setting config.reranker to "none" using setConfig(), you disable the TransformersRerankProvider entirely. The rerankResults function in src/search/rerank.ts checks this flag and passes results through unchanged when disabled, reducing memory and compute requirements to near zero while retaining embedding-based search capabilities.
How does local inference quality compare to cloud-based embedding APIs?
The cross-encoder reranker provides more discriminative scoring than simple embedding similarity by analyzing query-document interactions through the bge-reranker-v2-m3 model's logits. Combined with post-processing heuristics (recency, authority, and score-floor logic), Wigolo achieves relevance quality comparable to cloud APIs like OpenAI or Cohere, while avoiding latency and privacy concerns associated with external network calls.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →