How Wigolo's On-Device Embeddings and Reranker Avoid API Costs While Maintaining Quality

Wigolo eliminates recurring API fees by running the BGE-small-en-v1.5 embedding model and bge-reranker-v2-m3 cross-encoder locally using ONNX Runtime, downloading models once to ~/.wigolo/ and executing inference entirely on-device.

Wigolo, an open-source search tool from KnockOutEZ/wigolo, implements a zero-cost semantic search architecture by shipping native ONNX models for both embeddings and reranking. Unlike solutions that rely on OpenAI, Cohere, or other cloud APIs for every query, Wigolo's on-device embeddings and reranker perform all inference locally after a one-time model download. This approach removes per-token pricing while maintaining high-quality relevance scoring through cross-encoder reranking and intelligent post-processing heuristics.

On-Device Embeddings Architecture

FastembedEmbedProvider and fastembed-rs

At the core of Wigolo's embedding system is the FastembedEmbedProvider class in [src/embedding/fastembed-provider.ts](https://github.com/KnockOutEZ/wigolo/blob/main/src/embedding/fastembed-provider.ts). This provider wraps the fastembed-rs Rust bindings, which load the BGE-small-en-v1.5 model into memory and execute inference using the ONNX runtime. Because the model runs locally within the Node.js process via native bindings, no network requests are made to external AI services during embedding generation.

The provider handles cache-directory management automatically, storing model artifacts in ~/.wigolo/fastembed as verified in [src/cli/tui/verify.ts](https://github.com/KnockOutEZ/wigolo/blob/main/src/cli/tui/verify.ts). Subsequent application restarts reuse these cached weights, eliminating redundant downloads.

Lazy Initialization Pattern

Wigolo implements a lazy-on-first-use pattern to optimize startup performance. The embedInit function in [src/embedding/embed.ts](https://github.com/KnockOutEZ/wigolo/blob/main/src/embedding/embed.ts) initializes the embedding subsystem without loading the actual model into memory. The embedAsync method only triggers the download and model loading when the first embedding request occurs.

This architecture ensures that fresh installations incur no download cost or memory overhead until the user explicitly requests semantic search capabilities.

Provider Factory Abstraction

The getEmbedProvider factory in [src/providers/embed-provider.ts](https://github.com/KnockOutEZ/wigolo/blob/main/src/providers/embed-provider.ts) lazily imports the FastembedEmbedProvider module and returns a ready-or-lazy object. This abstraction allows the rest of the codebase to interact with a unified interface while the provider handles the complexities of native module loading and model lifecycle management.

Local Cross-Encoder Reranking

TransformersRerankProvider Implementation

For result ranking, Wigolo uses the TransformersRerankProvider located in [src/search/reranker/transformers-rerank-provider.ts](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/reranker/transformers-rerank-provider.ts). This provider loads the bge-reranker-v2-m3 cross-encoder model locally using the @huggingface/transformers library with the same ONNX runtime utilized by the embedding system.

The provider scores each (title, snippet, URL) tuple using the model's logits, producing fine-grained relevance scores that are significantly more discriminative than cosine similarity between embeddings alone.

Conditional Activation and Configuration

The rerankResults function in [src/search/rerank.ts](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/rerank.ts) checks the config.reranker flag before invoking the provider. When set to "onnx" (the default), the local reranker processes results; when set to "none", the pipeline passes results through unchanged.

This conditional logic makes it trivial to disable the reranker entirely for zero-cost operation on resource-constrained devices.

Warm-up and Failure Safety

During the CLI startup routine in [src/cli/warmup.ts](https://github.com/KnockOutEZ/wigolo/blob/main/src/cli/warmup.ts), the warmup function pre-loads both the embedding and reranker models. However, any failure during model initialization is caught and logged without aborting the process. This ensures the system never blocks users for remote service availability and continues with the best-available result ordering.

Quality Preservation Without Cloud APIs

Cross-Encoder Scoring vs. Embedding Similarity

While embedding similarity provides coarse semantic matching, the cross-encoder reranker analyzes the full interaction between query and candidate text. The TransformersRerankProvider computes logits that capture nuanced relevance signals, preserving top-rank quality comparable to cloud-based reranking APIs.

Post-Reranking Heuristics

After the reranker scores results, Wigolo applies additional lightweight heuristics entirely on-device:

These layers improve relevance without external API calls.

Score Floor Protection

The score-floor module in [src/search/core/score-floor.ts](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/core/score-floor.ts) implements a safety threshold ensuring that even if the reranker produces very low logits for a query, the final ranking never collapses to junk results. This maintains sensible user-facing output regardless of model confidence.

Practical Implementation Examples

Generating Embeddings Locally

import { getEmbeddingService } from './embedding/embed.js';

// First call triggers lazy download if model is absent
const service = getEmbeddingService();
const vectors = await service.embedAsync([
  'How does fastembed compare to sentence-transformers',
  'Wigolo local search architecture'
]);

// vectors contain 384-dimensional embeddings computed locally

The embedAsync method internally calls FastembedEmbedProvider.embed, executing the ONNX model without network requests.

Executing Searches with Reranking

import { search } from './search/search.js';

// Defaults to on-device reranker (onnx) and embeddings (fastembed)
const results = await search('rust onnx runtime performance');

// Pipeline execution:
// 1. Fetch raw engine results
// 2. Compute query embedding via embedAsync
// 3. Invoke rerankResults with TransformersRerankProvider
// 4. Apply recency/authority boosts
console.log(results.map(r => ({ 
  title: r.title, 
  finalScore: r.score 
})));

Disabling the Reranker for Zero-Cost Operation

import { setConfig } from './config.js';

// Disable local reranker entirely
setConfig({ reranker: 'none' });

// Search executes with embedding similarity only
// Zero ONNX inference cost, minimal memory footprint

Summary

  • Native ONNX Runtime: Wigolo uses fastembed-rs and @huggingface/transformers with ONNX to run BGE-small-en-v1.5 and bge-reranker-v2-m3 locally, eliminating per-query API costs.
  • Lazy Loading: Models download only on first use (embedAsync) and cache permanently in ~/.wigolo/, ensuring subsequent runs require no network access.
  • Conditional Execution: The config.reranker flag allows users to disable the TransformersRerankProvider entirely for resource-constrained environments.
  • Quality Maintenance: Cross-encoder logits provide fine-grained relevance scoring, supplemented by recency, authority, and score-floor heuristics to maintain result quality.
  • Failure Resilience: Warm-up routines in src/cli/warmup.ts handle model loading errors gracefully, ensuring the system remains operational even if model initialization fails.

Frequently Asked Questions

What specific models does Wigolo use for on-device embeddings and reranking?

Wigolo uses the BGE-small-en-v1.5 model for embeddings via the FastembedEmbedProvider, and the bge-reranker-v2-m3 cross-encoder for reranking via the TransformersRerankProvider. Both models run locally using ONNX Runtime through Rust and JavaScript bindings, as implemented in src/embedding/fastembed-provider.ts and src/search/reranker/transformers-rerank-provider.ts.

How does Wigolo handle the initial model download without blocking the user?

The system implements a lazy initialization pattern where embedInit prepares the service without loading models. The actual download and model loading occur only when embedAsync receives the first embedding request. Additionally, the CLI warmup command in src/cli/warmup.ts allows eager downloading during installation, while failures during warm-up are caught and logged without crashing the application.

Can I run Wigolo without the reranker to reduce resource usage?

Yes. By setting config.reranker to "none" using setConfig(), you disable the TransformersRerankProvider entirely. The rerankResults function in src/search/rerank.ts checks this flag and passes results through unchanged when disabled, reducing memory and compute requirements to near zero while retaining embedding-based search capabilities.

How does local inference quality compare to cloud-based embedding APIs?

The cross-encoder reranker provides more discriminative scoring than simple embedding similarity by analyzing query-document interactions through the bge-reranker-v2-m3 model's logits. Combined with post-processing heuristics (recency, authority, and score-floor logic), Wigolo achieves relevance quality comparable to cloud APIs like OpenAI or Cohere, while avoiding latency and privacy concerns associated with external network calls.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →