Architecture of Wigolo's On-Device ML Components for Embeddings and Reranking
Wigolo implements a modular, privacy-first RAG stack using Fastembed (Rust ONNX) for 384-dimensional BGE embeddings and Transformers.js for cross-encoder reranking, with both services lazy-loaded from ~/.wigolo/ caches and exposed via stable TypeScript provider interfaces.
The KnockOutEZ/wigolo repository ships a fully local retrieval-augmented generation (RAG) pipeline that runs entirely on-device without external API calls. At the heart of this wigolo on-device ML architecture are two high-performance ML services: an embedding provider for dense vector generation and a reranking provider for relevance scoring, both engineered for minimal footprint and maximum privacy.
Embedding Provider Architecture
Provider Interface and Contract
The embedding subsystem adheres to a strict EmbedProvider interface defined in src/providers/embed-provider.ts:
export interface EmbedProvider {
embed(texts: string[]): Promise<Float32Array[]>;
readonly dim: number;
readonly modelId: string;
}
This contract guarantees that any implementation—whether ONNX, TensorFlow.js, or native—returns standardized Float32Array vectors and exposes metadata about the model dimensions.
Fastembed Implementation with Rust ONNX
The concrete implementation resides in src/embedding/fastembed-provider.ts and leverages the fastembed NPM package, which wraps a Rust-based ONNX runtime.
Key architectural decisions:
- Lazy Loading: The heavy native dependency is imported dynamically via
import('fastembed')only whenembed()is first invoked, keeping startup times minimal. - Automatic Cache Management: The provider downloads
BGE-small-en-v1.5(384-dimensional) into~/.wigolo/fastembedand implements archive integrity checks viainitModelWithArchiveRetry(). If aTAR_BAD_ARCHIVEerror is detected, the corrupt cache is wiped and re-downloaded automatically. - Warm-up Pattern: The
FastembedEmbedProvider.warmup()method preloads the model into memory, ensuring subsequent embedding calls execute without initialization latency.
Singleton Factory Pattern
The factory in src/providers/embed-provider.ts manages provider lifecycle:
let cached: Promise<EmbedProvider> | null = null;
export function getEmbedProvider(): Promise<EmbedProvider> {
if (cached) return cached;
cached = import('../embedding/fastembed-provider.js')
.then(m => {
const p = new m.FastembedEmbedProvider();
return p.warmup().then(() => {
log.info('embed provider ready', {
provider: 'embed',
impl: 'fastembed',
modelId: p.modelId,
dim: p.dim,
});
return p;
});
})
.catch(err => {
cached = null;
throw err;
});
return cached;
}
This singleton ensures that the model is loaded exactly once per process and shared across all embedding requests.
Reranking Provider Architecture
Provider Interface
The reranking service implements the RerankProvider interface from src/providers/rerank-provider.ts:
export interface RerankProvider {
rerank(
query: string,
candidates: RerankCandidate[],
topK?: number,
): Promise<RerankResult[]>;
readonly modelId: string;
}
Transformers.js Cross-Encoder Implementation
Located in src/search/reranker/transformers-rerank-provider.ts, this provider loads the ms-marco-MiniLM-L-6-v2 cross-encoder using Transformers.js for in-process ONNX execution.
Critical implementation details:
- In-Process Execution: Both
AutoTokenizer.from_pretrainedandAutoModelForSequenceClassification.from_pretrainedrun inside the Node.js process via ONNX, eliminating Python runtime dependencies. - Network Resilience: The factory wraps initialization with
withFetchRetry()to handle transient network failures during model download, andwrapLoadError()translates obscure HuggingFace fetch errors into actionable messages suggestingwigolo warmup. - Resource Management: The
dispose()method explicitly releases the underlying ONNX session, preventing native-thread teardown crashes during process shutdown. - Scoring Logic: The cross-encoder returns single-logit regression scores representing relevance; results are sorted descending and truncated to
topK.
Reranker Factory with Retry Logic
The initialization pattern in src/providers/rerank-provider.ts mirrors the embedding factory but adds fetch resilience:
let cached: Promise<RerankProvider> | null = null;
export function getRerankProvider(): Promise<RerankProvider> {
if (cached) return cached;
cached = import('../search/reranker/transformers-rerank-provider.js')
.then(m => {
const p = new m.TransformersRerankProvider();
return withFetchRetry(() => p.warmup()).then(() => {
log.info('rerank provider ready', {
provider: 'rerank',
impl: 'transformers',
modelId: p.modelId,
});
return p;
});
})
.catch(err => {
cached = null;
throw err;
});
return cached;
}
Vector Store Integration
Embedding vectors persist in SQLite-Vec, a local vector extension for SQLite that enables approximate nearest-neighbor search without external databases.
The VectorStore abstraction in src/providers/vector-store.ts defines the contract, while src/cache/sqlite-vec-store.js provides the concrete implementation. The store is instantiated lazily via getVectorStore() and cached for the process lifetime, acting as the bridge between the embedding provider (which generates vectors) and the search tools (which query them).
Runtime Flow and Lifecycle Management
The wigolo on-device ML architecture follows a consistent lifecycle pattern:
- Optional Warm-up: Running
wigolo warmuptriggers bothgetEmbedProvider()andgetRerankProvider(), downloading models preemptively to~/.wigolo/subdirectories. - Document Indexing: When content is saved, the system calls
embed()to generate 384-dimensional vectors, then upserts them intoSqliteVecStorewith metadata including the model ID and content hash. - Hybrid Retrieval: Queries execute against both FTS5 (full-text) and the vector store (if
embedding_availableis true), merging results into a unified candidate list. - Reranking Stage: When
includeRerankis enabled, the system callsrerank()to score (query, document) pairs using the cross-encoder, returning a precision-ordered top-k list.
Practical Implementation Examples
Generating Embeddings
import { getEmbedProvider } from 'wigolo/providers/embed-provider.js';
async function embedTexts(texts: string[]) {
const provider = await getEmbedProvider(); // lazy init + warmup
const vectors = await provider.embed(texts); // Float32Array per text
console.log(`Got ${vectors.length} vectors, dim=${provider.dim}`);
return vectors;
}
Indexing Documents with Metadata
import { getVectorStore } from 'wigolo/providers/vector-store.js';
import { getEmbedProvider } from 'wigolo/providers/embed-provider.js';
async function indexPage(url: string, content: string) {
const embed = await getEmbedProvider();
const [vector] = await embed.embed([content]);
const store = await getVectorStore();
await store.upsert([{
id: url,
vector,
metadata: {
url,
contentHash: await hash(content),
modelId: embed.modelId
},
}]);
}
Reranking Search Results
import { getRerankProvider } from 'wigolo/providers/rerank-provider.js';
interface Candidate { id: string; text: string; }
async function rerank(query: string, candidates: Candidate[]) {
const provider = await getRerankProvider();
const ranked = await provider.rerank(query, candidates, 5);
console.log('Top result:', ranked[0]);
return ranked;
}
Summary
- Dual-Provider Design: Wigolo separates embedding (Fastembed/Rust ONNX) and reranking (Transformers.js) concerns behind stable TypeScript interfaces defined in
src/providers/. - Lazy-Loaded Efficiency: Both ML services use dynamic imports and singleton factories to avoid startup overhead, with explicit
warmup()methods for preloading. - Resilient Caching: Models cache to
~/.wigolo/fastembedand~/.wigolo/transformerswith automatic corruption detection and retry logic. - Local-First Storage: Vectors persist in SQLite-Vec via the
VectorStoreabstraction, enabling hybrid FTS5 + semantic search without network dependencies. - Resource Safety: The reranker implements explicit
dispose()patterns to prevent native memory leaks during shutdown.
Frequently Asked Questions
What models does Wigolo use for embeddings and reranking?
Wigolo uses BGE-small-en-v1.5 (384 dimensions) for embeddings via the Fastembed provider, and ms-marco-MiniLM-L-6-v2 for reranking via Transformers.js. Both are downloaded automatically on first use to ~/.wigolo/fastembed and ~/.wigolo/transformers respectively.
How does Wigolo handle model loading failures?
The architecture implements multiple resilience strategies. The embedding provider detects TAR_BAD_ARCHIVE errors and automatically wipes corrupt caches before retrying. The rerank provider wraps initialization with withFetchRetry() to handle transient network blips, and both factories clear their singleton cache on failure to allow fresh attempts.
Can I use Wigolo's ML components without internet access?
Yes, after the initial model download. Running wigolo warmup while online caches both the embedding and reranking models locally. Subsequent operations run entirely offline using the cached ONNX files in ~/.wigolo/, making the stack fully air-gappable for privacy-sensitive environments.
Why does Wigolo use different ONNX runtimes for embedding versus reranking?
The embedding provider uses Fastembed (Rust-based ONNX) for its superior performance with the small BGE model and minimal memory footprint. The reranker uses Transformers.js (JavaScript ONNX runtime) for its seamless integration with HuggingFace tokenizers and sequence classification APIs. Both run in-process without Python dependencies, but each leverages the optimal runtime for its specific workload characteristics.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →