Architecture of Wigolo's On-Device ML Components for Embeddings and Reranking

Wigolo implements a modular, privacy-first RAG stack using Fastembed (Rust ONNX) for 384-dimensional BGE embeddings and Transformers.js for cross-encoder reranking, with both services lazy-loaded from ~/.wigolo/ caches and exposed via stable TypeScript provider interfaces.

The KnockOutEZ/wigolo repository ships a fully local retrieval-augmented generation (RAG) pipeline that runs entirely on-device without external API calls. At the heart of this wigolo on-device ML architecture are two high-performance ML services: an embedding provider for dense vector generation and a reranking provider for relevance scoring, both engineered for minimal footprint and maximum privacy.

Embedding Provider Architecture

Provider Interface and Contract

The embedding subsystem adheres to a strict EmbedProvider interface defined in src/providers/embed-provider.ts:

export interface EmbedProvider {
  embed(texts: string[]): Promise<Float32Array[]>;
  readonly dim: number;
  readonly modelId: string;
}

This contract guarantees that any implementation—whether ONNX, TensorFlow.js, or native—returns standardized Float32Array vectors and exposes metadata about the model dimensions.

Fastembed Implementation with Rust ONNX

The concrete implementation resides in src/embedding/fastembed-provider.ts and leverages the fastembed NPM package, which wraps a Rust-based ONNX runtime.

Key architectural decisions:

  • Lazy Loading: The heavy native dependency is imported dynamically via import('fastembed') only when embed() is first invoked, keeping startup times minimal.
  • Automatic Cache Management: The provider downloads BGE-small-en-v1.5 (384-dimensional) into ~/.wigolo/fastembed and implements archive integrity checks via initModelWithArchiveRetry(). If a TAR_BAD_ARCHIVE error is detected, the corrupt cache is wiped and re-downloaded automatically.
  • Warm-up Pattern: The FastembedEmbedProvider.warmup() method preloads the model into memory, ensuring subsequent embedding calls execute without initialization latency.

Singleton Factory Pattern

The factory in src/providers/embed-provider.ts manages provider lifecycle:

let cached: Promise<EmbedProvider> | null = null;

export function getEmbedProvider(): Promise<EmbedProvider> {
  if (cached) return cached;
  cached = import('../embedding/fastembed-provider.js')
    .then(m => {
      const p = new m.FastembedEmbedProvider();
      return p.warmup().then(() => {
        log.info('embed provider ready', {
          provider: 'embed',
          impl: 'fastembed',
          modelId: p.modelId,
          dim: p.dim,
        });
        return p;
      });
    })
    .catch(err => {
      cached = null;
      throw err;
    });
  return cached;
}

This singleton ensures that the model is loaded exactly once per process and shared across all embedding requests.

Reranking Provider Architecture

Provider Interface

The reranking service implements the RerankProvider interface from src/providers/rerank-provider.ts:

export interface RerankProvider {
  rerank(
    query: string,
    candidates: RerankCandidate[],
    topK?: number,
  ): Promise<RerankResult[]>;
  readonly modelId: string;
}

Transformers.js Cross-Encoder Implementation

Located in src/search/reranker/transformers-rerank-provider.ts, this provider loads the ms-marco-MiniLM-L-6-v2 cross-encoder using Transformers.js for in-process ONNX execution.

Critical implementation details:

  • In-Process Execution: Both AutoTokenizer.from_pretrained and AutoModelForSequenceClassification.from_pretrained run inside the Node.js process via ONNX, eliminating Python runtime dependencies.
  • Network Resilience: The factory wraps initialization with withFetchRetry() to handle transient network failures during model download, and wrapLoadError() translates obscure HuggingFace fetch errors into actionable messages suggesting wigolo warmup.
  • Resource Management: The dispose() method explicitly releases the underlying ONNX session, preventing native-thread teardown crashes during process shutdown.
  • Scoring Logic: The cross-encoder returns single-logit regression scores representing relevance; results are sorted descending and truncated to topK.

Reranker Factory with Retry Logic

The initialization pattern in src/providers/rerank-provider.ts mirrors the embedding factory but adds fetch resilience:

let cached: Promise<RerankProvider> | null = null;

export function getRerankProvider(): Promise<RerankProvider> {
  if (cached) return cached;
  cached = import('../search/reranker/transformers-rerank-provider.js')
    .then(m => {
      const p = new m.TransformersRerankProvider();
      return withFetchRetry(() => p.warmup()).then(() => {
        log.info('rerank provider ready', {
          provider: 'rerank',
          impl: 'transformers',
          modelId: p.modelId,
        });
        return p;
      });
    })
    .catch(err => {
      cached = null;
      throw err;
    });
  return cached;
}

Vector Store Integration

Embedding vectors persist in SQLite-Vec, a local vector extension for SQLite that enables approximate nearest-neighbor search without external databases.

The VectorStore abstraction in src/providers/vector-store.ts defines the contract, while src/cache/sqlite-vec-store.js provides the concrete implementation. The store is instantiated lazily via getVectorStore() and cached for the process lifetime, acting as the bridge between the embedding provider (which generates vectors) and the search tools (which query them).

Runtime Flow and Lifecycle Management

The wigolo on-device ML architecture follows a consistent lifecycle pattern:

  1. Optional Warm-up: Running wigolo warmup triggers both getEmbedProvider() and getRerankProvider(), downloading models preemptively to ~/.wigolo/ subdirectories.
  2. Document Indexing: When content is saved, the system calls embed() to generate 384-dimensional vectors, then upserts them into SqliteVecStore with metadata including the model ID and content hash.
  3. Hybrid Retrieval: Queries execute against both FTS5 (full-text) and the vector store (if embedding_available is true), merging results into a unified candidate list.
  4. Reranking Stage: When includeRerank is enabled, the system calls rerank() to score (query, document) pairs using the cross-encoder, returning a precision-ordered top-k list.

Practical Implementation Examples

Generating Embeddings

import { getEmbedProvider } from 'wigolo/providers/embed-provider.js';

async function embedTexts(texts: string[]) {
  const provider = await getEmbedProvider();   // lazy init + warmup
  const vectors = await provider.embed(texts); // Float32Array per text
  console.log(`Got ${vectors.length} vectors, dim=${provider.dim}`);
  return vectors;
}

Indexing Documents with Metadata

import { getVectorStore } from 'wigolo/providers/vector-store.js';
import { getEmbedProvider } from 'wigolo/providers/embed-provider.js';

async function indexPage(url: string, content: string) {
  const embed = await getEmbedProvider();
  const [vector] = await embed.embed([content]);
  const store = await getVectorStore();
  await store.upsert([{
    id: url,
    vector,
    metadata: { 
      url, 
      contentHash: await hash(content), 
      modelId: embed.modelId 
    },
  }]);
}

Reranking Search Results

import { getRerankProvider } from 'wigolo/providers/rerank-provider.js';

interface Candidate { id: string; text: string; }

async function rerank(query: string, candidates: Candidate[]) {
  const provider = await getRerankProvider();
  const ranked = await provider.rerank(query, candidates, 5);
  console.log('Top result:', ranked[0]);
  return ranked;
}

Summary

  • Dual-Provider Design: Wigolo separates embedding (Fastembed/Rust ONNX) and reranking (Transformers.js) concerns behind stable TypeScript interfaces defined in src/providers/.
  • Lazy-Loaded Efficiency: Both ML services use dynamic imports and singleton factories to avoid startup overhead, with explicit warmup() methods for preloading.
  • Resilient Caching: Models cache to ~/.wigolo/fastembed and ~/.wigolo/transformers with automatic corruption detection and retry logic.
  • Local-First Storage: Vectors persist in SQLite-Vec via the VectorStore abstraction, enabling hybrid FTS5 + semantic search without network dependencies.
  • Resource Safety: The reranker implements explicit dispose() patterns to prevent native memory leaks during shutdown.

Frequently Asked Questions

What models does Wigolo use for embeddings and reranking?

Wigolo uses BGE-small-en-v1.5 (384 dimensions) for embeddings via the Fastembed provider, and ms-marco-MiniLM-L-6-v2 for reranking via Transformers.js. Both are downloaded automatically on first use to ~/.wigolo/fastembed and ~/.wigolo/transformers respectively.

How does Wigolo handle model loading failures?

The architecture implements multiple resilience strategies. The embedding provider detects TAR_BAD_ARCHIVE errors and automatically wipes corrupt caches before retrying. The rerank provider wraps initialization with withFetchRetry() to handle transient network blips, and both factories clear their singleton cache on failure to allow fresh attempts.

Can I use Wigolo's ML components without internet access?

Yes, after the initial model download. Running wigolo warmup while online caches both the embedding and reranking models locally. Subsequent operations run entirely offline using the cached ONNX files in ~/.wigolo/, making the stack fully air-gappable for privacy-sensitive environments.

Why does Wigolo use different ONNX runtimes for embedding versus reranking?

The embedding provider uses Fastembed (Rust-based ONNX) for its superior performance with the small BGE model and minimal memory footprint. The reranker uses Transformers.js (JavaScript ONNX runtime) for its seamless integration with HuggingFace tokenizers and sequence classification APIs. Both run in-process without Python dependencies, but each leverages the optimal runtime for its specific workload characteristics.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →