# Architecture of Wigolo's On-Device ML Components for Embeddings and Reranking

> Explore Wigolo's on-device ML architecture for embeddings and reranking. Discover its modular, privacy-first RAG stack powered by Fastembed and Transformers.js.

- Repository: [Towhid Khan/wigolo](https://github.com/KnockOutEZ/wigolo)
- Tags: architecture
- Published: 2026-07-29

---

**Wigolo implements a modular, privacy-first RAG stack using Fastembed (Rust ONNX) for 384-dimensional BGE embeddings and Transformers.js for cross-encoder reranking, with both services lazy-loaded from `~/.wigolo/` caches and exposed via stable TypeScript provider interfaces.**

The [KnockOutEZ/wigolo](https://github.com/KnockOutEZ/wigolo) repository ships a fully local retrieval-augmented generation (RAG) pipeline that runs entirely on-device without external API calls. At the heart of this **wigolo on-device ML architecture** are two high-performance ML services: an embedding provider for dense vector generation and a reranking provider for relevance scoring, both engineered for minimal footprint and maximum privacy.

## Embedding Provider Architecture

### Provider Interface and Contract

The embedding subsystem adheres to a strict `EmbedProvider` interface defined in [`src/providers/embed-provider.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/providers/embed-provider.ts):

```ts
export interface EmbedProvider {
  embed(texts: string[]): Promise<Float32Array[]>;
  readonly dim: number;
  readonly modelId: string;
}

```

This contract guarantees that any implementation—whether ONNX, TensorFlow.js, or native—returns standardized `Float32Array` vectors and exposes metadata about the model dimensions.

### Fastembed Implementation with Rust ONNX

The concrete implementation resides in [`src/embedding/fastembed-provider.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/embedding/fastembed-provider.ts) and leverages the `fastembed` NPM package, which wraps a Rust-based ONNX runtime.

Key architectural decisions:

- **Lazy Loading**: The heavy native dependency is imported dynamically via `import('fastembed')` only when `embed()` is first invoked, keeping startup times minimal.
- **Automatic Cache Management**: The provider downloads `BGE-small-en-v1.5` (384-dimensional) into `~/.wigolo/fastembed` and implements archive integrity checks via `initModelWithArchiveRetry()`. If a `TAR_BAD_ARCHIVE` error is detected, the corrupt cache is wiped and re-downloaded automatically.
- **Warm-up Pattern**: The `FastembedEmbedProvider.warmup()` method preloads the model into memory, ensuring subsequent embedding calls execute without initialization latency.

### Singleton Factory Pattern

The factory in [`src/providers/embed-provider.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/providers/embed-provider.ts) manages provider lifecycle:

```ts
let cached: Promise<EmbedProvider> | null = null;

export function getEmbedProvider(): Promise<EmbedProvider> {
  if (cached) return cached;
  cached = import('../embedding/fastembed-provider.js')
    .then(m => {
      const p = new m.FastembedEmbedProvider();
      return p.warmup().then(() => {
        log.info('embed provider ready', {
          provider: 'embed',
          impl: 'fastembed',
          modelId: p.modelId,
          dim: p.dim,
        });
        return p;
      });
    })
    .catch(err => {
      cached = null;
      throw err;
    });
  return cached;
}

```

This singleton ensures that the model is loaded exactly once per process and shared across all embedding requests.

## Reranking Provider Architecture

### Provider Interface

The reranking service implements the `RerankProvider` interface from [`src/providers/rerank-provider.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/providers/rerank-provider.ts):

```ts
export interface RerankProvider {
  rerank(
    query: string,
    candidates: RerankCandidate[],
    topK?: number,
  ): Promise<RerankResult[]>;
  readonly modelId: string;
}

```

### Transformers.js Cross-Encoder Implementation

Located in [`src/search/reranker/transformers-rerank-provider.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/reranker/transformers-rerank-provider.ts), this provider loads the `ms-marco-MiniLM-L-6-v2` cross-encoder using [`Transformers.js`](https://github.com/KnockOutEZ/wigolo/blob/main/Transformers.js) for in-process ONNX execution.

Critical implementation details:

- **In-Process Execution**: Both `AutoTokenizer.from_pretrained` and `AutoModelForSequenceClassification.from_pretrained` run inside the Node.js process via ONNX, eliminating Python runtime dependencies.
- **Network Resilience**: The factory wraps initialization with `withFetchRetry()` to handle transient network failures during model download, and `wrapLoadError()` translates obscure HuggingFace fetch errors into actionable messages suggesting `wigolo warmup`.
- **Resource Management**: The `dispose()` method explicitly releases the underlying ONNX session, preventing native-thread teardown crashes during process shutdown.
- **Scoring Logic**: The cross-encoder returns single-logit regression scores representing relevance; results are sorted descending and truncated to `topK`.

### Reranker Factory with Retry Logic

The initialization pattern in [`src/providers/rerank-provider.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/providers/rerank-provider.ts) mirrors the embedding factory but adds fetch resilience:

```ts
let cached: Promise<RerankProvider> | null = null;

export function getRerankProvider(): Promise<RerankProvider> {
  if (cached) return cached;
  cached = import('../search/reranker/transformers-rerank-provider.js')
    .then(m => {
      const p = new m.TransformersRerankProvider();
      return withFetchRetry(() => p.warmup()).then(() => {
        log.info('rerank provider ready', {
          provider: 'rerank',
          impl: 'transformers',
          modelId: p.modelId,
        });
        return p;
      });
    })
    .catch(err => {
      cached = null;
      throw err;
    });
  return cached;
}

```

## Vector Store Integration

Embedding vectors persist in **SQLite-Vec**, a local vector extension for SQLite that enables approximate nearest-neighbor search without external databases.

The `VectorStore` abstraction in [`src/providers/vector-store.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/providers/vector-store.ts) defines the contract, while [`src/cache/sqlite-vec-store.js`](https://github.com/KnockOutEZ/wigolo/blob/main/src/cache/sqlite-vec-store.js) provides the concrete implementation. The store is instantiated lazily via `getVectorStore()` and cached for the process lifetime, acting as the bridge between the embedding provider (which generates vectors) and the search tools (which query them).

## Runtime Flow and Lifecycle Management

The **wigolo on-device ML architecture** follows a consistent lifecycle pattern:

1. **Optional Warm-up**: Running `wigolo warmup` triggers both `getEmbedProvider()` and `getRerankProvider()`, downloading models preemptively to `~/.wigolo/` subdirectories.
2. **Document Indexing**: When content is saved, the system calls `embed()` to generate 384-dimensional vectors, then upserts them into `SqliteVecStore` with metadata including the model ID and content hash.
3. **Hybrid Retrieval**: Queries execute against both FTS5 (full-text) and the vector store (if `embedding_available` is true), merging results into a unified candidate list.
4. **Reranking Stage**: When `includeRerank` is enabled, the system calls `rerank()` to score (query, document) pairs using the cross-encoder, returning a precision-ordered top-k list.

## Practical Implementation Examples

### Generating Embeddings

```ts
import { getEmbedProvider } from 'wigolo/providers/embed-provider.js';

async function embedTexts(texts: string[]) {
  const provider = await getEmbedProvider();   // lazy init + warmup
  const vectors = await provider.embed(texts); // Float32Array per text
  console.log(`Got ${vectors.length} vectors, dim=${provider.dim}`);
  return vectors;
}

```

### Indexing Documents with Metadata

```ts
import { getVectorStore } from 'wigolo/providers/vector-store.js';
import { getEmbedProvider } from 'wigolo/providers/embed-provider.js';

async function indexPage(url: string, content: string) {
  const embed = await getEmbedProvider();
  const [vector] = await embed.embed([content]);
  const store = await getVectorStore();
  await store.upsert([{
    id: url,
    vector,
    metadata: { 
      url, 
      contentHash: await hash(content), 
      modelId: embed.modelId 
    },
  }]);
}

```

### Reranking Search Results

```ts
import { getRerankProvider } from 'wigolo/providers/rerank-provider.js';

interface Candidate { id: string; text: string; }

async function rerank(query: string, candidates: Candidate[]) {
  const provider = await getRerankProvider();
  const ranked = await provider.rerank(query, candidates, 5);
  console.log('Top result:', ranked[0]);
  return ranked;
}

```

## Summary

- **Dual-Provider Design**: Wigolo separates embedding (Fastembed/Rust ONNX) and reranking (Transformers.js) concerns behind stable TypeScript interfaces defined in `src/providers/`.
- **Lazy-Loaded Efficiency**: Both ML services use dynamic imports and singleton factories to avoid startup overhead, with explicit `warmup()` methods for preloading.
- **Resilient Caching**: Models cache to `~/.wigolo/fastembed` and `~/.wigolo/transformers` with automatic corruption detection and retry logic.
- **Local-First Storage**: Vectors persist in SQLite-Vec via the `VectorStore` abstraction, enabling hybrid FTS5 + semantic search without network dependencies.
- **Resource Safety**: The reranker implements explicit `dispose()` patterns to prevent native memory leaks during shutdown.

## Frequently Asked Questions

### What models does Wigolo use for embeddings and reranking?

Wigolo uses `BGE-small-en-v1.5` (384 dimensions) for embeddings via the Fastembed provider, and `ms-marco-MiniLM-L-6-v2` for reranking via Transformers.js. Both are downloaded automatically on first use to `~/.wigolo/fastembed` and `~/.wigolo/transformers` respectively.

### How does Wigolo handle model loading failures?

The architecture implements multiple resilience strategies. The embedding provider detects `TAR_BAD_ARCHIVE` errors and automatically wipes corrupt caches before retrying. The rerank provider wraps initialization with `withFetchRetry()` to handle transient network blips, and both factories clear their singleton cache on failure to allow fresh attempts.

### Can I use Wigolo's ML components without internet access?

Yes, after the initial model download. Running `wigolo warmup` while online caches both the embedding and reranking models locally. Subsequent operations run entirely offline using the cached ONNX files in `~/.wigolo/`, making the stack fully air-gappable for privacy-sensitive environments.

### Why does Wigolo use different ONNX runtimes for embedding versus reranking?

The embedding provider uses Fastembed (Rust-based ONNX) for its superior performance with the small BGE model and minimal memory footprint. The reranker uses Transformers.js (JavaScript ONNX runtime) for its seamless integration with HuggingFace tokenizers and sequence classification APIs. Both run in-process without Python dependencies, but each leverages the optimal runtime for its specific workload characteristics.