# How 5ire Executes the bge-m3 Embedding Model for Local Inference Without Subprocesses

> Discover how 5ire executes the bge-m3 embedding model in-process with Xenova Transformers and ONNX Runtime for faster local inference without subprocess overhead.

- Repository: [Ironben/5ire](https://github.com/nanbingxyz/5ire)
- Tags: how-to-guide
- Published: 2026-03-07

---

**The 5ire application runs the bge-m3 embedding model entirely in-process using the @xenova/transformers JavaScript library and ONNX Runtime, eliminating subprocess overhead by performing tensor calculations directly inside the Electron main process.**

The 5ire knowledge-base application implements local semantic search capabilities using the bge-m3 embedding model. Unlike traditional architectures that spawn separate Python processes for machine learning inference, 5ire keeps all embedding operations within the Node.js/Electron runtime. This approach leverages the JavaScript-native transformer pipeline to load and execute the ONNX model directly, ensuring low-latency vector generation without inter-process communication overhead.

## Why 5ire Avoids Subprocesses for bge-m3 Inference

Traditional local AI implementations often spawn Python subprocesses via `child_process.spawn` to execute PyTorch or TensorFlow models. In 5ire, the architecture deliberately avoids this pattern. A comprehensive search of the codebase reveals no active subprocess invocations for model inference—the only reference to child processes appears as a comment in [`src/main/setup.ts`](https://github.com/nanbingxyz/5ire/blob/main/src/main/setup.ts) at line 12, with no actual implementation following it.

By keeping the bge-m3 model execution inside the main Electron process, 5ire eliminates serialization overhead and reduces memory footprint. The application uses the **@xenova/transformers** library, which provides a JavaScript interface to transformer models converted to ONNX format. This library internally loads the native **onnxruntime-node** binary to perform hardware-accelerated tensor operations without leaving the Node.js environment.

## Architecture Overview: In-Process ONNX Execution

The execution flow follows a strict in-process pipeline that loads the model once and reuses it for multiple inference calls:

1. **Model Definition**: The bge-m3 model name and required file manifests are declared in [`src/main/constants.ts`](https://github.com/nanbingxyz/5ire/blob/main/src/main/constants.ts) (lines 1-18), specifying the ONNX model file, tokenizer configuration, and model config.
2. **Environment Setup**: The `Embedder.init()` method in [`src/main/services/embedder.ts`](https://github.com/nanbingxyz/5ire/blob/main/src/main/services/embedder.ts) (lines 97-104) configures the transformer environment to disallow remote models and points `env.localModelPath` to the local model storage directory.
3. **Pipeline Creation**: The `pipeline("feature-extraction", DOCUMENT_EMBEDDING_MODEL_NAME)` call instantiates a FeatureExtractionPipeline that loads the ONNX graph via `onnxruntime-node`.
4. **Inference**: The `embed()` method executes `extractor(text, { pooling: "mean", normalize: true })` synchronously (as a Promise) within the same process, returning dense vector representations.

## Model Configuration and Constants

All metadata required to locate and validate the bge-m3 model resides in [`src/main/constants.ts`](https://github.com/nanbingxyz/5ire/blob/main/src/main/constants.ts). This file defines the model identifier as `Xenova/bge-m3` and enumerates the specific files the downloader must retrieve: the quantized ONNX weights, tokenizer JSON, and model configuration.

The constants also specify the local storage path relative to the user's data directory. When packaged, the application expects these files to reside under `<user-data>/Embedding/Models/Xenova/bge-m3`, with the primary inference graph stored at `onnx/model_quantized.onnx`.

## Initializing the Embedding Pipeline

The `Embedder` class in [`src/main/services/embedder.ts`](https://github.com/nanbingxyz/5ire/blob/main/src/main/services/embedder.ts) manages the lifecycle of the embedding pipeline. During initialization, it strictly enforces local-only model loading to prevent automatic downloads from Hugging Face during inference.

```typescript
import { pipeline, env } from "@xenova/transformers";

async function initializeEmbedder(modelPath: string) {
  // Restrict to local files only
  env.allowRemoteModels = false;
  env.allowLocalModels = true;
  
  // Override the default local model path
  Object.defineProperty(env, "localModelPath", {
    value: modelPath,
    writable: false,
  });

  // Create the feature extraction pipeline for bge-m3
  const extractor = await pipeline(
    "feature-extraction", 
    "Xenova/bge-m3"
  );
  
  return extractor;
}

```

This initialization sequence ensures the pipeline loads the ONNX model from the local filesystem rather than attempting remote fetching. The [`rsbuild.config.ts`](https://github.com/nanbingxyz/5ire/blob/main/rsbuild.config.ts) file includes specific copy rules to bundle the `onnxruntime-node` native binaries with the application, making them available to the pipeline at runtime.

## Running Inference with the embed() Method

Once initialized, the embedder processes text inputs through the `embed()` method (lines 82-98 in [`embedder.ts`](https://github.com/nanbingxyz/5ire/blob/main/embedder.ts)). This method accepts an array of strings and returns normalized, pooled embedding vectors.

```typescript
async function generateEmbeddings(
  extractor: any, 
  texts: string[]
): Promise<number[][]> {
  // pooling: "mean" averages token embeddings into a single vector
  // normalize: true applies L2 normalization for cosine similarity
  const output = await extractor(texts, { 
    pooling: "mean", 
    normalize: true 
  });
  
  // Convert tensor to native JavaScript array
  return output.tolist();
}

// Usage example
const extractor = await initializeEmbedder("/path/to/Embedding/Models");
const vectors = await generateEmbeddings(extractor, [
  "Local inference architecture",
  "Subprocess vs in-process embedding"
]);

```

The inference executes immediately within the Electron main process. Because `onnxruntime-node` uses native bindings to execute the ONNX graph, the computation benefits from hardware acceleration where available, yet remains entirely within the Node.js memory space.

## Model Download and Storage Strategy

While inference happens in-process, the initial model acquisition uses the `Downloader` service in [`src/main/services/downloader.ts`](https://github.com/nanbingxyz/5ire/blob/main/src/main/services/downloader.ts). The `Embedder.downloadModel()` method triggers a one-time fetch of the bge-m3 files from Hugging Face, storing them in the user's data directory.

After download, the complete model resides at:

```

<user-data>/Embedding/Models/Xenova/bge-m3/
├── config.json
├── tokenizer.json
├── onnx/
│   └── model_quantized.onnx

```

Subsequent application launches load these files directly from disk, ensuring the embedding service remains available offline without requiring network calls or external process spawning.

## Summary

- **No subprocess overhead**: 5ire executes bge-m3 entirely within the Electron main process using `@xenova/transformers`, avoiding `child_process.spawn` or Python interpreters.
- **ONNX Runtime backbone**: The pipeline relies on `onnxruntime-node` native binaries to perform hardware-accelerated tensor operations locally.
- **Strict local loading**: `env.allowRemoteModels = false` ensures the embedder never fetches models during inference, loading exclusively from `<user-data>/Embedding/Models`.
- **Mean pooling and normalization**: The `embed()` method configures the pipeline with `pooling: "mean"` and `normalize: true` to produce standardized dense vectors suitable for vector database storage.
- **One-time download**: The `Downloader` service fetches model files once, after which all embedding operations use the cached ONNX graph.

## Frequently Asked Questions

### Does 5ire spawn a Python subprocess to run the bge-m3 model?

No. According to the source code analysis of `nanbingxyz/5ire`, the application contains no active `child_process.spawn` or similar APIs for model execution. The only reference to subprocesses appears as a comment in [`src/main/setup.ts`](https://github.com/nanbingxyz/5ire/blob/main/src/main/setup.ts) at line 12, with no implementation following it. Instead, 5ire uses the JavaScript-native `@xenova/transformers` library to execute the ONNX model directly within the Node.js process.

### What library handles the actual tensor computation for bge-m3?

The **@xenova/transformers** library provides the high-level API, while **onnxruntime-node** performs the low-level tensor computations. When the `pipeline()` function creates a FeatureExtractionPipeline, it loads the quantized ONNX model (`model_quantized.onnx`) using the native ONNX Runtime bindings bundled with the application via [`rsbuild.config.ts`](https://github.com/nanbingxyz/5ire/blob/main/rsbuild.config.ts) copy rules.

### Where are the bge-m3 model files stored locally?

The model files reside in the application's user data directory under `Embedding/Models/Xenova/bge-m3/`. This path is configured at runtime in [`src/main/services/embedder.ts`](https://github.com/nanbingxyz/5ire/blob/main/src/main/services/embedder.ts) by setting `env.localModelPath` to the `embedderModelsFolder` value. The directory contains the ONNX model graph, tokenizer configuration, and model JSON config required for inference.

### Is remote model loading allowed during embedding operations?

No. The `Embedder.init()` method explicitly sets `env.allowRemoteModels = false` and `env.allowLocalModels = true` before creating the pipeline. This configuration prevents the transformer library from attempting to download missing files from Hugging Face during inference, ensuring that embedding operations fail safely if the local model is absent rather than introducing network latency or external dependencies during runtime.