# How the 5ire Embedding Model Processes Documents for the Knowledge Base and Stores Embeddings

> Discover how the 5ire embedding model transforms documents into 1024-dimensional vectors using Xenova/bge-m3. Learn about Embedder services and pgvector storage for knowledge base similarity search.

- Repository: [Ironben/5ire](https://github.com/nanbingxyz/5ire)
- Tags: how-to-guide
- Published: 2026-03-07

---

**The 5ire application converts raw documents into 1024-dimensional embedding vectors using the Xenova/bge-m3 transformer model, orchestrating the workflow through the Embedder and DocumentEmbedder services before persisting vectors in a PostgreSQL database with pgvector for efficient similarity search.**

The open-source 5ire project (nanbingxyz/5ire) implements a retrieval-augmented generation (RAG) pipeline that transforms unstructured documents into searchable vector representations. Understanding how the embedding model processes documents for the knowledge base and stores embeddings reveals the architecture behind its semantic search capabilities.

## Core Architecture: Embedder and DocumentEmbedder Services

The system separates concerns between low-level vector generation and high-level document workflow management. The **Embedder** service ([`src/main/services/embedder.ts`](https://github.com/nanbingxyz/5ire/blob/main/src/main/services/embedder.ts)) handles model loading and text-to-vector conversion, while the **DocumentEmbedder** service ([`src/main/services/document-embedder.ts`](https://github.com/nanbingxyz/5ire/blob/main/src/main/services/document-embedder.ts)) orchestrates document fetching, text extraction, and database persistence.

## Text-to-Vector Conversion with the Embedder Service

### Model Initialization and File Verification

Before processing begins, the Embedder verifies that all required model files exist locally. The constructor initializes state as `idle` and records the model name `Xenova/bge-m3` along with required file paths defined in [`src/main/constants.ts`](https://github.com/nanbingxyz/5ire/blob/main/src/main/constants.ts).

```typescript
// src/main/services/embedder.ts
constructor() {
  this.model = DOCUMENT_EMBEDDING_MODEL_NAME; // "Xenova/bge-m3"
  this.files = DOCUMENT_EMBEDDING_MODEL_FILES;
  this.status = "idle";
}

```

The `init()` method checks for missing files. If any are absent, the service transitions to `unavailable`; otherwise, it imports `@xenova/transformers` and creates a feature-extraction pipeline.

```typescript
// src/main/services/embedder.ts
env.allowRemoteModels = false;
env.allowLocalModels = true;
this.pipeline = await pipeline("feature-extraction", this.model);

```

### The Embedding Pipeline and Concurrency Control

The `embed(texts: string[])` method converts input strings into 1024-dimensional **Float32Array** vectors. It validates that the service status is `ready`, then processes each text through the transformer pipeline.

```typescript
// src/main/services/embedder.ts
async embed(texts: string[]): Promise<Float32Array[]> {
  if (this.status !== "ready") throw new Error("Embedder not ready");
  
  const outputs = await this.pipeline(texts, { pooling: "mean", normalize: true });
  return outputs.map((output: any) => output.data);
}

```

To prevent memory exhaustion, a private `#concurrentRequests` counter limits parallel embedding operations. The counter increments before each pipeline call and decrements afterward, ensuring the service remains stable under load.

## Document Workflow Orchestration

### Pending Document Selection and Worker Management

The DocumentEmbedder manages the end-to-end flow from raw document to stored vectors. It queries the database for rows where `status = "pending"` or `failed` (for retries), then processes them using a worker pool capped at `MAX_WORKERS` (equal to `os.cpus().length`).

```typescript
// src/main/services/document-embedder.ts
const pendingDocs = await db.select()
  .from(schema.document)
  .where(eq(schema.document.status, "pending"));

if (this.workers < MAX_WORKERS) {
  await this.#process(doc.id, doc.url);
}

```

### Text Extraction and Vector Generation

For each document, the service downloads the file and extracts plain text using helper modules. The extracted text is split into chunks, then passed to `Embedder.embed()` to generate vectors.

```typescript
// src/main/services/document-embedder.ts
const texts = await this.#extractText(url);
const vectors = await embedder.embed(texts); // Float32Array[1024] per chunk

```

### Persisting Embeddings to PostgreSQL

The system stores each chunk as a row in the `documentChunk` table. The `embedding` column accepts the raw `Float32Array`, which PostgreSQL stores as a `vector(1024)` type via the **pgvector** extension.

```typescript
// src/main/services/document-embedder.ts
await db.insert(schema.documentChunk).values({
  documentId: id,
  text: chunk,
  embedding: vector, // 1024-dimensional Float32Array
});

```

After all chunks are persisted, the parent document row is updated to `status = "completed"`. Failures set `status = "failed"` and emit events for the UI to handle.

## Database Schema and Vector Storage

### The documentChunk Table and Vector Type

The storage layer relies on PostgreSQL with the **pgvector** extension. The schema definition in [`src/main/database/schema/tables.ts`](https://github.com/nanbingxyz/5ire/blob/main/src/main/database/schema/tables.ts) declares the `documentChunk` table with a dedicated vector column:

```typescript
// src/main/database/schema/tables.ts
export const documentChunk = table("documentChunk", {
  id: serial("id").primaryKey(),
  documentId: integer("documentId").notNull(),
  embedding: vector({ dimensions: 1024 }).notNull(), // pgvector type
  text: text("text").notNull(),
});

```

### HNSW Indexing for Similarity Search

To enable fast retrieval, the schema creates an **HNSW (Hierarchical Navigable Small World)** index on the embedding column using cosine similarity operations:

```typescript
// src/main/database/schema/tables.ts
createIndex("documentChunk_embedding_idx")
  .on(documentChunk)
  .using("hnsw", documentChunk.embedding.op("vector_cosine_ops"));

```

This index allows the system to execute approximate nearest-neighbor searches in milliseconds, even with millions of chunks.

## Query-Time Retrieval Flow

When a user submits a question, the system follows this retrieval path:

1. **Embed the query**: The question text is passed to `Embedder.embed([question])`, producing a 1024-dimensional vector.
2. **Similarity search**: The query vector is compared against the `documentChunk.embedding` column using the `cosine_distance` function provided by pgvector.
3. **Return context**: The top-k most similar chunks are retrieved and passed to the LLM as contextual grounding.

```typescript
// Conceptual query implementation
const queryVector = (await embedder.embed([userQuestion]))[0];
const results = await db
  .select()
  .from(schema.documentChunk)
  .orderBy(
    sql`cosine_distance(${schema.documentChunk.embedding}, ${queryVector})`
  )
  .limit(5);

```

## Summary

- The **Embedder** service ([`src/main/services/embedder.ts`](https://github.com/nanbingxyz/5ire/blob/main/src/main/services/embedder.ts)) loads the `Xenova/bge-m3` transformer model and converts text chunks into 1024-dimensional **Float32Array** vectors.
- The **DocumentEmbedder** service ([`src/main/services/document-embedder.ts`](https://github.com/nanbingxyz/5ire/blob/main/src/main/services/document-embedder.ts)) orchestrates the workflow: selecting pending documents, extracting text, generating embeddings, and persisting data.
- Vectors are stored in **PostgreSQL** using the **pgvector** extension in the `documentChunk` table, with an **HNSW index** enabling fast cosine-similarity search.
- Concurrency guards in both services prevent memory exhaustion: the Embedder limits parallel requests with a private counter, while the DocumentEmbedder uses a CPU-count-based worker pool.
- At query time, user questions are embedded and matched against stored vectors using `cosine_distance` to retrieve semantically relevant context for the LLM.

## Frequently Asked Questions

### What embedding model does 5ire use for knowledge base documents?

5ire uses the **Xenova/bge-m3** transformer model, a BERT-based architecture optimized for semantic search. The model is loaded via the `@xenova/transformers` library in [`src/main/services/embedder.ts`](https://github.com/nanbingxyz/5ire/blob/main/src/main/services/embedder.ts) and generates 1024-dimensional dense vectors for each text chunk. All model files are verified locally before initialization to ensure offline capability.

### How does 5ire prevent memory issues when embedding large documents?

The system implements two layers of concurrency control. The **Embedder** service maintains a private `#concurrentRequests` counter that limits parallel pipeline executions to avoid GPU/CPU memory exhaustion. Separately, the **DocumentEmbedder** uses a worker pool capped at `os.cpus().length` to limit simultaneous document processing, ensuring stable performance even with high-volume imports.

### What database extension enables vector storage in 5ire?

5ire relies on **pgvector**, a PostgreSQL extension that provides the `vector` data type and similarity search operators. The schema in [`src/main/database/schema/tables.ts`](https://github.com/nanbingxyz/5ire/blob/main/src/main/database/schema/tables.ts) defines the `embedding` column as `vector({ dimensions: 1024 })` and creates an **HNSW index** using `vector_cosine_ops` for efficient approximate nearest-neighbor queries.

### How does 5ire retrieve relevant context when a user asks a question?

At query time, the application embeds the user's question using the same `Embedder.embed()` method to produce a 1024-dimensional vector. It then executes a SQL query using the `cosine_distance` function against the `documentChunk.embeddings` column, ordering results by similarity and returning the top-k chunks. These retrieved chunks provide contextual grounding for the LLM's response.