How OmniRoute's Memory System Works with Int8 Vector Quantization
OmniRoute's memory system uses int8 vector quantization to compress 32-bit float embeddings into 8-bit integers, reducing storage and memory bandwidth by ~75% while preserving search accuracy through SQLite's built-in vec_quantize_int8 function or Qdrant's scalar quantization with rescoring.
The int8 vector quantization feature in OmniRoute enables efficient storage of user memories without sacrificing retrieval quality. When enabled, the system automatically converts dense float embeddings into compact signed 8-bit representations, cutting per-dimension storage from 4 bytes to 1 byte. This guide examines the implementation across SQLite-vec and Qdrant backends based on the v3.8.50 source code.
Enabling Int8 Quantization Mode
OmniRoute controls quantization through environment variables. Set MEMORY_VEC_QUANTIZATION=int8 to activate compression for the default SQLite-vec store, or use QDRANT_QUANTIZATION=int8 when targeting Qdrant.
# SQLite-vec backend (default)
export MEMORY_VEC_QUANTIZATION=int8
# Qdrant backend
export MEMORY_VECTOR_STORE=qdrant
export QDRANT_QUANTIZATION=int8
The system validates these settings through requestedVecQuantization() in src/lib/memory/vectorStore.ts (lines 92-94) and QdrantQuantization enum in src/lib/memory/qdrant.ts (lines 9-14).
SQLite-Vec Int8 Pipeline: From Insertion to Search
Vector Insertion and Encoding
When VectorStore.upsertVector() receives a Float32Array embedding, it delegates encoding to encodeVector() before storage. For int8-quantized tables, the raw bytes pass through SQLite's vec_quantize_int8(?, 'unit') function:
// src/lib/memory/vectorStore.ts lines 79-81
const quantization = this.requestedVecQuantization();
const vec = quantization === "int8"
? sql`vec_quantize_int8(${sql.raw(`?`)}, 'unit')`
: sql.raw(`?`);
The 'unit' parameter instructs the extension to unit-normalize vectors before quantization, preserving cosine-similarity relationships essential for semantic search.
Table Signature and Schema Migration
The ensureReady() method constructs a table signature that embeds the quantization mode. The helper addInt8SuffixToSignature() appends :int8 when active (lines 105-108):
// Signature examples
"memories:vec" // float32 storage
"memories:vec:int8" // int8 quantized storage
Signature changes trigger automatic re-indexing. When OmniRoute detects a mismatch between the stored signature and current configuration, it:
- Drops and recreates the
vec_memoriesvirtual table - Sets
needs_reindex=1on all memory rows - Queues background regeneration through
memory/reindex.ts
Search Behavior with Int8 Vectors
OmniRoute offers two search paths, both transparent to quantization state:
| Method | Function | Int8 Handling |
|---|---|---|
| Pure K-NN | searchVector() |
SQLite de-quantizes automatically via vec_search |
| Hybrid RRF | searchHybrid() |
Combines int8 vector results with FTS5 scores |
The caller receives distance scores without managing quantization details. The SQLite-vec extension handles de-quantization internally during vec_search execution.
Qdrant Backend: Scalar Quantization with Rescoring
When configured for Qdrant, OmniRoute builds a scalar quantization config that enables higher-accuracy retrieval:
// src/lib/memory/qdrant.ts lines 38-41
quantization: {
scalar: {
type: "int8",
always_ram: true, // Keep quantized vectors in RAM
quantile: 0.99 // Exclude 1% outlier values
}
}
The always_ram: true setting pins quantized vectors in memory for fast candidate retrieval. Qdrant then rescores the shortlist using original float vectors, recovering precision lost to 8-bit compression. This two-phase approach balances speed and accuracy.
Re-Index Flow: Handling Configuration Changes
The memory/reindex.ts worker processes memories marked for rebuild. At line 92, it regenerates embeddings with the current provider and calls vec.upsertVector(), ensuring new quantization settings take effect:
// src/lib/memory/reindex.ts line 92
await vectorStore.upsertVector(memory.id, newVector);
This mechanism enables non-disruptive mode switches—admins can toggle MEMORY_VEC_QUANTIZATION and let the background worker migrate existing memories incrementally.
Code Examples
Basic Int8 Memory Storage
import { embedText } from "@/lib/memory/embedding";
import { getVectorStore } from "@/lib/memory/vectorStore";
// Configuration (typically in .env)
process.env.MEMORY_VEC_QUANTIZATION = "int8";
async function storeMemory(id: string, content: string) {
const { vector } = await embedText({
model: "all-MiniLM-L6-v2",
text: content
});
const store = await getVectorStore();
// Automatically uses vec_quantize_int8
await store.upsertVector(id, vector);
}
Hybrid Search with Int8 Backend
async function findMemories(query: string, limit = 10) {
const { vector } = await embedText({
model: "all-MiniLM-L6-v2",
text: query
});
const store = await getVectorStore();
// Ranks memories by vector similarity
const semanticResults = await store.searchVector(vector, limit);
// Combines semantic + keyword signals
const combinedResults = await store.searchHybrid(vector, query, limit);
return { semanticResults, combinedResults };
}
Initializing Qdrant with Int8
process.env.MEMORY_VECTOR_STORE = "qdrant";
process.env.QDRANT_QUANTIZATION = "int8";
process.env.QDRANT_URL = "http://localhost:6333";
const store = await getVectorStore();
// Configures scalar quantization with 0.99 quantile rescoring
Key Implementation Files
| File | Purpose |
|---|---|
src/lib/memory/vectorStore.ts |
Core SQLite-vec logic; vec_quantize_int8 integration, signature management |
src/lib/memory/qdrant.ts |
Qdrant client; QdrantQuantization enum, scalar config building |
src/lib/memory/reindex.ts |
Background worker for embedding regeneration |
tests/unit/memory-vectorstore-int8-quant.test.ts |
Validation of :int8 signatures and search correctness |
Summary
- Int8 quantization activates via
MEMORY_VEC_QUANTIZATION=int8(SQLite) orQDRANT_QUANTIZATION=int8(Qdrant) - SQLite-vec uses
vec_quantize_int8(?, 'unit')with automatic de-quantization during search - Table signatures encode the quantization mode; changes force re-indexing via
needs_reindexflags - Qdrant employs scalar quantization with
always_ram: trueand float-vector rescoring for accuracy - Both backends expose identical APIs—callers work with
Float32Arrayinputs and distance scores regardless of internal representation
Frequently Asked Questions
How much storage does int8 quantization save?
Int8 reduces per-dimension storage from 4 bytes to 1 byte, yielding approximately 75% savings on vector data. For a 1536-dimensional embedding (common with OpenAI models), this drops from ~6 KB to ~1.5 KB per memory.
Does int8 quantization hurt search accuracy?
For most LLM-driven retrieval tasks, the accuracy loss is negligible. SQLite-vec's unit-normalized quantization preserves cosine-similarity relationships well. Qdrant's configuration adds rescoring with original floats to recover any precision loss on the top-K candidates.
What happens if I switch quantization modes on existing data?
OmniRoute detects signature mismatches in ensureReady() and triggers automatic re-indexing. All existing memories receive needs_reindex=1, and the memory/reindex.ts worker regenerates embeddings with the new quantization format. Search remains available during migration, though results may mix formats temporarily.
Can I use int8 quantization with any embedding model?
Yes—the quantization happens after embedding generation. Any model producing float vectors (OpenAI, local transformers, etc.) works with int8 mode. The vec_quantize_int8 function handles the conversion regardless of source model.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →