How to Configure the Embedding Wrapper for Different Embedder Models in MemPalace
MemPalace abstracts embedding generation behind a unified wrapper that supports multiple ONNX models—MiniLM for fast CPU inference and EmbeddingGemma for multilingual tasks—configurable via environment variables or a JSON config file.
The MemPalace vector memory system delegates all embedding vector creation to a centralized wrapper architecture. Located in mempalace/backends/embedding_wrapper.py, this abstraction allows storage backends like SQLite-exact or pgvector to work with any supported embedder model without code changes. The actual model instantiation occurs through the get_embedding_function factory in mempalace/embedding.py, which handles model selection, device resolution, and caching.
How the Embedding Wrapper Architecture Works
When a backend requires explicit vectors—indicated by the requires_explicit_embeddings flag—MemPalace wraps the underlying collection with an EmbeddingCollection instance.
The wrapper intercepts calls to add() or upsert() where embeddings are omitted. In these cases, it delegates to _embed_texts, which invokes get_embedding_function() to obtain the appropriate embedder【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/backends/embedding_wrapper.py#L10-L18】. This factory function reads configuration from MempalaceConfig, checking for model and device specifications passed via initialization, environment variables, or the ~/.mempalace/config.json file【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/embedding.py#L46-L53】.
Supported Embedder Models
MemPalace currently supports two distinct ONNX-based embedding models, selectable via the embedding_model parameter:
-
MiniLM (
minilm): The default model, implemented as a subclass ofchromadb.utils.embedding_functions.ONNXMiniLM_L6_V2named "default". This model offers fast inference on CPU and is ideal for general-purpose English text embedding【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/embedding.py#L10-L18】【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/embedding.py#L110-L118】. -
EmbeddingGemma (
embeddinggemma): A multilingual ONNX model based ononnx-community/embeddinggemma-300m-ONNX. The factory creates anEmbeddinggemmaONNXinstance that downloads the model weights lazily on first use, supporting non-English languages and GPU acceleration【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/embedding.py#L30-L38】【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/embedding.py#L40-L66】.
Configuration Methods
You can configure the embedding wrapper through two primary mechanisms, evaluated in order of precedence:
Environment Variables
Set the following variables to override all other configuration:
export MEMPALACE_EMBEDDING_MODEL=embeddinggemma
export MEMPALACE_EMBEDDING_DEVICE=cuda
Valid values for MEMPALACE_EMBEDDING_MODEL are minilm or embeddinggemma. For MEMPALACE_EMBEDDING_DEVICE, use auto, cpu, cuda, coreml, or dml.
Configuration File
For persistent settings across sessions, modify ~/.mempalace/config.json:
{
"embedding_model": "embeddinggemma",
"embedding_device": "cuda"
}
The MempalaceConfig class reads these fields when the environment variables or explicit parameters are None【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/embedding.py#L46-L53】.
Device Resolution and ONNX Providers
The _resolve_providers function in mempalace/embedding.py determines the ONNX execution providers based on your device selection:
- Auto-detection: When
device="auto", the system attempts providers in order: CUDA → CoreML → DirectML (DML) → CPU. - Explicit selection: Maps requests via
_PROVIDER_MAPto specific ONNX providers. - Fallback: If the requested provider is unavailable, the system warns once and falls back to CPU【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/embedding.py#L68-L88】.
The factory caches created embedding functions in _EF_CACHE keyed by the tuple (model, providers), ensuring that subsequent requests reuse the already-loaded model without reinitialization【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/embedding.py#L55-L60】【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/embedding.py#L66-L74】.
Practical Code Examples
Example 1: Using the Default MiniLM Model
from mempalace.backends.embedding_wrapper import EmbeddingCollection
from mempalace.backends.sqlite_exact import SqliteExactCollection
inner = SqliteExactCollection(name="my_palace")
palace = EmbeddingCollection(inner)
# Embeddings generated automatically using MiniLM
palace.add(documents=["Hello world", "Machine learning"], ids=["1", "2"])
Example 2: Switching to EmbeddingGemma via Environment Variables
import os
os.environ["MEMPALACE_EMBEDDING_MODEL"] = "embeddinggemma"
os.environ["MEMPALACE_EMBEDDING_DEVICE"] = "cpu"
from mempalace.backends.embedding_wrapper import EmbeddingCollection
from mempalace.backends.pgvector import PgVectorCollection
inner = PgVectorCollection(name="multilingual_palace")
palace = EmbeddingCollection(inner)
# First call triggers lazy download of the EmbeddingGemma ONNX model
palace.add(documents=["¡Hola!", "你好!", "Bonjour"], ids=["a", "b", "c"])
Example 3: Manual Embedder Instantiation
from mempalace.embedding import get_embedding_function
ef = get_embedding_function(model="embeddinggemma", device="auto")
vectors = ef(["Sample sentence", "Another example"])
# Returns list of 384-dimensional float vectors
Critical Considerations for Model Switching
Changing the embedding_model on an existing palace requires rebuilding the entire index because vector dimensions and semantic spaces differ between models. After modifying the configuration, execute:
mempalace repair rebuild-index
This command regenerates all embeddings using the newly configured model and updates the vector store accordingly.
Summary
- Architecture: The
EmbeddingCollectionwrapper inmempalace/backends/embedding_wrapper.pyautomatically injects embeddings for backends requiring explicit vectors. - Models: Choose between MiniLM (fast, CPU-optimized) and EmbeddingGemma (multilingual, GPU-capable) via the
embedding_modelparameter. - Configuration: Set
MEMPALACE_EMBEDDING_MODELandMEMPALACE_EMBEDDING_DEVICEenvironment variables, or persist choices in~/.mempalace/config.json. - Performance: Embedding functions are cached by
(model, providers)tuple to avoid reloading ONNX models unnecessarily. - Migration: Switching models necessitates running
mempalace repair rebuild-indexto re-embed existing documents.
Frequently Asked Questions
What embedding models does MemPalace support?
MemPalace supports two ONNX models: MiniLM (the default, based on all-MiniLM-L6-v2 via ChromaDB's implementation) for efficient CPU inference, and EmbeddingGemma (300M parameter multilingual model) for cross-lingual applications. You select between them using the embedding_model configuration key with values minilm or embeddinggemma.
How do I configure GPU acceleration for embeddings?
Set MEMPALACE_EMBEDDING_DEVICE=cuda (or coreml for Apple Silicon, dml for DirectML on Windows) either as an environment variable or in your config file. When device is set to auto, MemPalace attempts CUDA first, then CoreML, then DirectML, falling back to CPU if none are available.
Do I need to rebuild my data when changing the embedding model?
Yes. Because different models produce vectors in incompatible semantic spaces with varying dimensions, you must run mempalace repair rebuild-index after switching models. This command iterates through all documents and regenerates embeddings using the newly configured embedder.
How does MemPalace cache embedding models?
The get_embedding_function factory stores instantiated embedders in an internal _EF_CACHE dictionary keyed by the tuple (model, providers). This ensures that multiple collections or reconnection attempts reuse the same ONNX model instance in memory, eliminating redundant loading and initialization overhead.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →