MemPalace Multi-Language Embedding Support and Configuration: A Deep Architectural Guide
MemPalace provides a configurable embedding layer that defaults to the multilingual Embeddinggemma ONNX model and automatically falls back to CPU execution when hardware accelerators are unavailable.
MemPalace stores every word a user says verbatim and relies on semantic embeddings to enable efficient retrieval. The develop branch introduces a flexible embedding system capable of running on CPU or hardware accelerators while supporting multilingual text through a lazy-loaded ONNX model. This article examines the architecture and configuration of MemPalace multi-language embedding support and configuration based on the source code in the MemPalace/mempalace repository.
How Embedding Configuration Works in MemPalace
The MemPalace configuration system exposes two primary settings that control the embedding pipeline.
Supported Embedding Models
MemPalace supports two distinct embedding functions selectable via the MEMPALACE_EMBEDDING_MODEL environment variable or the embedding_model key in ~/.mempalace/config.json:
minilm— The historic default based on all-MiniLM-L6-v2, producing 384-dimensional vectors optimized for English-only content.embeddinggemma— A multilingual ONNX model (onnx-community/embeddinggemma-300m-ONNX, quantized q8) that supports over 100 languages. After Matryoshka truncation, it yields 384-dimensional vectors, making it compatible with existing ChromaDB collections without schema changes.
According to the source code in [mempalace/embedding.py](https://github.com/MemPalace/mempalace/blob/develop/mempalace/embedding.py#L6-L18), cross-lingual cosine similarity for the multilingual model measures approximately 0.88 compared to roughly 0.35 for MiniLM.
Device and Execution Provider Selection
The MEMPALACE_EMBEDDING_DEVICE setting (or embedding_device in config) controls hardware acceleration. The _resolve_providers helper in [mempalace/embedding.py](https://github.com/MemPalace/mempalace/blob/develop/mempalace/embedding.py#L62-L81) inspects installed onnxruntime providers and resolves the following priority:
- CUDA
- CoreML
- DirectML
- CPU
If the requested accelerator is unavailable, the system falls back to CPU with a one-time warning. This ensures mining works on any laptop regardless of GPU availability.
Core Architecture of the Embedding Layer
The Embedding Factory
[mempalace/embedding.py](https://github.com/MemPalace/mempalace/blob/develop/mempalace/embedding.py) serves as the public factory that instantiates the correct embedding function based on user configuration. Lines 6–18 define the supported models, while lines 110–127 create a Chroma-compatible subclass for the default MiniLM embedding function. The module reads configuration values set in [mempalace/config.py](https://github.com/MemPalace/mempalace/blob/develop/mempalace/config.py#L610-L628) through getters exposed at lines 610–628.
Multilingual Model Implementation
The EmbeddinggemmaONNX class implements the multilingual backend. Its name() method returns the stable identifier "embeddinggemma_300m", which ChromaDB stores on the collection to detect mismatched embedding functions. The class lazily loads the model, tokenizer, and ONNX session on first call, downloading the artifact via huggingface_hub if not already cached.
This implementation is defined at lines 54–70 of [mempalace/embedding.py](https://github.com/MemPalace/mempalace/blob/develop/mempalace/embedding.py#L54-L70), with lazy loading and download logic spanning lines 73–99.
Backend Integration
[mempalace/backends/embedding_wrapper.py](https://github.com/MemPalace/mempalace/blob/develop/mempalace/backends/embedding_wrapper.py#L38) defines EmbeddingCollection, a thin wrapper that makes any embedding function conform to the ChromaDB collection interface. The core palace class in [mempalace/palace.py](https://github.com/MemPalace/mempalace/blob/develop/mempalace/palace.py) instantiates this wrapper at startup, injecting the user-selected embedding function.
Onboarding and Persistent Configuration
When MemPalace runs for the first time, [mempalace/onboarding.py](https://github.com/MemPalace/mempalace/blob/develop/mempalace/onboarding.py) prompts the user to select an embedding model, defaulting new installations to the multilingual option. The selection is persisted through [mempalace/config.py](https://github.com/MemPalace/mempalace/blob/develop/mempalace/config.py#L630), which provides the setter at line 630.
Practical Configuration Examples
Configure the Embedding Model via Python
You can programmatically switch to the multilingual model using the configuration API:
from mempalace.config import Config
cfg = Config()
cfg.set_embedding_model("embeddinggemma")
cfg.set_embedding_device("auto")
This persists the values to ~/.mempalace/config.json and will be respected across subsequent sessions.
Change Models Using the CLI
The MemPalace CLI exposes commands to inspect and modify embedding settings:
# Display current embedding configuration
mempalace config show
# Switch to the multilingual model
mempalace config set embedding_model embeddinggemma
# Re-index existing data after switching models
mempalace repair rebuild-index
Because MiniLM and Embeddinggemma operate in different vector spaces, you must run mempalace repair rebuild-index after changing the model.
Use the Embedding Function Directly
For advanced integrations, instantiate the multilingual embedding function directly:
from mempalace.embedding import EmbeddinggemmaONNX, _resolve_providers
# Resolve the best available provider list
providers, device = _resolve_providers("auto")
# Initialize the embedding function (triggers lazy download on first call)
ef = EmbeddinggemmaONNX(preferred_providers=providers)
# Generate a 384-dimensional vector for non-English text
vector = ef(["task: sentence similarity | query: こんにちは世界"])
print(vector.shape) # (1, 384)
Search with the Configured Embedding Model
Once configured, the MemPalace class automatically loads the correct embedding function:
from mempalace import MemPalace
pal = MemPalace()
pal.add_documents(["Hello world"], ids=["doc1"])
# Query using the same multilingual embedding space
results = pal.query("こんにちは", top_k=5)
print(results)
Key Source Files
- [
mempalace/embedding.py](https://github.com/MemPalace/mempalace/blob/develop/mempalace/embedding.py) — Factory, provider resolution via_resolve_providers, and theEmbeddinggemmaONNXimplementation. - [
mempalace/config.py](https://github.com/MemPalace/mempalace/blob/develop/mempalace/config.py#L610-L630) — Reads and writesembedding_modelandembedding_deviceproperties. - [
mempalace/backends/embedding_wrapper.py](https://github.com/MemPalace/mempalace/blob/develop/mempalace/backends/embedding_wrapper.py#L38) — Wraps embedding functions as ChromaDB-compatible collections. - [
mempalace/onboarding.py](https://github.com/MemPalace/mempalace/blob/develop/mempalace/onboarding.py) — First-run prompt that defaults new users to the multilingual model. - [
mempalace/palace.py](https://github.com/MemPalace/mempalace/blob/develop/mempalace/palace.py) — Core palace class that wires the selected embedding function into the storage backend.
Summary
- MemPalace supports two embedding models: the English-only MiniLM and the multilingual Embeddinggemma ONNX model.
- Device selection is automatic, falling back from CUDA to CoreML to DirectML to CPU.
- The
EmbeddinggemmaONNXclass lazily downloads and caches the multilingual model viahuggingface_hub. - Configuration is managed through environment variables, the CLI, or
~/.mempalace/config.jsonviaconfig.py. - Existing palaces must be re-indexed with
mempalace repair rebuild-indexafter changing embedding models. - The architecture separates concerns between factory logic, ONNX inference, ChromaDB wrapper compatibility, and user onboarding.
Frequently Asked Questions
What embedding models does MemPalace support?
MemPalace supports minilm (all-MiniLM-L6-v2, English-only, 384-dim) and embeddinggemma (multilingual ONNX model supporting over 100 languages, also 384-dim after Matryoshka truncation). The multilingual model is available on the develop branch and is the default for new installations.
How do I switch from MiniLM to the multilingual embedding model?
You can switch by running mempalace config set embedding_model embeddinggemma in the CLI or by calling cfg.set_embedding_model("embeddinggemma") on a Config instance in Python. After switching, execute mempalace repair rebuild-index to re-embed existing documents in the new vector space.
Does MemPalace require a GPU for multilingual embeddings?
No. The _resolve_providers function in mempalace/embedding.py automatically falls back to CPU execution with a one-time warning if CUDA, CoreML, or DirectML are unavailable. The multilingual ONNX model runs efficiently on CPU.
Why do I need to rebuild my index after changing the embedding model?
MiniLM and Embeddinggemma map text into different vector spaces with incompatible geometries. Searching across mixed embeddings would return meaningless results, so MemPalace requires a full mempalace repair rebuild-index to ensure all vectors are generated by the same embedding function.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →