Switching Embedding Models in MemPalace: Gemma-300m vs all-MiniLM-L6-v2
To switch embedding models in MemPalace, set the MEMPALACE_EMBEDDING_MODEL environment variable to embeddinggemma or minilm, then run mempalace repair rebuild-index to rebuild your vector index with the new model.
MemPalace abstracts vector generation behind a home-grown embedding factory in mempalace/embedding.py, allowing you to swap between the lightweight English-only MiniLM and the multilingual EmbeddingGemma-300m. Both models produce 384-dimensional vectors but differ in language support and computational requirements, making the switch valuable when moving from English-only to multilingual document collections.
Model Comparison: MiniLM vs EmbeddingGemma
MemPalace supports two distinct embedding architectures through the factory pattern implemented in get_embedding_function().
all-MiniLM-L6-v2 (Default)
The MiniLM model (identifier: minilm) is the default English-only embedding model shipped with ChromaDB. It generates 384-dimensional vectors and executes in the else branch of get_embedding_function (lines 64-66 in mempalace/embedding.py). This model is optimized for speed and memory footprint, making it ideal for monolingual English document collections.
EmbeddingGemma-300m (Multilingual)
The EmbeddingGemma model (identifier: embeddinggemma) uses an ONNX-quantized 300MB checkpoint supporting 100+ languages. It produces the same 384-dimensional output via Matryoshka truncation but requires the EmbeddinggemmaONNX class (lines 30-59 in mempalace/embedding.py). When model == "embeddinggemma", the factory instantiates this class rather than the default MiniLM path.
How Model Selection Works
MemPalace resolves embedding configuration through a hierarchy of environment variables and persistent config files.
Configuration Sources
The model identifier is read from either:
- The
MEMPALACE_EMBEDDING_MODELenvironment variable - The
embedding_modelkey in~/.mempalace/config.json(accessed viaMempalaceConfig.embedding_modelproperty inmempalace/config.pylines 6-18)
Factory Resolution
The get_embedding_function() factory (lines 38-68 in mempalace/embedding.py) resolves the identifier, builds a cached embedding-function instance, and returns it to the rest of the codebase. For backends requiring explicit vectors (ChromaDB, SQLite, Qdrant, pgvector), the EmbeddingCollection wrapper in backends/embedding_wrapper.py (lines 10-19) automatically invokes the factory via _embed_texts when callers supply documents without pre-computed vectors.
Switching Models Step-by-Step
Method 1: Temporary Switch via Environment Variable
Set the environment variable before running MemPalace commands:
export MEMPALACE_EMBEDDING_MODEL=embeddinggemma
mempalace mine ./my_project
Method 2: Persistent Switch via Config File
Use the Python API to write the configuration permanently:
from mempalace.config import MempalaceConfig
cfg = MempalaceConfig()
cfg.set_embedding_model("embeddinggemma") # Writes to ~/.mempalace/config.json
Rebuild the Index
When you switch models, the vector space changes because different embeddings produce different numeric representations. ChromaDB and other backends store the embedding function name on each collection; if the persisted name does not match the factory-supplied name, reads will fail with a mismatch error (as noted in the comment at lines 15-17 of mempalace/embedding.py).
Re-embed all existing drawers with:
mempalace repair rebuild-index
Hardware Acceleration
Both models support hardware acceleration through the MEMPALACE_EMBEDDING_DEVICE environment variable or the embedding_device config key. The _resolve_providers function (lines 62-84 in mempalace/embedding.py) handles device selection, supporting CUDA, CoreML, DirectML, or CPU fallback.
export MEMPALACE_EMBEDDING_MODEL=embeddinggemma
export MEMPALACE_EMBEDDING_DEVICE=cuda
mempalace repair rebuild-index
Verifying the Active Model
Inspect the currently configured embedding function at runtime:
from mempalace.embedding import get_embedding_function
ef = get_embedding_function()
print("Active EF name:", ef.name())
# Returns "embeddinggemma_300m" or "default" (MiniLM)
For custom backend integrations, use the wrapper directly:
from mempalace.backends.embedding_wrapper import EmbeddingCollection
from mempalace.backends.chromadb import ChromaCollection
raw_collection = ChromaCollection(name="my_drawers")
emb_collection = EmbeddingCollection(raw_collection)
# Documents are auto-embedded with the currently configured model
emb_collection.add(ids=["doc1"], documents=["Hello world!"])
Summary
- Model identifiers: Use
minilmfor all-MiniLM-L6-v2 (default) orembeddinggemmafor the multilingual 300m model. - Configuration: Set via
MEMPALACE_EMBEDDING_MODELenvironment variable orembedding_modelin~/.mempalace/config.json. - Factory location:
get_embedding_function()inmempalace/embedding.pyhandles model instantiation and caching. - Required action: Always run
mempalace repair rebuild-indexafter switching models to regenerate vectors in the new semantic space. - Hardware support: Configure accelerators via
MEMPALACE_EMBEDDING_DEVICEor theembedding_deviceconfig key. - Backend integration: The
EmbeddingCollectionwrapper inbackends/embedding_wrapper.pyautomatically applies the active model to add/upsert/query operations.
Frequently Asked Questions
What happens if I don't rebuild the index after switching embedding models?
If you skip the rebuild, ChromaDB and other backends will detect an embedding function name mismatch between the stored metadata and the active factory configuration, causing read operations to fail. The repository explicitly warns about this in mempalace/embedding.py (lines 13-17) because the vector space changes completely when switching from MiniLM to EmbeddingGemma or vice versa.
Can I use both MiniLM and EmbeddingGemma simultaneously in the same MemPalace instance?
No. MemPalace uses a singleton factory pattern via get_embedding_function() that returns a single cached embedding function instance per process. While you can instantiate different models in separate Python processes or environment contexts, a single MemPalace collection cannot mix vectors from different embedding spaces.
Which model should I choose for non-English content?
Choose EmbeddingGemma (identifier: embeddinggemma). The all-MiniLM-L6-v2 model is English-only, whereas EmbeddingGemma supports 100+ languages through its multilingual training. Both models output 384-dimensional vectors, so they are compatible with the same storage backends, but only EmbeddingGemma will produce meaningful semantic similarities for non-English text.
How do I check which embedding model is currently active?
Call get_embedding_function().name() from mempalace/embedding.py. This returns either "embeddinggemma_300m" when using the Gemma model or "default" when using MiniLM. You can also inspect the embedding_model property on your MempalaceConfig instance to see the configured value before factory instantiation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →