MemPalace Multi-Language Embedding Support and Configuration: A Deep Architectural Guide

MemPalace provides a configurable embedding layer that defaults to the multilingual Embeddinggemma ONNX model and automatically falls back to CPU execution when hardware accelerators are unavailable.

MemPalace stores every word a user says verbatim and relies on semantic embeddings to enable efficient retrieval. The develop branch introduces a flexible embedding system capable of running on CPU or hardware accelerators while supporting multilingual text through a lazy-loaded ONNX model. This article examines the architecture and configuration of MemPalace multi-language embedding support and configuration based on the source code in the MemPalace/mempalace repository.

How Embedding Configuration Works in MemPalace

The MemPalace configuration system exposes two primary settings that control the embedding pipeline.

Supported Embedding Models

MemPalace supports two distinct embedding functions selectable via the MEMPALACE_EMBEDDING_MODEL environment variable or the embedding_model key in ~/.mempalace/config.json:

  • minilm — The historic default based on all-MiniLM-L6-v2, producing 384-dimensional vectors optimized for English-only content.
  • embeddinggemma — A multilingual ONNX model (onnx-community/embeddinggemma-300m-ONNX, quantized q8) that supports over 100 languages. After Matryoshka truncation, it yields 384-dimensional vectors, making it compatible with existing ChromaDB collections without schema changes.

According to the source code in [mempalace/embedding.py](https://github.com/MemPalace/mempalace/blob/develop/mempalace/embedding.py#L6-L18), cross-lingual cosine similarity for the multilingual model measures approximately 0.88 compared to roughly 0.35 for MiniLM.

Device and Execution Provider Selection

The MEMPALACE_EMBEDDING_DEVICE setting (or embedding_device in config) controls hardware acceleration. The _resolve_providers helper in [mempalace/embedding.py](https://github.com/MemPalace/mempalace/blob/develop/mempalace/embedding.py#L62-L81) inspects installed onnxruntime providers and resolves the following priority:

  1. CUDA
  2. CoreML
  3. DirectML
  4. CPU

If the requested accelerator is unavailable, the system falls back to CPU with a one-time warning. This ensures mining works on any laptop regardless of GPU availability.

Core Architecture of the Embedding Layer

The Embedding Factory

[mempalace/embedding.py](https://github.com/MemPalace/mempalace/blob/develop/mempalace/embedding.py) serves as the public factory that instantiates the correct embedding function based on user configuration. Lines 6–18 define the supported models, while lines 110–127 create a Chroma-compatible subclass for the default MiniLM embedding function. The module reads configuration values set in [mempalace/config.py](https://github.com/MemPalace/mempalace/blob/develop/mempalace/config.py#L610-L628) through getters exposed at lines 610–628.

Multilingual Model Implementation

The EmbeddinggemmaONNX class implements the multilingual backend. Its name() method returns the stable identifier "embeddinggemma_300m", which ChromaDB stores on the collection to detect mismatched embedding functions. The class lazily loads the model, tokenizer, and ONNX session on first call, downloading the artifact via huggingface_hub if not already cached.

This implementation is defined at lines 54–70 of [mempalace/embedding.py](https://github.com/MemPalace/mempalace/blob/develop/mempalace/embedding.py#L54-L70), with lazy loading and download logic spanning lines 73–99.

Backend Integration

[mempalace/backends/embedding_wrapper.py](https://github.com/MemPalace/mempalace/blob/develop/mempalace/backends/embedding_wrapper.py#L38) defines EmbeddingCollection, a thin wrapper that makes any embedding function conform to the ChromaDB collection interface. The core palace class in [mempalace/palace.py](https://github.com/MemPalace/mempalace/blob/develop/mempalace/palace.py) instantiates this wrapper at startup, injecting the user-selected embedding function.

Onboarding and Persistent Configuration

When MemPalace runs for the first time, [mempalace/onboarding.py](https://github.com/MemPalace/mempalace/blob/develop/mempalace/onboarding.py) prompts the user to select an embedding model, defaulting new installations to the multilingual option. The selection is persisted through [mempalace/config.py](https://github.com/MemPalace/mempalace/blob/develop/mempalace/config.py#L630), which provides the setter at line 630.

Practical Configuration Examples

Configure the Embedding Model via Python

You can programmatically switch to the multilingual model using the configuration API:

from mempalace.config import Config

cfg = Config()
cfg.set_embedding_model("embeddinggemma")
cfg.set_embedding_device("auto")

This persists the values to ~/.mempalace/config.json and will be respected across subsequent sessions.

Change Models Using the CLI

The MemPalace CLI exposes commands to inspect and modify embedding settings:


# Display current embedding configuration

mempalace config show

# Switch to the multilingual model

mempalace config set embedding_model embeddinggemma

# Re-index existing data after switching models

mempalace repair rebuild-index

Because MiniLM and Embeddinggemma operate in different vector spaces, you must run mempalace repair rebuild-index after changing the model.

Use the Embedding Function Directly

For advanced integrations, instantiate the multilingual embedding function directly:

from mempalace.embedding import EmbeddinggemmaONNX, _resolve_providers

# Resolve the best available provider list

providers, device = _resolve_providers("auto")

# Initialize the embedding function (triggers lazy download on first call)

ef = EmbeddinggemmaONNX(preferred_providers=providers)

# Generate a 384-dimensional vector for non-English text

vector = ef(["task: sentence similarity | query: こんにちは世界"])
print(vector.shape)  # (1, 384)

Search with the Configured Embedding Model

Once configured, the MemPalace class automatically loads the correct embedding function:

from mempalace import MemPalace

pal = MemPalace()
pal.add_documents(["Hello world"], ids=["doc1"])

# Query using the same multilingual embedding space

results = pal.query("こんにちは", top_k=5)
print(results)

Key Source Files

Summary

  • MemPalace supports two embedding models: the English-only MiniLM and the multilingual Embeddinggemma ONNX model.
  • Device selection is automatic, falling back from CUDA to CoreML to DirectML to CPU.
  • The EmbeddinggemmaONNX class lazily downloads and caches the multilingual model via huggingface_hub.
  • Configuration is managed through environment variables, the CLI, or ~/.mempalace/config.json via config.py.
  • Existing palaces must be re-indexed with mempalace repair rebuild-index after changing embedding models.
  • The architecture separates concerns between factory logic, ONNX inference, ChromaDB wrapper compatibility, and user onboarding.

Frequently Asked Questions

What embedding models does MemPalace support?

MemPalace supports minilm (all-MiniLM-L6-v2, English-only, 384-dim) and embeddinggemma (multilingual ONNX model supporting over 100 languages, also 384-dim after Matryoshka truncation). The multilingual model is available on the develop branch and is the default for new installations.

How do I switch from MiniLM to the multilingual embedding model?

You can switch by running mempalace config set embedding_model embeddinggemma in the CLI or by calling cfg.set_embedding_model("embeddinggemma") on a Config instance in Python. After switching, execute mempalace repair rebuild-index to re-embed existing documents in the new vector space.

Does MemPalace require a GPU for multilingual embeddings?

No. The _resolve_providers function in mempalace/embedding.py automatically falls back to CPU execution with a one-time warning if CUDA, CoreML, or DirectML are unavailable. The multilingual ONNX model runs efficiently on CPU.

Why do I need to rebuild my index after changing the embedding model?

MiniLM and Embeddinggemma map text into different vector spaces with incompatible geometries. Searching across mixed embeddings would return meaningless results, so MemPalace requires a full mempalace repair rebuild-index to ensure all vectors are generated by the same embedding function.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →