# MemPalace Multi-Language Embedding Support and Configuration: A Deep Architectural Guide

> Explore MemPalace multi-language embedding support and configuration. Learn how to leverage the Embeddinggemma ONNX model and configure CPU fallback for efficient NLP.

- Repository: [MemPalace/mempalace](https://github.com/MemPalace/mempalace)
- Tags: deep-dive
- Published: 2026-06-07

---

**MemPalace provides a configurable embedding layer that defaults to the multilingual Embeddinggemma ONNX model and automatically falls back to CPU execution when hardware accelerators are unavailable.**

MemPalace stores every word a user says verbatim and relies on semantic embeddings to enable efficient retrieval. The *develop* branch introduces a flexible embedding system capable of running on CPU or hardware accelerators while supporting multilingual text through a lazy-loaded ONNX model. This article examines the architecture and configuration of MemPalace multi-language embedding support and configuration based on the source code in the `MemPalace/mempalace` repository.

## How Embedding Configuration Works in MemPalace

The MemPalace configuration system exposes two primary settings that control the embedding pipeline.

### Supported Embedding Models

MemPalace supports two distinct embedding functions selectable via the `MEMPALACE_EMBEDDING_MODEL` environment variable or the `embedding_model` key in `~/.mempalace/config.json`:

- **`minilm`** — The historic default based on *all-MiniLM-L6-v2*, producing 384-dimensional vectors optimized for English-only content.
- **`embeddinggemma`** — A multilingual ONNX model (`onnx-community/embeddinggemma-300m-ONNX`, quantized q8) that supports over 100 languages. After Matryoshka truncation, it yields 384-dimensional vectors, making it compatible with existing ChromaDB collections without schema changes.

According to the source code in [[`mempalace/embedding.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/embedding.py)](https://github.com/MemPalace/mempalace/blob/develop/mempalace/embedding.py#L6-L18), cross-lingual cosine similarity for the multilingual model measures approximately 0.88 compared to roughly 0.35 for MiniLM.

### Device and Execution Provider Selection

The `MEMPALACE_EMBEDDING_DEVICE` setting (or `embedding_device` in config) controls hardware acceleration. The `_resolve_providers` helper in [[`mempalace/embedding.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/embedding.py)](https://github.com/MemPalace/mempalace/blob/develop/mempalace/embedding.py#L62-L81) inspects installed `onnxruntime` providers and resolves the following priority:

1. CUDA
2. CoreML
3. DirectML
4. CPU

If the requested accelerator is unavailable, the system falls back to CPU with a one-time warning. This ensures mining works on any laptop regardless of GPU availability.

## Core Architecture of the Embedding Layer

### The Embedding Factory

[[`mempalace/embedding.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/embedding.py)](https://github.com/MemPalace/mempalace/blob/develop/mempalace/embedding.py) serves as the public factory that instantiates the correct embedding function based on user configuration. Lines 6–18 define the supported models, while lines 110–127 create a Chroma-compatible subclass for the default MiniLM embedding function. The module reads configuration values set in [[`mempalace/config.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/config.py)](https://github.com/MemPalace/mempalace/blob/develop/mempalace/config.py#L610-L628) through getters exposed at lines 610–628.

### Multilingual Model Implementation

The `EmbeddinggemmaONNX` class implements the multilingual backend. Its `name()` method returns the stable identifier `"embeddinggemma_300m"`, which ChromaDB stores on the collection to detect mismatched embedding functions. The class lazily loads the model, tokenizer, and ONNX session on first call, downloading the artifact via `huggingface_hub` if not already cached.

This implementation is defined at lines 54–70 of [[`mempalace/embedding.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/embedding.py)](https://github.com/MemPalace/mempalace/blob/develop/mempalace/embedding.py#L54-L70), with lazy loading and download logic spanning lines 73–99.

### Backend Integration

[[`mempalace/backends/embedding_wrapper.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/backends/embedding_wrapper.py)](https://github.com/MemPalace/mempalace/blob/develop/mempalace/backends/embedding_wrapper.py#L38) defines `EmbeddingCollection`, a thin wrapper that makes any embedding function conform to the ChromaDB collection interface. The core palace class in [[`mempalace/palace.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/palace.py)](https://github.com/MemPalace/mempalace/blob/develop/mempalace/palace.py) instantiates this wrapper at startup, injecting the user-selected embedding function.

### Onboarding and Persistent Configuration

When MemPalace runs for the first time, [[`mempalace/onboarding.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/onboarding.py)](https://github.com/MemPalace/mempalace/blob/develop/mempalace/onboarding.py) prompts the user to select an embedding model, defaulting new installations to the multilingual option. The selection is persisted through [[`mempalace/config.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/config.py)](https://github.com/MemPalace/mempalace/blob/develop/mempalace/config.py#L630), which provides the setter at line 630.

## Practical Configuration Examples

### Configure the Embedding Model via Python

You can programmatically switch to the multilingual model using the configuration API:

```python
from mempalace.config import Config

cfg = Config()
cfg.set_embedding_model("embeddinggemma")
cfg.set_embedding_device("auto")

```

This persists the values to `~/.mempalace/config.json` and will be respected across subsequent sessions.

### Change Models Using the CLI

The MemPalace CLI exposes commands to inspect and modify embedding settings:

```bash

# Display current embedding configuration

mempalace config show

# Switch to the multilingual model

mempalace config set embedding_model embeddinggemma

# Re-index existing data after switching models

mempalace repair rebuild-index

```

Because MiniLM and Embeddinggemma operate in different vector spaces, you must run `mempalace repair rebuild-index` after changing the model.

### Use the Embedding Function Directly

For advanced integrations, instantiate the multilingual embedding function directly:

```python
from mempalace.embedding import EmbeddinggemmaONNX, _resolve_providers

# Resolve the best available provider list

providers, device = _resolve_providers("auto")

# Initialize the embedding function (triggers lazy download on first call)

ef = EmbeddinggemmaONNX(preferred_providers=providers)

# Generate a 384-dimensional vector for non-English text

vector = ef(["task: sentence similarity | query: こんにちは世界"])
print(vector.shape)  # (1, 384)

```

### Search with the Configured Embedding Model

Once configured, the `MemPalace` class automatically loads the correct embedding function:

```python
from mempalace import MemPalace

pal = MemPalace()
pal.add_documents(["Hello world"], ids=["doc1"])

# Query using the same multilingual embedding space

results = pal.query("こんにちは", top_k=5)
print(results)

```

## Key Source Files

- **[[`mempalace/embedding.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/embedding.py)](https://github.com/MemPalace/mempalace/blob/develop/mempalace/embedding.py)** — Factory, provider resolution via `_resolve_providers`, and the `EmbeddinggemmaONNX` implementation.
- **[[`mempalace/config.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/config.py)](https://github.com/MemPalace/mempalace/blob/develop/mempalace/config.py#L610-L630)** — Reads and writes `embedding_model` and `embedding_device` properties.
- **[[`mempalace/backends/embedding_wrapper.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/backends/embedding_wrapper.py)](https://github.com/MemPalace/mempalace/blob/develop/mempalace/backends/embedding_wrapper.py#L38)** — Wraps embedding functions as ChromaDB-compatible collections.
- **[[`mempalace/onboarding.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/onboarding.py)](https://github.com/MemPalace/mempalace/blob/develop/mempalace/onboarding.py)** — First-run prompt that defaults new users to the multilingual model.
- **[[`mempalace/palace.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/palace.py)](https://github.com/MemPalace/mempalace/blob/develop/mempalace/palace.py)** — Core palace class that wires the selected embedding function into the storage backend.

## Summary

- MemPalace supports two embedding models: the English-only **MiniLM** and the multilingual **Embeddinggemma** ONNX model.
- **Device selection is automatic**, falling back from CUDA to CoreML to DirectML to CPU.
- The **`EmbeddinggemmaONNX`** class lazily downloads and caches the multilingual model via `huggingface_hub`.
- Configuration is managed through **environment variables**, the **CLI**, or **`~/.mempalace/config.json`** via [`config.py`](https://github.com/MemPalace/mempalace/blob/main/config.py).
- Existing palaces must be **re-indexed** with `mempalace repair rebuild-index` after changing embedding models.
- The architecture separates concerns between factory logic, ONNX inference, ChromaDB wrapper compatibility, and user onboarding.

## Frequently Asked Questions

### What embedding models does MemPalace support?

MemPalace supports `minilm` (all-MiniLM-L6-v2, English-only, 384-dim) and `embeddinggemma` (multilingual ONNX model supporting over 100 languages, also 384-dim after Matryoshka truncation). The multilingual model is available on the *develop* branch and is the default for new installations.

### How do I switch from MiniLM to the multilingual embedding model?

You can switch by running `mempalace config set embedding_model embeddinggemma` in the CLI or by calling `cfg.set_embedding_model("embeddinggemma")` on a `Config` instance in Python. After switching, execute `mempalace repair rebuild-index` to re-embed existing documents in the new vector space.

### Does MemPalace require a GPU for multilingual embeddings?

No. The `_resolve_providers` function in [`mempalace/embedding.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/embedding.py) automatically falls back to CPU execution with a one-time warning if CUDA, CoreML, or DirectML are unavailable. The multilingual ONNX model runs efficiently on CPU.

### Why do I need to rebuild my index after changing the embedding model?

MiniLM and Embeddinggemma map text into different vector spaces with incompatible geometries. Searching across mixed embeddings would return meaningless results, so MemPalace requires a full `mempalace repair rebuild-index` to ensure all vectors are generated by the same embedding function.