# How to Configure the Embedding Wrapper for Different Embedder Models in MemPalace

> Learn to configure the MemPalace embedding wrapper for diverse ONNX models like MiniLM and EmbeddingGemma using environment variables or a config file for efficient embedding generation.

- Repository: [MemPalace/mempalace](https://github.com/MemPalace/mempalace)
- Tags: how-to-guide
- Published: 2026-06-06

---

**MemPalace abstracts embedding generation behind a unified wrapper that supports multiple ONNX models—MiniLM for fast CPU inference and EmbeddingGemma for multilingual tasks—configurable via environment variables or a JSON config file.**

The MemPalace vector memory system delegates all embedding vector creation to a centralized wrapper architecture. Located in [`mempalace/backends/embedding_wrapper.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/backends/embedding_wrapper.py), this abstraction allows storage backends like SQLite-exact or pgvector to work with any supported embedder model without code changes. The actual model instantiation occurs through the `get_embedding_function` factory in [`mempalace/embedding.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/embedding.py), which handles model selection, device resolution, and caching.

## How the Embedding Wrapper Architecture Works

When a backend requires explicit vectors—indicated by the `requires_explicit_embeddings` flag—MemPalace wraps the underlying collection with an `EmbeddingCollection` instance.

The wrapper intercepts calls to `add()` or `upsert()` where embeddings are omitted. In these cases, it delegates to `_embed_texts`, which invokes `get_embedding_function()` to obtain the appropriate embedder【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/backends/embedding_wrapper.py#L10-L18】. This factory function reads configuration from `MempalaceConfig`, checking for model and device specifications passed via initialization, environment variables, or the `~/.mempalace/config.json` file【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/embedding.py#L46-L53】.

## Supported Embedder Models

MemPalace currently supports two distinct ONNX-based embedding models, selectable via the `embedding_model` parameter:

- **MiniLM (`minilm`)**: The default model, implemented as a subclass of `chromadb.utils.embedding_functions.ONNXMiniLM_L6_V2` named "default". This model offers fast inference on CPU and is ideal for general-purpose English text embedding【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/embedding.py#L10-L18】【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/embedding.py#L110-L118】.

- **EmbeddingGemma (`embeddinggemma`)**: A multilingual ONNX model based on `onnx-community/embeddinggemma-300m-ONNX`. The factory creates an `EmbeddinggemmaONNX` instance that downloads the model weights lazily on first use, supporting non-English languages and GPU acceleration【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/embedding.py#L30-L38】【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/embedding.py#L40-L66】.

## Configuration Methods

You can configure the embedding wrapper through two primary mechanisms, evaluated in order of precedence:

### Environment Variables

Set the following variables to override all other configuration:

```bash
export MEMPALACE_EMBEDDING_MODEL=embeddinggemma
export MEMPALACE_EMBEDDING_DEVICE=cuda

```

Valid values for `MEMPALACE_EMBEDDING_MODEL` are `minilm` or `embeddinggemma`. For `MEMPALACE_EMBEDDING_DEVICE`, use `auto`, `cpu`, `cuda`, `coreml`, or `dml`.

### Configuration File

For persistent settings across sessions, modify `~/.mempalace/config.json`:

```json
{
  "embedding_model": "embeddinggemma",
  "embedding_device": "cuda"
}

```

The `MempalaceConfig` class reads these fields when the environment variables or explicit parameters are `None`【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/embedding.py#L46-L53】.

## Device Resolution and ONNX Providers

The `_resolve_providers` function in [`mempalace/embedding.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/embedding.py) determines the ONNX execution providers based on your device selection:

1. **Auto-detection**: When `device="auto"`, the system attempts providers in order: CUDA → CoreML → DirectML (DML) → CPU.
2. **Explicit selection**: Maps requests via `_PROVIDER_MAP` to specific ONNX providers.
3. **Fallback**: If the requested provider is unavailable, the system warns once and falls back to CPU【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/embedding.py#L68-L88】.

The factory caches created embedding functions in `_EF_CACHE` keyed by the tuple `(model, providers)`, ensuring that subsequent requests reuse the already-loaded model without reinitialization【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/embedding.py#L55-L60】【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/embedding.py#L66-L74】.

## Practical Code Examples

### Example 1: Using the Default MiniLM Model

```python
from mempalace.backends.embedding_wrapper import EmbeddingCollection
from mempalace.backends.sqlite_exact import SqliteExactCollection

inner = SqliteExactCollection(name="my_palace")
palace = EmbeddingCollection(inner)

# Embeddings generated automatically using MiniLM

palace.add(documents=["Hello world", "Machine learning"], ids=["1", "2"])

```

### Example 2: Switching to EmbeddingGemma via Environment Variables

```python
import os

os.environ["MEMPALACE_EMBEDDING_MODEL"] = "embeddinggemma"
os.environ["MEMPALACE_EMBEDDING_DEVICE"] = "cpu"

from mempalace.backends.embedding_wrapper import EmbeddingCollection
from mempalace.backends.pgvector import PgVectorCollection

inner = PgVectorCollection(name="multilingual_palace")
palace = EmbeddingCollection(inner)

# First call triggers lazy download of the EmbeddingGemma ONNX model

palace.add(documents=["¡Hola!", "你好！", "Bonjour"], ids=["a", "b", "c"])

```

### Example 3: Manual Embedder Instantiation

```python
from mempalace.embedding import get_embedding_function

ef = get_embedding_function(model="embeddinggemma", device="auto")
vectors = ef(["Sample sentence", "Another example"])

# Returns list of 384-dimensional float vectors

```

## Critical Considerations for Model Switching

Changing the `embedding_model` on an existing palace requires rebuilding the entire index because vector dimensions and semantic spaces differ between models. After modifying the configuration, execute:

```bash
mempalace repair rebuild-index

```

This command regenerates all embeddings using the newly configured model and updates the vector store accordingly.

## Summary

- **Architecture**: The `EmbeddingCollection` wrapper in [`mempalace/backends/embedding_wrapper.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/backends/embedding_wrapper.py) automatically injects embeddings for backends requiring explicit vectors.
- **Models**: Choose between **MiniLM** (fast, CPU-optimized) and **EmbeddingGemma** (multilingual, GPU-capable) via the `embedding_model` parameter.
- **Configuration**: Set `MEMPALACE_EMBEDDING_MODEL` and `MEMPALACE_EMBEDDING_DEVICE` environment variables, or persist choices in `~/.mempalace/config.json`.
- **Performance**: Embedding functions are cached by `(model, providers)` tuple to avoid reloading ONNX models unnecessarily.
- **Migration**: Switching models necessitates running `mempalace repair rebuild-index` to re-embed existing documents.

## Frequently Asked Questions

### What embedding models does MemPalace support?

MemPalace supports two ONNX models: **MiniLM** (the default, based on `all-MiniLM-L6-v2` via ChromaDB's implementation) for efficient CPU inference, and **EmbeddingGemma** (300M parameter multilingual model) for cross-lingual applications. You select between them using the `embedding_model` configuration key with values `minilm` or `embeddinggemma`.

### How do I configure GPU acceleration for embeddings?

Set `MEMPALACE_EMBEDDING_DEVICE=cuda` (or `coreml` for Apple Silicon, `dml` for DirectML on Windows) either as an environment variable or in your config file. When `device` is set to `auto`, MemPalace attempts CUDA first, then CoreML, then DirectML, falling back to CPU if none are available.

### Do I need to rebuild my data when changing the embedding model?

Yes. Because different models produce vectors in incompatible semantic spaces with varying dimensions, you must run `mempalace repair rebuild-index` after switching models. This command iterates through all documents and regenerates embeddings using the newly configured embedder.

### How does MemPalace cache embedding models?

The `get_embedding_function` factory stores instantiated embedders in an internal `_EF_CACHE` dictionary keyed by the tuple `(model, providers)`. This ensures that multiple collections or reconnection attempts reuse the same ONNX model instance in memory, eliminating redundant loading and initialization overhead.