How to Configure the Embedding Wrapper for Different Embedder Models in MemPalace

MemPalace abstracts embedding generation behind a unified wrapper that supports multiple ONNX models—MiniLM for fast CPU inference and EmbeddingGemma for multilingual tasks—configurable via environment variables or a JSON config file.

The MemPalace vector memory system delegates all embedding vector creation to a centralized wrapper architecture. Located in mempalace/backends/embedding_wrapper.py, this abstraction allows storage backends like SQLite-exact or pgvector to work with any supported embedder model without code changes. The actual model instantiation occurs through the get_embedding_function factory in mempalace/embedding.py, which handles model selection, device resolution, and caching.

How the Embedding Wrapper Architecture Works

When a backend requires explicit vectors—indicated by the requires_explicit_embeddings flag—MemPalace wraps the underlying collection with an EmbeddingCollection instance.

The wrapper intercepts calls to add() or upsert() where embeddings are omitted. In these cases, it delegates to _embed_texts, which invokes get_embedding_function() to obtain the appropriate embedder【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/backends/embedding_wrapper.py#L10-L18】. This factory function reads configuration from MempalaceConfig, checking for model and device specifications passed via initialization, environment variables, or the ~/.mempalace/config.json file【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/embedding.py#L46-L53】.

Supported Embedder Models

MemPalace currently supports two distinct ONNX-based embedding models, selectable via the embedding_model parameter:

  • MiniLM (minilm): The default model, implemented as a subclass of chromadb.utils.embedding_functions.ONNXMiniLM_L6_V2 named "default". This model offers fast inference on CPU and is ideal for general-purpose English text embedding【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/embedding.py#L10-L18】【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/embedding.py#L110-L118】.

  • EmbeddingGemma (embeddinggemma): A multilingual ONNX model based on onnx-community/embeddinggemma-300m-ONNX. The factory creates an EmbeddinggemmaONNX instance that downloads the model weights lazily on first use, supporting non-English languages and GPU acceleration【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/embedding.py#L30-L38】【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/embedding.py#L40-L66】.

Configuration Methods

You can configure the embedding wrapper through two primary mechanisms, evaluated in order of precedence:

Environment Variables

Set the following variables to override all other configuration:

export MEMPALACE_EMBEDDING_MODEL=embeddinggemma
export MEMPALACE_EMBEDDING_DEVICE=cuda

Valid values for MEMPALACE_EMBEDDING_MODEL are minilm or embeddinggemma. For MEMPALACE_EMBEDDING_DEVICE, use auto, cpu, cuda, coreml, or dml.

Configuration File

For persistent settings across sessions, modify ~/.mempalace/config.json:

{
  "embedding_model": "embeddinggemma",
  "embedding_device": "cuda"
}

The MempalaceConfig class reads these fields when the environment variables or explicit parameters are None【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/embedding.py#L46-L53】.

Device Resolution and ONNX Providers

The _resolve_providers function in mempalace/embedding.py determines the ONNX execution providers based on your device selection:

  1. Auto-detection: When device="auto", the system attempts providers in order: CUDA → CoreML → DirectML (DML) → CPU.
  2. Explicit selection: Maps requests via _PROVIDER_MAP to specific ONNX providers.
  3. Fallback: If the requested provider is unavailable, the system warns once and falls back to CPU【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/embedding.py#L68-L88】.

The factory caches created embedding functions in _EF_CACHE keyed by the tuple (model, providers), ensuring that subsequent requests reuse the already-loaded model without reinitialization【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/embedding.py#L55-L60】【/cache/repos/github.com/MemPalace/mempalace/develop/mempalace/embedding.py#L66-L74】.

Practical Code Examples

Example 1: Using the Default MiniLM Model

from mempalace.backends.embedding_wrapper import EmbeddingCollection
from mempalace.backends.sqlite_exact import SqliteExactCollection

inner = SqliteExactCollection(name="my_palace")
palace = EmbeddingCollection(inner)

# Embeddings generated automatically using MiniLM

palace.add(documents=["Hello world", "Machine learning"], ids=["1", "2"])

Example 2: Switching to EmbeddingGemma via Environment Variables

import os

os.environ["MEMPALACE_EMBEDDING_MODEL"] = "embeddinggemma"
os.environ["MEMPALACE_EMBEDDING_DEVICE"] = "cpu"

from mempalace.backends.embedding_wrapper import EmbeddingCollection
from mempalace.backends.pgvector import PgVectorCollection

inner = PgVectorCollection(name="multilingual_palace")
palace = EmbeddingCollection(inner)

# First call triggers lazy download of the EmbeddingGemma ONNX model

palace.add(documents=["¡Hola!", "你好!", "Bonjour"], ids=["a", "b", "c"])

Example 3: Manual Embedder Instantiation

from mempalace.embedding import get_embedding_function

ef = get_embedding_function(model="embeddinggemma", device="auto")
vectors = ef(["Sample sentence", "Another example"])

# Returns list of 384-dimensional float vectors

Critical Considerations for Model Switching

Changing the embedding_model on an existing palace requires rebuilding the entire index because vector dimensions and semantic spaces differ between models. After modifying the configuration, execute:

mempalace repair rebuild-index

This command regenerates all embeddings using the newly configured model and updates the vector store accordingly.

Summary

  • Architecture: The EmbeddingCollection wrapper in mempalace/backends/embedding_wrapper.py automatically injects embeddings for backends requiring explicit vectors.
  • Models: Choose between MiniLM (fast, CPU-optimized) and EmbeddingGemma (multilingual, GPU-capable) via the embedding_model parameter.
  • Configuration: Set MEMPALACE_EMBEDDING_MODEL and MEMPALACE_EMBEDDING_DEVICE environment variables, or persist choices in ~/.mempalace/config.json.
  • Performance: Embedding functions are cached by (model, providers) tuple to avoid reloading ONNX models unnecessarily.
  • Migration: Switching models necessitates running mempalace repair rebuild-index to re-embed existing documents.

Frequently Asked Questions

What embedding models does MemPalace support?

MemPalace supports two ONNX models: MiniLM (the default, based on all-MiniLM-L6-v2 via ChromaDB's implementation) for efficient CPU inference, and EmbeddingGemma (300M parameter multilingual model) for cross-lingual applications. You select between them using the embedding_model configuration key with values minilm or embeddinggemma.

How do I configure GPU acceleration for embeddings?

Set MEMPALACE_EMBEDDING_DEVICE=cuda (or coreml for Apple Silicon, dml for DirectML on Windows) either as an environment variable or in your config file. When device is set to auto, MemPalace attempts CUDA first, then CoreML, then DirectML, falling back to CPU if none are available.

Do I need to rebuild my data when changing the embedding model?

Yes. Because different models produce vectors in incompatible semantic spaces with varying dimensions, you must run mempalace repair rebuild-index after switching models. This command iterates through all documents and regenerates embeddings using the newly configured embedder.

How does MemPalace cache embedding models?

The get_embedding_function factory stores instantiated embedders in an internal _EF_CACHE dictionary keyed by the tuple (model, providers). This ensures that multiple collections or reconnection attempts reuse the same ONNX model instance in memory, eliminating redundant loading and initialization overhead.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →