How to Subclass BaseBackend and BaseCollection for Custom Storage in MemPalace

To create a custom storage backend in MemPalace, inherit from BaseCollection to implement CRUD operations and BaseBackend to act as a factory, then register your implementation via the registry or Python entry points.

MemPalace isolates the storage layer behind two abstract contracts defined in mempalace/backends/base.py. When you subclass BaseBackend and BaseCollection for custom storage in MemPalace, you can plug in any persistence mechanism—from in-memory dictionaries to cloud-native vector databases—without changing the rest of the application.

Understanding the Core Contracts

MemPalace uses a two-level abstraction: a long-lived backend factory that manages collections, and per-collection collection instances that handle actual data operations.

BaseCollection: The Per-Collection API

BaseCollection (lines 33–69 in mempalace/backends/base.py) defines the read/write contract for individual collections. Your subclass must implement every abstract method to satisfy the storage protocol.

The required method signatures are:

def add(self, *, documents: list[str], ids: list[str],
        metadatas: Optional[list[dict]] = None,
        embeddings: Optional[list[list[float]]] = None) -> None: ...

def upsert(self, *, documents: list[str], ids: list[str],
           metadatas: Optional[list[dict]] = None,
           embeddings: Optional[list[list[float]]] = None) -> None: ...

def query(self, *, query_texts: Optional[list[str]] = None,
          query_embeddings: Optional[list[list[float]]] = None,
          n_results: int = 10,
          where: Optional[dict] = None,
          where_document: Optional[dict] = None,
          include: Optional[list[str]] = None) -> QueryResult: ...

def get(self, *, ids: Optional[list[str]] = None,
        where: Optional[dict] = None,
        where_document: Optional[dict] = None,
        limit: Optional[int] = None,
        offset: Optional[int] = None,
        include: Optional[list[str]] = None) -> GetResult: ...

def delete(self, *, ids: Optional[list[str]] = None,
           where: Optional[dict] = None) -> None: ...

def count(self) -> int: ...

query must return a QueryResult and get must return a GetResult. Both are typed dataclasses defined in the same file. You may also override optional methods such as estimated_count, lexical_search, or update if your backend provides native optimizations.

BaseBackend: The Factory and Lifecycle Manager

BaseBackend (lines 65–93 in mempalace/backends/base.py) acts as a factory that instantiates collections. The only required method is get_collection:

def get_collection(self, *, palace: PalaceRef,
                   collection_name: str,
                   create: bool = False,
                   options: Optional[dict] = None) -> BaseCollection: ...

Optional lifecycle hooks include close_palace (per-palace cleanup), close (process-wide shutdown), health (status reporting), and the static detect method for on-disk auto-detection.

Implementing a Custom Collection

When implementing BaseCollection, map each abstract method to your storage engine's native operations. For example, if building a Redis backend, your add method would serialize documents and write them to Redis hashes keyed by the provided ids.

Key implementation details:

  • Result types: QueryResult expects lists of lists (outer list = queries, inner list = results per query) for ids, documents, metadatas, and distances.
  • Metadata filtering: If your backend does not support the where clause natively, implement filtering in Python after retrieval, though this may impact performance.
  • Embeddings: The embeddings parameter in add and upsert may be None if the user expects the backend to compute embeddings; otherwise, store the provided vectors.

Implementing the Backend Factory

The get_collection method receives a PalaceRef object containing the palace identifier and local path. Typical implementation steps include:

  1. Validate the palace path using palace.local_path.
  2. Initialize storage if create=True (e.g., create directories or database tables).
  3. Return a collection instance configured for that specific palace and collection name.

Thread safety is the backend's responsibility. If your storage connection is not thread-safe, initialize a connection pool or use locking inside get_collection.

Registering Your Backend

MemPalace discovers backends through a registry defined in mempalace/backends/registry.py. You have two registration options:

Programmatic registration (ideal for testing or single-file scripts):

from mempalace.backends.registry import register
from my_module import MyBackend

register("mybackend", MyBackend)

Entry-point registration (preferred for distributable packages). Add this to your pyproject.toml:

[project.entry-points."mempalace.backends"]
mybackend = "my_package:MyBackend"

MemPalace automatically discovers entry points at startup (see registry.py lines 56–90).

Complete Minimal Example

Below is a functional in-memory backend that stores collections in module-level dictionaries. It demonstrates all required methods and registration.


# memory_backend.py

import threading
from typing import Optional, Dict, Tuple

from mempalace.backends.base import (
    BaseBackend, BaseCollection, PalaceRef, 
    QueryResult, GetResult, HealthStatus
)
from mempalace.backends.registry import register


class MemoryCollection(BaseCollection):
    """In-memory collection storing documents, metadata, and embeddings."""
    
    def __init__(self, store: Dict[str, Tuple]):
        self._store = store  # id -> (document, metadata, embedding)

    def add(self, *, documents, ids, metadatas=None, embeddings=None):
        for i, did in enumerate(ids):
            self._store[did] = (
                documents[i],
                metadatas[i] if metadatas else {},
                embeddings[i] if embeddings else None,
            )

    def upsert(self, *, documents, ids, metadatas=None, embeddings=None):
        self.add(documents=documents, ids=ids, metadatas=metadatas, embeddings=embeddings)

    def query(self, *, query_texts=None, query_embeddings=None,
              n_results=10, where=None, where_document=None, include=None):
        # Naive implementation: return first n_results

        ids = list(self._store.keys())[:n_results]
        docs = [self._store[i][0] for i in ids]
        metas = [self._store[i][1] for i in ids]
        
        return QueryResult(
            ids=[[i] for i in ids],
            documents=[[d] for d in docs],
            metadatas=[[m] for m in metas],
            distances=[[0.0] for _ in ids],
            embeddings=None,
        )

    def get(self, *, ids=None, where=None, where_document=None,
            limit=None, offset=None, include=None):
        if ids is None:
            ids = list(self._store.keys())
        docs = [self._store[i][0] for i in ids]
        metas = [self._store[i][1] for i in ids]
        embeds = [self._store[i][2] for i in ids]
        
        return GetResult(
            ids=ids,
            documents=docs,
            metadatas=metas,
            embeddings=embeds,
        )

    def delete(self, *, ids=None, where=None):
        if ids:
            for did in ids:
                self._store.pop(did, None)

    def count(self) -> int:
        return len(self._store)


class MemoryBackend(BaseBackend):
    """Factory for in-memory collections."""
    
    name = "memory"
    capabilities = frozenset({"supports_embeddings_in", "supports_metadata_filters"})

    def __init__(self):
        self._stores: Dict[Tuple[str, str], Dict] = {}
        self._lock = threading.Lock()

    def get_collection(self, *, palace: PalaceRef,
                       collection_name: str,
                       create: bool = False,
                       options: Optional[dict] = None) -> BaseCollection:
        key = (palace.id, collection_name)
        with self._lock:
            if key not in self._stores:
                if not create:
                    raise CollectionNotInitializedError(palace.id)
                self._stores[key] = {}
            store = self._stores[key]
        return MemoryCollection(store)


# Register the backend

register("memory", MemoryBackend)

After importing this module, select your backend via:

mempalace --backend memory ...

# or

export MEMPALACE_BACKEND=memory

Summary

  • BaseCollection requires six abstract methods (add, upsert, query, get, delete, count) that must return QueryResult or GetResult dataclasses defined in mempalace/backends/base.py.
  • BaseBackend requires only get_collection to act as a factory, with optional hooks for close_palace, close, health, and static detect for auto-discovery.
  • Register backends programmatically with register() or declaratively via pyproject.toml entry points under the group mempalace.backends.
  • The create parameter in get_collection determines whether to initialize new storage resources or raise errors for missing collections.

Frequently Asked Questions

Do I need to implement all methods in BaseCollection?

Yes. BaseCollection is an abstract base class requiring implementations for add, upsert, query, get, delete, and count. Omitting any will raise TypeError on instantiation. Optional methods like lexical_search provide default implementations that raise NotImplementedError, so you only override them when your storage engine supports the capability natively.

How does MemPalace auto-detect my backend from disk?

Implement the static detect(cls, path: str) -> bool method in your BaseBackend subclass. This method should inspect the directory at path and return True if it contains the specific artifacts (files, subdirectories, or metadata) belonging to your backend. MemPalace calls this during backend resolution in registry.py (lines 68–76) when no explicit backend is specified.

Can I use asynchronous I/O in my backend implementation?

The current MemPalace contracts use synchronous method signatures. If your storage driver is async-only, run the async calls using asyncio.run() or bridge patterns inside the synchronous methods, though this may block the event loop if MemPalace adds async support later. Check the base class definitions in mempalace/backends/base.py for the exact synchronous signatures required.

What is the difference between add and upsert?

add inserts new documents and should raise an error if an ID already exists, while upsert inserts new documents or updates existing ones based on the provided IDs. In the example above, upsert simply delegates to add because the dictionary assignment overwrites existing keys, but production backends should handle this distinction according to their storage semantics.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →