# How to Subclass BaseBackend and BaseCollection for Custom Storage in MemPalace

> Learn to subclass BaseBackend and BaseCollection for custom storage in MemPalace. Implement CRUD operations and a factory, then register your custom storage solution.

- Repository: [MemPalace/mempalace](https://github.com/MemPalace/mempalace)
- Tags: how-to-guide
- Published: 2026-06-06

---

**To create a custom storage backend in MemPalace, inherit from `BaseCollection` to implement CRUD operations and `BaseBackend` to act as a factory, then register your implementation via the registry or Python entry points.**

MemPalace isolates the storage layer behind two abstract contracts defined in [`mempalace/backends/base.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/backends/base.py). When you subclass `BaseBackend` and `BaseCollection` for custom storage in MemPalace, you can plug in any persistence mechanism—from in-memory dictionaries to cloud-native vector databases—without changing the rest of the application.

## Understanding the Core Contracts

MemPalace uses a two-level abstraction: a long-lived **backend** factory that manages collections, and per-collection **collection** instances that handle actual data operations.

### BaseCollection: The Per-Collection API

`BaseCollection` (lines 33–69 in [`mempalace/backends/base.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/backends/base.py)) defines the read/write contract for individual collections. Your subclass must implement every abstract method to satisfy the storage protocol.

The required method signatures are:

```python
def add(self, *, documents: list[str], ids: list[str],
        metadatas: Optional[list[dict]] = None,
        embeddings: Optional[list[list[float]]] = None) -> None: ...

def upsert(self, *, documents: list[str], ids: list[str],
           metadatas: Optional[list[dict]] = None,
           embeddings: Optional[list[list[float]]] = None) -> None: ...

def query(self, *, query_texts: Optional[list[str]] = None,
          query_embeddings: Optional[list[list[float]]] = None,
          n_results: int = 10,
          where: Optional[dict] = None,
          where_document: Optional[dict] = None,
          include: Optional[list[str]] = None) -> QueryResult: ...

def get(self, *, ids: Optional[list[str]] = None,
        where: Optional[dict] = None,
        where_document: Optional[dict] = None,
        limit: Optional[int] = None,
        offset: Optional[int] = None,
        include: Optional[list[str]] = None) -> GetResult: ...

def delete(self, *, ids: Optional[list[str]] = None,
           where: Optional[dict] = None) -> None: ...

def count(self) -> int: ...

```

`query` must return a `QueryResult` and `get` must return a `GetResult`. Both are typed dataclasses defined in the same file. You may also override optional methods such as `estimated_count`, `lexical_search`, or `update` if your backend provides native optimizations.

### BaseBackend: The Factory and Lifecycle Manager

`BaseBackend` (lines 65–93 in [`mempalace/backends/base.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/backends/base.py)) acts as a factory that instantiates collections. The only required method is `get_collection`:

```python
def get_collection(self, *, palace: PalaceRef,
                   collection_name: str,
                   create: bool = False,
                   options: Optional[dict] = None) -> BaseCollection: ...

```

Optional lifecycle hooks include `close_palace` (per-palace cleanup), `close` (process-wide shutdown), `health` (status reporting), and the static `detect` method for on-disk auto-detection.

## Implementing a Custom Collection

When implementing `BaseCollection`, map each abstract method to your storage engine's native operations. For example, if building a Redis backend, your `add` method would serialize documents and write them to Redis hashes keyed by the provided `ids`.

Key implementation details:

- **Result types**: `QueryResult` expects lists of lists (outer list = queries, inner list = results per query) for `ids`, `documents`, `metadatas`, and `distances`.
- **Metadata filtering**: If your backend does not support the `where` clause natively, implement filtering in Python after retrieval, though this may impact performance.
- **Embeddings**: The `embeddings` parameter in `add` and `upsert` may be `None` if the user expects the backend to compute embeddings; otherwise, store the provided vectors.

## Implementing the Backend Factory

The `get_collection` method receives a `PalaceRef` object containing the palace identifier and local path. Typical implementation steps include:

1. **Validate the palace path** using `palace.local_path`.
2. **Initialize storage** if `create=True` (e.g., create directories or database tables).
3. **Return a collection instance** configured for that specific palace and collection name.

Thread safety is the backend's responsibility. If your storage connection is not thread-safe, initialize a connection pool or use locking inside `get_collection`.

## Registering Your Backend

MemPalace discovers backends through a registry defined in [`mempalace/backends/registry.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/backends/registry.py). You have two registration options:

**Programmatic registration** (ideal for testing or single-file scripts):

```python
from mempalace.backends.registry import register
from my_module import MyBackend

register("mybackend", MyBackend)

```

**Entry-point registration** (preferred for distributable packages). Add this to your [`pyproject.toml`](https://github.com/MemPalace/mempalace/blob/main/pyproject.toml):

```toml
[project.entry-points."mempalace.backends"]
mybackend = "my_package:MyBackend"

```

MemPalace automatically discovers entry points at startup (see registry.py lines 56–90).

## Complete Minimal Example

Below is a functional in-memory backend that stores collections in module-level dictionaries. It demonstrates all required methods and registration.

```python

# memory_backend.py

import threading
from typing import Optional, Dict, Tuple

from mempalace.backends.base import (
    BaseBackend, BaseCollection, PalaceRef, 
    QueryResult, GetResult, HealthStatus
)
from mempalace.backends.registry import register


class MemoryCollection(BaseCollection):
    """In-memory collection storing documents, metadata, and embeddings."""
    
    def __init__(self, store: Dict[str, Tuple]):
        self._store = store  # id -> (document, metadata, embedding)

    def add(self, *, documents, ids, metadatas=None, embeddings=None):
        for i, did in enumerate(ids):
            self._store[did] = (
                documents[i],
                metadatas[i] if metadatas else {},
                embeddings[i] if embeddings else None,
            )

    def upsert(self, *, documents, ids, metadatas=None, embeddings=None):
        self.add(documents=documents, ids=ids, metadatas=metadatas, embeddings=embeddings)

    def query(self, *, query_texts=None, query_embeddings=None,
              n_results=10, where=None, where_document=None, include=None):
        # Naive implementation: return first n_results

        ids = list(self._store.keys())[:n_results]
        docs = [self._store[i][0] for i in ids]
        metas = [self._store[i][1] for i in ids]
        
        return QueryResult(
            ids=[[i] for i in ids],
            documents=[[d] for d in docs],
            metadatas=[[m] for m in metas],
            distances=[[0.0] for _ in ids],
            embeddings=None,
        )

    def get(self, *, ids=None, where=None, where_document=None,
            limit=None, offset=None, include=None):
        if ids is None:
            ids = list(self._store.keys())
        docs = [self._store[i][0] for i in ids]
        metas = [self._store[i][1] for i in ids]
        embeds = [self._store[i][2] for i in ids]
        
        return GetResult(
            ids=ids,
            documents=docs,
            metadatas=metas,
            embeddings=embeds,
        )

    def delete(self, *, ids=None, where=None):
        if ids:
            for did in ids:
                self._store.pop(did, None)

    def count(self) -> int:
        return len(self._store)


class MemoryBackend(BaseBackend):
    """Factory for in-memory collections."""
    
    name = "memory"
    capabilities = frozenset({"supports_embeddings_in", "supports_metadata_filters"})

    def __init__(self):
        self._stores: Dict[Tuple[str, str], Dict] = {}
        self._lock = threading.Lock()

    def get_collection(self, *, palace: PalaceRef,
                       collection_name: str,
                       create: bool = False,
                       options: Optional[dict] = None) -> BaseCollection:
        key = (palace.id, collection_name)
        with self._lock:
            if key not in self._stores:
                if not create:
                    raise CollectionNotInitializedError(palace.id)
                self._stores[key] = {}
            store = self._stores[key]
        return MemoryCollection(store)


# Register the backend

register("memory", MemoryBackend)

```

After importing this module, select your backend via:

```bash
mempalace --backend memory ...

# or

export MEMPALACE_BACKEND=memory

```

## Summary

- **BaseCollection** requires six abstract methods (`add`, `upsert`, `query`, `get`, `delete`, `count`) that must return `QueryResult` or `GetResult` dataclasses defined in [`mempalace/backends/base.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/backends/base.py).
- **BaseBackend** requires only `get_collection` to act as a factory, with optional hooks for `close_palace`, `close`, `health`, and static `detect` for auto-discovery.
- Register backends programmatically with `register()` or declaratively via [`pyproject.toml`](https://github.com/MemPalace/mempalace/blob/main/pyproject.toml) entry points under the group `mempalace.backends`.
- The `create` parameter in `get_collection` determines whether to initialize new storage resources or raise errors for missing collections.

## Frequently Asked Questions

### Do I need to implement all methods in BaseCollection?

Yes. `BaseCollection` is an abstract base class requiring implementations for `add`, `upsert`, `query`, `get`, `delete`, and `count`. Omitting any will raise `TypeError` on instantiation. Optional methods like `lexical_search` provide default implementations that raise `NotImplementedError`, so you only override them when your storage engine supports the capability natively.

### How does MemPalace auto-detect my backend from disk?

Implement the static `detect(cls, path: str) -> bool` method in your `BaseBackend` subclass. This method should inspect the directory at `path` and return `True` if it contains the specific artifacts (files, subdirectories, or metadata) belonging to your backend. MemPalace calls this during backend resolution in [`registry.py`](https://github.com/MemPalace/mempalace/blob/main/registry.py) (lines 68–76) when no explicit backend is specified.

### Can I use asynchronous I/O in my backend implementation?

The current MemPalace contracts use synchronous method signatures. If your storage driver is async-only, run the async calls using `asyncio.run()` or bridge patterns inside the synchronous methods, though this may block the event loop if MemPalace adds async support later. Check the base class definitions in [`mempalace/backends/base.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/backends/base.py) for the exact synchronous signatures required.

### What is the difference between add and upsert?

`add` inserts new documents and should raise an error if an ID already exists, while `upsert` inserts new documents or updates existing ones based on the provided IDs. In the example above, `upsert` simply delegates to `add` because the dictionary assignment overwrites existing keys, but production backends should handle this distinction according to their storage semantics.