# How to Migrate MemPalace Between ChromaDB Versions Without Data Loss

> Safely migrate MemPalace between ChromaDB versions. Our built-in tool rebuilds your palace extracting data from SQLite, ensuring no data loss. Learn more.

- Repository: [MemPalace/mempalace](https://github.com/MemPalace/mempalace)
- Tags: migration-guide
- Published: 2026-06-07

---

**MemPalace provides a built-in migration tool in [`mempalace/migrate.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/migrate.py) that safely rebuilds your palace by extracting drawer data directly from SQLite, creating a fresh ChromaDB collection compatible with your current library version, and atomically swapping the directories after a timestamped backup.**

When the ChromaDB dependency is upgraded or downgraded, the on-disk SQLite schema and HNSW vector index inside your palace directory can diverge and crash the client. The `mempalace migrate` command avoids the broken ChromaDB API entirely by reading raw SQL records and reconstructing your collection from scratch.

## Why ChromaDB Version Changes Break MemPalace

MemPalace stores every drawer—verbatim text chunks with embeddings and metadata—inside a ChromaDB collection that lives on disk as a directory called the **palace**. Because ChromaDB persists both a SQLite database and native HNSW segment files, any change in the underlying library version risks three distinct failure modes:

- **SQLite schema drift** – Upgrading from ChromaDB 0.6 to 1.x introduces new columns such as `schema_str`; downgrading removes them. Opening the old database with mismatched code raises schema errors or silently drops writes.
- **HNSW metadata incompatibility** – Newer versions write `index_metadata.pickle` files that older libraries cannot deserialize, causing the native Rust segment loader to crash before Python handles the error.
- **Corrupt vector segments** – A damaged `link_lists.bin` or oversized HNSW segment can trigger a segfault in the Rust backend during initialization.

The migration script in [`mempalace/migrate.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/migrate.py) was built to handle all three cases without ever calling the broken ChromaDB collection API.

## How the Migration Tool Works

The migration follows a ten-step pipeline that prioritizes read-only inspection, backup safety, and atomic replacement.

### Detect Source Version and Test Compatibility

First, `detect_chromadb_version()` inspects the palace’s SQLite file directly. It checks the `collections` table for the `schema_str` column, which signals ChromaDB 1.x, and falls back to looking for an `embeddings_queue` table to identify 0.6.x. Next, `collection_write_roundtrip_works()` performs a tiny upsert-delete cycle on the existing collection. If the round-trip fails, the tool knows the collection must be rebuilt from extracted records rather than opened normally.

### Extract Drawers Directly from SQLite

To avoid broken ChromaDB APIs, `extract_drawers_from_sqlite()` runs raw SQL against `chroma.sqlite3` inside the palace. The query pulls every embedding ID, its document content (`chroma:document`), and all associated metadata. Because this bypasses the Chroma client entirely, it works regardless of which ChromaDB version originally created the file.

### Summarize Data and Optional Dry-Run

Before modifying anything, the script builds a wing-room histogram (`defaultdict(lambda: defaultdict(int))`) so you can inspect what will be migrated. Running `mempalace migrate --dry-run` stops after this summary, reporting exactly how many drawers exist per wing and room without touching disk.

### Backup Before Any Changes

If the migration proceeds, the tool creates a timestamped copy of the entire palace directory using `shutil.copytree`. The backup is named `<palace>.pre-migrate.<timestamp>` and is preserved after the migration finishes so you can recover manually if needed.

### Build a Fresh Collection with the Current ChromaDB Version

The migration creates a temporary working directory to avoid stale state. It calls `ChromaBackend().get_or_create_collection()` to build a brand-new collection that uses whatever ChromaDB version is currently installed in the environment. This guarantees the new palace format matches the active library.

### Batch Import and Atomic Swap

Drawers are inserted in batches of 500 via `col.add(...)` to keep memory usage modest. Once every record is imported, the tool renames the old palace to `<palace>.old` and moves the new directory into place with `os.replace`. If the move crosses filesystem boundaries and raises `errno.EXDEV`, a fallback `shutil.move` completes the transfer.

### Rollback Safety

Should any step after the backup fail, `_restore_stale_palace()` cleans up partial swap artifacts. The original timestamped backup remains untouched, ensuring the operator always has a recoverable copy of the pre-migration state.

## Wing-Name Normalization for Legacy Palaces

Older MemPalace releases stored wing names with leading and trailing underscores, such as `_-home-user-`. The `migrate_wing_names()` function in [`mempalace/migrate.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/migrate.py) rewrites the `wing` metadata field to the normalized form—for example, `home_user`—while leaving drawer IDs untouched. It also updates the `topics_by_wing` registry so search results remain consistent after the migration. This step is **idempotent** and runs automatically after the main data migration when legacy naming is detected.

## HNSW Health Checks and Automatic Repairs

The Chroma backend at [`mempalace/backends/chroma.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/backends/chroma.py) contains several low-level helpers that guard against HNSW corruption before and after a migration:

- **`quarantine_stale_hnsw()`** – Renames any segment whose `index_metadata.pickle` fails a byte-sniff validity check or where the ratio of `link_lists.bin` size to `data_level0.bin` size exceeds a 10× safety threshold.
- **`hnsw_capacity_status()`** – Compares the SQLite embedding count with the HNSW element count by reading `index_metadata.pickle` through a tightly controlled unpickler, flagging large divergences that indicate index drift.
- **`_fix_blob_seq_ids()`** and **`_fix_missing_collection_type()`** – Run once before any Chroma client is instantiated to repair legacy SQLite schema issues that would otherwise cause immediate crashes.

These checks are invoked automatically by the MCP server and the `repair` CLI command, ensuring that a palace that survives migration is also safe for future operations.

## Running the Migration CLI

The `migrate` and `migrate-wings` commands are registered in [`mempalace/cli.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/cli.py) and operate on the default palace unless you specify a path.

### Dry-Run a Migration

Use `--dry-run` to inspect the palace without making changes:

```bash
mempalace migrate --dry-run

```

The output prints the detected ChromaDB source version, target version, total drawer count, and a per-wing, per-room histogram.

### Migrate the Default Palace

Run an interactive migration with automatic backup:

```bash
mempalace migrate

```

The tool prompts for confirmation, then shows progress:

```bash
...
Backing up to /home/user/.mempalace.pre-migrate.20240607_154212...
Creating fresh palace in /tmp/mempalace_migrate_abc123...
Imported 5,000/12,845 drawers...
Imported 10,000/12,845 drawers...
Imported 12,845/12,845 drawers...
Swapping old palace for migrated version...
Migration complete.
Drawers migrated: 12,845
Backup at: /home/user/.mempalace.pre-migrate.20240607_154212

```

### Migrate a Specific Palace Non-Interactively

```bash
mempalace migrate --palace /path/to/palace --yes

```

### Normalize Wing Names After Upgrade

If you upgraded from a pre-#1675 release, run the wing-name normalizer:

```bash
mempalace migrate-wings

```

This prints the migration plan, applies the rewrite, and updates the internal registry:

```bash
Wing-name migration plan:
  '-home-user-' -> 'home_user': 3120 drawer(s), 0 closet(s)
  'work-'       -> 'work'      : 9725 drawer(s), 0 closet(s)

Migrated 12,845 drawer(s).

```

## Summary

- MemPalace migration is required because ChromaDB upgrades and downgrades can break SQLite schemas and HNSW vector segment compatibility.
- The migration tool in [`mempalace/migrate.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/migrate.py) reads drawer data directly from SQLite to bypass incompatible ChromaDB APIs.
- A timestamped backup is created before any destructive operation, and an atomic directory swap replaces the old palace.
- HNSW health checks in [`mempalace/backends/chroma.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/backends/chroma.py) automatically quarantine corrupt segments and repair legacy schema issues.
- Use `mempalace migrate --dry-run` to preview changes, and `mempalace migrate-wings` to fix legacy wing-name formatting.

## Frequently Asked Questions

### Can I downgrade ChromaDB after upgrading MemPalace?

Yes. The migration tool detects the source version from the SQLite schema and extracts all data with raw SQL, so it works in both directions. The rebuilt palace always matches the ChromaDB version currently installed in your environment, whether that is newer or older than the original.

### What happens if the migration is interrupted mid-swap?

The migration creates a full timestamped backup with `shutil.copytree` before modifying anything. If the atomic swap fails, `_restore_stale_palace()` cleans up partial artifacts and leaves the original backup intact. You can manually restore the backup directory to resume operations.

### Does the migration tool modify my original drawer IDs?

No. `extract_drawers_from_sqlite()` preserves every embedding ID, document, and metadata field. Only the `wing` metadata value is rewritten if you run `mempalace migrate-wings` to normalize legacy underscore-wrapped names, and even then the drawer IDs themselves remain unchanged.

### How do I know if my HNSW index is corrupt before migrating?

Start the MCP server or run `mempalace repair`. Both invoke `hnsw_capacity_status()` and `quarantine_stale_hnsw()` from [`mempalace/backends/chroma.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/backends/chroma.py). These helpers flag segments where the SQLite embedding count diverges from the HNSW count or where `index_metadata.pickle` and binary size ratios exceed safety thresholds.