How to Migrate MemPalace Between ChromaDB Versions Without Data Loss

MemPalace provides a built-in migration tool in mempalace/migrate.py that safely rebuilds your palace by extracting drawer data directly from SQLite, creating a fresh ChromaDB collection compatible with your current library version, and atomically swapping the directories after a timestamped backup.

When the ChromaDB dependency is upgraded or downgraded, the on-disk SQLite schema and HNSW vector index inside your palace directory can diverge and crash the client. The mempalace migrate command avoids the broken ChromaDB API entirely by reading raw SQL records and reconstructing your collection from scratch.

Why ChromaDB Version Changes Break MemPalace

MemPalace stores every drawer—verbatim text chunks with embeddings and metadata—inside a ChromaDB collection that lives on disk as a directory called the palace. Because ChromaDB persists both a SQLite database and native HNSW segment files, any change in the underlying library version risks three distinct failure modes:

  • SQLite schema drift – Upgrading from ChromaDB 0.6 to 1.x introduces new columns such as schema_str; downgrading removes them. Opening the old database with mismatched code raises schema errors or silently drops writes.
  • HNSW metadata incompatibility – Newer versions write index_metadata.pickle files that older libraries cannot deserialize, causing the native Rust segment loader to crash before Python handles the error.
  • Corrupt vector segments – A damaged link_lists.bin or oversized HNSW segment can trigger a segfault in the Rust backend during initialization.

The migration script in mempalace/migrate.py was built to handle all three cases without ever calling the broken ChromaDB collection API.

How the Migration Tool Works

The migration follows a ten-step pipeline that prioritizes read-only inspection, backup safety, and atomic replacement.

Detect Source Version and Test Compatibility

First, detect_chromadb_version() inspects the palace’s SQLite file directly. It checks the collections table for the schema_str column, which signals ChromaDB 1.x, and falls back to looking for an embeddings_queue table to identify 0.6.x. Next, collection_write_roundtrip_works() performs a tiny upsert-delete cycle on the existing collection. If the round-trip fails, the tool knows the collection must be rebuilt from extracted records rather than opened normally.

Extract Drawers Directly from SQLite

To avoid broken ChromaDB APIs, extract_drawers_from_sqlite() runs raw SQL against chroma.sqlite3 inside the palace. The query pulls every embedding ID, its document content (chroma:document), and all associated metadata. Because this bypasses the Chroma client entirely, it works regardless of which ChromaDB version originally created the file.

Summarize Data and Optional Dry-Run

Before modifying anything, the script builds a wing-room histogram (defaultdict(lambda: defaultdict(int))) so you can inspect what will be migrated. Running mempalace migrate --dry-run stops after this summary, reporting exactly how many drawers exist per wing and room without touching disk.

Backup Before Any Changes

If the migration proceeds, the tool creates a timestamped copy of the entire palace directory using shutil.copytree. The backup is named <palace>.pre-migrate.<timestamp> and is preserved after the migration finishes so you can recover manually if needed.

Build a Fresh Collection with the Current ChromaDB Version

The migration creates a temporary working directory to avoid stale state. It calls ChromaBackend().get_or_create_collection() to build a brand-new collection that uses whatever ChromaDB version is currently installed in the environment. This guarantees the new palace format matches the active library.

Batch Import and Atomic Swap

Drawers are inserted in batches of 500 via col.add(...) to keep memory usage modest. Once every record is imported, the tool renames the old palace to <palace>.old and moves the new directory into place with os.replace. If the move crosses filesystem boundaries and raises errno.EXDEV, a fallback shutil.move completes the transfer.

Rollback Safety

Should any step after the backup fail, _restore_stale_palace() cleans up partial swap artifacts. The original timestamped backup remains untouched, ensuring the operator always has a recoverable copy of the pre-migration state.

Wing-Name Normalization for Legacy Palaces

Older MemPalace releases stored wing names with leading and trailing underscores, such as _-home-user-. The migrate_wing_names() function in mempalace/migrate.py rewrites the wing metadata field to the normalized form—for example, home_user—while leaving drawer IDs untouched. It also updates the topics_by_wing registry so search results remain consistent after the migration. This step is idempotent and runs automatically after the main data migration when legacy naming is detected.

HNSW Health Checks and Automatic Repairs

The Chroma backend at mempalace/backends/chroma.py contains several low-level helpers that guard against HNSW corruption before and after a migration:

  • quarantine_stale_hnsw() – Renames any segment whose index_metadata.pickle fails a byte-sniff validity check or where the ratio of link_lists.bin size to data_level0.bin size exceeds a 10× safety threshold.
  • hnsw_capacity_status() – Compares the SQLite embedding count with the HNSW element count by reading index_metadata.pickle through a tightly controlled unpickler, flagging large divergences that indicate index drift.
  • _fix_blob_seq_ids() and _fix_missing_collection_type() – Run once before any Chroma client is instantiated to repair legacy SQLite schema issues that would otherwise cause immediate crashes.

These checks are invoked automatically by the MCP server and the repair CLI command, ensuring that a palace that survives migration is also safe for future operations.

Running the Migration CLI

The migrate and migrate-wings commands are registered in mempalace/cli.py and operate on the default palace unless you specify a path.

Dry-Run a Migration

Use --dry-run to inspect the palace without making changes:

mempalace migrate --dry-run

The output prints the detected ChromaDB source version, target version, total drawer count, and a per-wing, per-room histogram.

Migrate the Default Palace

Run an interactive migration with automatic backup:

mempalace migrate

The tool prompts for confirmation, then shows progress:

...
Backing up to /home/user/.mempalace.pre-migrate.20240607_154212...
Creating fresh palace in /tmp/mempalace_migrate_abc123...
Imported 5,000/12,845 drawers...
Imported 10,000/12,845 drawers...
Imported 12,845/12,845 drawers...
Swapping old palace for migrated version...
Migration complete.
Drawers migrated: 12,845
Backup at: /home/user/.mempalace.pre-migrate.20240607_154212

Migrate a Specific Palace Non-Interactively

mempalace migrate --palace /path/to/palace --yes

Normalize Wing Names After Upgrade

If you upgraded from a pre-#1675 release, run the wing-name normalizer:

mempalace migrate-wings

This prints the migration plan, applies the rewrite, and updates the internal registry:

Wing-name migration plan:
  '-home-user-' -> 'home_user': 3120 drawer(s), 0 closet(s)
  'work-'       -> 'work'      : 9725 drawer(s), 0 closet(s)

Migrated 12,845 drawer(s).

Summary

  • MemPalace migration is required because ChromaDB upgrades and downgrades can break SQLite schemas and HNSW vector segment compatibility.
  • The migration tool in mempalace/migrate.py reads drawer data directly from SQLite to bypass incompatible ChromaDB APIs.
  • A timestamped backup is created before any destructive operation, and an atomic directory swap replaces the old palace.
  • HNSW health checks in mempalace/backends/chroma.py automatically quarantine corrupt segments and repair legacy schema issues.
  • Use mempalace migrate --dry-run to preview changes, and mempalace migrate-wings to fix legacy wing-name formatting.

Frequently Asked Questions

Can I downgrade ChromaDB after upgrading MemPalace?

Yes. The migration tool detects the source version from the SQLite schema and extracts all data with raw SQL, so it works in both directions. The rebuilt palace always matches the ChromaDB version currently installed in your environment, whether that is newer or older than the original.

What happens if the migration is interrupted mid-swap?

The migration creates a full timestamped backup with shutil.copytree before modifying anything. If the atomic swap fails, _restore_stale_palace() cleans up partial artifacts and leaves the original backup intact. You can manually restore the backup directory to resume operations.

Does the migration tool modify my original drawer IDs?

No. extract_drawers_from_sqlite() preserves every embedding ID, document, and metadata field. Only the wing metadata value is rewritten if you run mempalace migrate-wings to normalize legacy underscore-wrapped names, and even then the drawer IDs themselves remain unchanged.

How do I know if my HNSW index is corrupt before migrating?

Start the MCP server or run mempalace repair. Both invoke hnsw_capacity_status() and quarantine_stale_hnsw() from mempalace/backends/chroma.py. These helpers flag segments where the SQLite embedding count diverges from the HNSW count or where index_metadata.pickle and binary size ratios exceed safety thresholds.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →