# How the MemPalace Migrate Module Handles ChromaDB Version Migration

> Learn how MemPalace's migrate module safely handles ChromaDB version migration. It detects schema versions, verifies compatibility, and performs atomic swaps with full backups.

- Repository: [MemPalace/mempalace](https://github.com/MemPalace/mempalace)
- Tags: migration-guide
- Published: 2026-06-06

---

**MemPalace's `migrate` module enables safe upgrades and downgrades across ChromaDB schema versions by detecting the source database format, verifying write compatibility, and performing atomic directory swaps with full backup protection.**

When ChromaDB releases introduce breaking schema changes, existing MemPalace databases require careful migration to remain accessible. The `migrate` module in the [MemPalace/mempalace](https://github.com/MemPalace/mempalace) repository provides a robust **ChromaDB version migration** pipeline that preserves your palace data while transitioning between database formats.

## How the Migration Pipeline Works

The migration process operates in three distinct stages, each designed to minimize data loss risk while handling breaking changes between ChromaDB 0.6.x and 1.x releases.

### Detecting the Source Database Version

In [`mempalace/migrate.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/migrate.py), the `detect_chromadb_version()` function inspects the raw SQLite schema to determine whether a palace was created with ChromaDB 0.6.x or 1.x. This detection examines the `collections` table columns and checks for the presence of an `embeddings_queue` table, allowing the module to select the appropriate extraction strategy without relying on potentially incompatible API calls.

### Verifying Current Backend Compatibility

Before attempting data extraction, the module verifies whether the current ChromaDB installation can already read and write to the existing collection. The `collection_write_roundtrip_works()` function performs a complete read-write-delete cycle to confirm compatibility. If this verification succeeds—meaning `col.count()` returns valid results and write operations persist—the migration is skipped to avoid unnecessary data handling. If the round-trip fails (common with legacy 0.6.x collections that silently ignore writes), the process falls back to raw SQLite extraction.

### Extracting Data via Raw SQLite

When the compatibility check fails, `extract_drawers_from_sqlite()` bypasses the ChromaDB API entirely and reads directly from the `chroma.sqlite3` file. This function returns a list of dictionaries containing `{id, document, metadata}` for every drawer in the palace, ensuring data accessibility even when the ChromaDB client cannot open the collection. This low-level extraction is the critical bridge that allows migration across incompatible schema versions.

### Rebuilding and Atomic Directory Swapping

The migration creates a fresh temporary palace using `tempfile.mkdtemp()` and initializes a new collection via `ChromaBackend().get_or_create_collection()`. Extracted drawers are re-imported in batches using `col.add(...)` to prevent memory spikes when handling large datasets. 

Once the new database is populated, the module performs an atomic swap: the original palace directory is renamed to `<palace>.old`, and the temporary directory is moved into place using `os.replace()`. If this operation crosses filesystem boundaries (raising `EXDEV`), the code falls back to `shutil.move()`. Should any step fail, `_restore_stale_palace()` automatically rolls back to the original state, ensuring the user never loses access to their data.

## Safety Mechanisms and Rollback Protection

The migrate module implements multiple safeguards to protect against data corruption during the migration process.

### Automatic Backup Creation

Before any destructive operation, the module creates a timestamped backup of the entire original palace using `shutil.copytree()`. This backup persists after migration completion, allowing users to manually recover their pre-migration state if needed.

### Dry-Run Mode

The `--dry-run` flag allows administrators to preview the migration steps without modifying the filesystem. When enabled, the module reports what actions would be taken—version detection, extraction targets, and swap destinations—enabling validation of the migration plan before committing changes.

### Interactive Confirmation

To prevent accidental data loss, the `confirm_destructive_action()` function requires explicit user confirmation before proceeding with destructive operations. Users can bypass the interactive prompt by passing the `--yes` flag to the CLI, or programmatically via the `confirm=True` parameter in the Python API.

### Wing Name Normalization

After the main database migration completes, `migrate_wing_names()` runs to fix legacy metadata inconsistencies. This function utilizes `normalize_wing_name` from [`mempalace/config.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/config.py) to ensure wing identifiers conform to current naming conventions, resolving potential compatibility issues with newer palace management features.

## Using the Migration Tool

The migrate module can be invoked via command-line interface or programmatically from Python scripts.

Migrate the default palace with interactive confirmation:

```bash
mempalace migrate

```

Migrate a specific palace path automatically accepting defaults:

```bash
mempalace migrate --palace /path/to/palace --yes

```

Preview changes without modifying data:

```bash
mempalace migrate --dry-run

```

Programmatic usage from Python:

```python
from mempalace.migrate import migrate

# Migrate with interactive confirmation

migrate("/home/user/.mempalace")

# Dry-run only

migrate("/home/user/.mempalace", dry_run=True)

# Force migration without prompts

migrate("/home/user/.mempalace", confirm=True)

```

## Summary

- **Version Detection**: The `detect_chromadb_version()` function in [`mempalace/migrate.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/migrate.py) inspects SQLite schema details to identify whether the source palace uses ChromaDB 0.6.x or 1.x formats.
- **Compatibility Verification**: `collection_write_roundtrip_works()` tests read-write-delete cycles to determine if the current backend can access the existing data, skipping unnecessary migrations when possible.
- **Raw Extraction**: `extract_drawers_from_sqlite()` reads legacy databases directly from the SQLite file, bypassing incompatible APIs to retrieve `{id, document, metadata}` records.
- **Atomic Swapping**: The module creates temporary palaces, imports data in batches, and uses `os.replace()` or `shutil.move()` to atomically swap directories, with `_restore_stale_palace()` providing automatic rollback on failure.
- **Safety Features**: Automatic timestamped backups via `shutil.copytree()`, dry-run preview mode, and interactive confirmation through `confirm_destructive_action()` protect against accidental data loss.
- **Metadata Cleanup**: `migrate_wing_names()` normalizes legacy wing identifiers using `normalize_wing_name` from [`mempalace/config.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/config.py) to ensure metadata consistency.

## Frequently Asked Questions

### Does the migration support downgrading ChromaDB versions?

Yes, the migration process is bidirectional. The `extract_drawers_from_sqlite()` function reads data at the SQLite level regardless of the original ChromaDB version, and the rebuild process always uses the currently installed ChromaDB backend. This means you can migrate a palace from 1.x back to 0.6.x compatibility (or vice versa) by running the tool against the target ChromaDB installation.

### What happens if the migration is interrupted?

The atomic directory swap design ensures that interruption at any point before the final `os.replace()` call leaves the original palace untouched. If the swap fails or validation checks trigger a rollback, `_restore_stale_palace()` reverts any partial changes by restoring the original directory from the `<palace>.old` backup. The timestamped backup created by `shutil.copytree()` at the start of the process provides an additional recovery point.

### How do I verify a migration succeeded without modifying data?

Use the `--dry-run` flag when invoking `mempalace migrate`. This mode executes the detection and extraction logic but stops before creating backups or moving directories, reporting exactly which drawers would be migrated and what schema transformations would occur. For programmatic verification, call `migrate(path, dry_run=True)` and inspect the returned migration report object.

### Where are the backup files stored during migration?

Before the atomic swap occurs, `shutil.copytree()` creates a timestamped backup of the entire palace directory at a location adjacent to the original. After successful migration, the original directory is preserved as `<palace>.old` in the same parent directory. Both the timestamped backup and the `.old` directory persist until manually deleted, providing multiple recovery options if issues are discovered post-migration.