How to Repair MemPalace Palace Inconsistencies with the repair Module
The mempalace.repair module provides a safety-first toolkit for detecting and fixing HNSW index corruption through four coordinated operations—status checks, scanning, pruning, and rebuilding—that work entirely on-disk before engaging the ChromaDB client.
The mempalace.repair module serves as the single entry point for resolving corruption in a MemPalace, implementing a defensive architecture that treats the SQLite database as ground truth while protecting against segfaults from damaged HNSW indices. When you need to fix repairing MemPalace palace inconsistencies, this module coordinates read-only probes, batch validation, and non-destructive rebuilds to recover data without service interruption.
Understanding the Repair Architecture
MemPalace maintains two distinct data layers that can diverge during hardware failures or crashes. The repair module addresses this split by validating the relationship between the SQLite canonical store and the HNSW vector index before performing any write operations.
Ground Truth and Index Divergence
The SQLite file (chroma.sqlite3) contains the authoritative drawer data, while ChromaDB stores a separate HNSW index (data_level0.bin, link_lists.bin, index_metadata.pickle) for fast vector search. When these layers become inconsistent, the HNSW index can cause segfaults before Python code executes. The repair module treats the SQLite count as the source of truth, comparing it against the HNSW metadata to detect divergence.
Safety-First Design
Before opening any ChromaDB client connection, the module probes on-disk metadata to verify index integrity. This approach protects the MCP server and tooling from crashes when opening corrupted palaces, ensuring that repairing MemPalace palace inconsistencies never risks additional system instability.
Core Repair Operations
The mempalace/repair.py file implements four distinct operations that operators can execute sequentially or independently. Each function targets a specific phase of corruption detection and remediation.
Checking Status with repair.status()
The repair.status() function (source lines 135-164) performs a read-only health check by comparing the SQLite drawer row count against the HNSW index capacity. It calls hnsw_capacity_status in mempalace/backends/chroma.py (source lines 645-674) to probe the on-disk HNSW element count without loading the native segment. If the counts diverge beyond a safe margin, the function reports a diverged status, signaling the need for remediation.
Scanning for Corrupt IDs with repair.scan_palace()
The repair.scan_palace() function (source lines 46-78) walks every drawer ID in the collection, batching probes in groups of 100. Unfetchable IDs—those that exist in the SQLite metadata but cannot be retrieved from the index—are written to corrupt_ids.txt for later processing. This operation identifies specific records requiring removal without modifying the live dataset.
Pruning Corrupt Data with repair.prune_corrupt()
The repair.prune_corrupt() function (source lines 120-157) reads the corrupt_ids.txt file generated by the scan phase and deletes those specific IDs from the collection. By default, this function performs a dry-run, displaying the count of affected records without executing deletions. Supplying confirm=True (or --confirm via CLI) executes the actual removal, eliminating corrupted pointers while preserving valid data.
Rebuilding the Index with repair.rebuild_index()
The repair.rebuild_index() function (source lines 184-222) performs the most aggressive repair by reconstructing the entire HNSW index from the SQLite ground truth. The process extracts all drawers using _extract_drawers, validates the extraction count against SQLite via check_extraction_safety, backs up the existing database, creates a temporary collection (<collection>__repair_tmp), streams valid drawers back in batches, atomically swaps the live collection, and finally vacuums the SQLite file and rebuilds the FTS5 index.
Implementation Details and Safety Guards
Two critical validation mechanisms ensure that repairing MemPalace palace inconsistencies never results in data loss or system crashes.
HNSW Capacity Probing
The hnsw_capacity_status function in mempalace/backends/chroma.py (source lines 645-674) reads HNSW metadata directly from disk without initializing the ChromaDB client. This probe detects index truncation or corruption that would cause immediate segfaults upon client initialization, allowing the repair module to exit gracefully with a diagnostic message rather than crashing the Python process.
Extraction Safety Validation
The check_extraction_safety function (source lines 226-259) prevents a documented failure mode where the SQLite count exceeds the extracted drawer count (issue #1208). Before any destructive operation, this guard validates that the number of drawers extracted by the repair code matches the SQLite count, aborting the process if discrepancies are detected. Additionally, the _rebuild_collection_via_temp function (source lines 108-138) implements exception handling that restores the original SQLite backup if any rebuild step fails, ensuring the palace remains untouched on error.
CLI and API Usage
The repair workflow is exposed through the CLI entry points defined in cli.py, providing operators with predictable commands for each phase.
Command-Line Workflow
# 1. Quick health check – safe even on corrupted palaces
mempalace repair-status
# 2. Find unfetchable IDs; outputs to corrupt_ids.txt
mempalace repair-scan --wing=John
# 3. Remove bad IDs (dry-run first, then confirm)
mempalace repair-prune
mempalace repair-prune --confirm
# 4. Full index reconstruction from SQLite ground truth
mempalace repair-rebuild
Programmatic API
from mempalace import repair
# Validate palace integrity without opening client
status = repair.status()
print(status["message"])
# Identify corruption
good, bad = repair.scan_palace()
print(f"Found {len(bad)} corrupt IDs")
# Execute repairs
repair.prune_corrupt(confirm=True)
repair.rebuild_index()
Summary
- The
mempalace.repairmodule provides the sole interface for repairing MemPalace palace inconsistencies, coordinating four distinct operations:status(),scan_palace(),prune_corrupt(), andrebuild_index(). - All operations begin with on-disk validation (
hnsw_capacity_statusinmempalace/backends/chroma.py) to prevent segfaults from corrupted HNSW indices before opening ChromaDB clients. - SQLite serves as the immutable ground truth; the
check_extraction_safetyguard (source lines 226-259) ensures extraction counts match database records before any destructive operation. - The rebuild process creates temporary collections for staging, performs atomic swaps, and restores from backup on failure, guaranteeing data safety during index reconstruction.
- CLI commands (
repair-status,repair-scan,repair-prune,repair-rebuild) and Python API offer equivalent functionality for both automation and interactive debugging.
Frequently Asked Questions
What causes palace inconsistencies in MemPalace?
Palace inconsistencies typically occur when the HNSW vector index (data_level0.bin, link_lists.bin) becomes desynchronized from the SQLite metadata (chroma.sqlite3) due to unclean shutdowns, disk full conditions, or process crashes during write operations. Because ChromaDB maintains separate storage for the vector index and relational metadata, hardware failures can corrupt the binary index while leaving the SQLite records intact, resulting in unfetchable drawer IDs or count mismatches.
How does the repair module prevent crashes when checking corrupted palaces?
The module implements a defensive "read-only first" strategy through the hnsw_capacity_status probe in mempalace/backends/chroma.py (source lines 645-674). This function parses the HNSW metadata files directly on disk without loading the native segment into memory or instantiating a ChromaDB client. By detecting truncation or corruption at the filesystem level, the module can report status and exit cleanly rather than triggering the segfaults that occur when a corrupted HNSW index loads into the ChromaDB C++ backend.
What is the difference between pruning and rebuilding a palace?
Pruning (repair.prune_corrupt()) is a surgical operation that removes only the specific drawer IDs identified as corrupt during the scan phase, preserving all valid data and index structure. Rebuilding (repair.rebuild_index()) is a comprehensive operation that extracts all valid drawers from SQLite, drops the existing HNSW index entirely, and constructs a new vector search structure from scratch. Use pruning for minor corruption affecting specific records; use rebuilding when the HNSW index itself is damaged or when count divergences indicate systemic index failure.
Is the repair process safe for production data?
Yes, the repair module implements multiple safety mechanisms to protect production data. The check_extraction_safety validation ensures data completeness before any modifications occur. The rebuild process creates a full SQLite backup before starting, writes to a temporary collection (<collection>__repair_tmp) rather than the live index, and only performs the atomic swap after verification. If any step raises an exception, the module restores the original database from backup, leaving the production palace in its pre-repair state.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →