How the MemPalace Repair Module Performs Palace Consistency Checks
The MemPalace repair module performs palace consistency checks through a three-stage pipeline—read-only status verification, corrupt ID scanning, and safety-guarded index rebuilding—that compares SQLite drawer counts against HNSW index capacity and validates data retrievability before executing atomic repairs.
The mempalace.repair subsystem serves as the safety net for the MemPalace vector memory system, guarding against the index corruption that occurs when the ChromaDB HNSW segment desynchronizes from the SQLite backing store. Understanding how this module executes palace consistency checks is essential for maintaining data integrity in production deployments where HNSW capacity drift or un-fetchable IDs might otherwise cause crashes or silent data loss.
Understanding Palace Corruption Failure Modes
The repair module specifically targets two primary failure modes that can corrupt the on-disk index:
| Failure Mode | Symptom | Consequence |
|---|---|---|
| HNSW capacity drift | The ChromaDB HNSW segment freezes at a stale max_elements while SQLite continues accepting new drawers |
The palace crashes when tools attempt to load the undersized HNSW index |
| Corrupt/un-fetchable IDs | Duplicate add() calls or a broken WAL leave entries that the collection layer cannot retrieve |
Collection.get() raises errors and link-list files can grow to terabytes |
The Three-Stage Consistency Verification Pipeline
The module implements three complementary checks that together provide a complete health picture without risking data integrity.
Stage 1: Read-Only Status Verification
The status() function (mempalace/repair.py lines [35‑71]) performs a fast, non-destructive health check that never opens a ChromaDB client. Instead, it directly reads chroma.sqlite3 and the HNSW metadata pickle via mempalace.backends.chroma.hnsw_capacity_status.
This check compares the SQLite drawer row count (sqlite_drawer_count) against the HNSW element count. When these values diverge, the function returns a DIVERGED status string, signaling that the palace requires a full repair. This read-only approach allows frequent monitoring without locking the database or loading the full vector index into memory.
Stage 2: Deep Corruption Scanning
The scan_palace() function (mempalace/repair.py lines [47‑55]) executes a comprehensive audit of drawer retrievability. It lists every drawer ID using pagination via _paginate_ids, then probes IDs in batches of 100. Any ID that cannot be fetched is added to a bad_set.
The function writes these corrupt IDs to corrupt_ids.txt inside the palace directory, creating a precise manifest for later pruning operations. This scan identifies entries that exist in the SQLite embeddings table but have become orphaned from the HNSW index or vector storage.
Stage 3: Safety-Guarded Index Rebuild
The rebuild_index() function (mempalace/repair.py lines [117‑129]) executes the final reconstruction phase only after passing multiple safety validations. It first extracts all drawers using _extract_drawers, then invokes check_extraction_safety() (mempalace/repair.py lines [98‑119]) to cross-check the extracted count against the ground-truth SQLite count.
Two guard conditions abort the rebuild immediately if triggered:
- Strong guard: SQLite reports more rows than were extracted, indicating data truncation
- Weak guard: Extraction stopped exactly at ChromaDB’s default
10,000limit and the SQLite check could not run
If either guard fires, the system raises a TruncationDetected exception with a detailed abort message, preventing silent data loss from partial extractions.
Critical Safety Mechanisms
Beyond the three-stage pipeline, the repair module implements additional integrity safeguards:
sqlite_drawer_count()(mempalace/repair.pylines [60‑71]): Reads the SQLiteembeddingstable directly via SQL query, bypassing the ChromaDB client entirely to avoid client-side caching or connection issuessqlite_integrity_errors()(mempalace/repair.pylines [101‑115]): ExecutesPRAGMA quick_checkon the SQLite file before any destructive operation, aborting early if the database file itself shows corruption_rebuild_collection_via_temp()(mempalace/repair.pylines [89‑124]): Constructs a temporary collection first, verifies its size matches expectations, then performs an atomic swap to replace the live collection, ensuring partially-written indexes never activate
Running Palace Consistency Checks
Command-Line Interface
Execute consistency checks using the MemPalace CLI:
# Fast health check (read-only)
$ mempalace repair-status
# Outputs SQLite vs HNSW counts and divergence status
# Scan for corrupt IDs (creates corrupt_ids.txt)
$ mempalace repair-scan --wing <wing-name>
# Bad IDs written to <palace>/corrupt_ids.txt
# Prune corrupt IDs (dry-run first recommended)
$ mempalace repair-prune # Shows dry-run summary
$ mempalace repair-prune --confirm # Executes deletion
# Full automated repair pipeline
$ mempalace repair
# Runs scan → prune → rebuild with automatic backup of chroma.sqlite3
Programmatic Usage
Integrate consistency checks directly into Python applications:
from mempalace.repair import status, scan_palace, prune_corrupt, rebuild_index
# Verify palace health without side effects
report = status()
print("Palace health:", report["drawers"]["status"])
# Detect unreachable IDs
good_ids, bad_ids = scan_palace()
print(f"Found {len(bad_ids)} corrupt IDs")
# Remove corrupted entries
prune_corrupt(confirm=True)
# Safely rebuild HNSW index from SQLite source
rebuild_index()
Summary
- Read-only verification: The
status()function inmempalace/repair.pycompares SQLite drawer counts against HNSW capacity by directly querying the database files without loading the ChromaDB client - Corruption detection: The
scan_palace()function probes drawer IDs in batches of 100, documenting un-fetchable entries incorrupt_ids.txtfor surgical removal - Safe reconstruction: The
rebuild_index()function enforces extraction safety guards and atomic collection swaps via_rebuild_collection_via_temp(), ensuring that SQLite integrity errors or truncation events abort before data loss occurs
Frequently Asked Questions
What triggers a DIVERGED status during palace consistency checks?
A DIVERGED status occurs when status() detects that the SQLite embeddings table contains more rows than the HNSW index has capacity for, or when the element counts otherwise mismatch. This typically happens when the HNSW segment stops accepting new elements due to a frozen max_elements value while SQLite continues receiving new drawers, creating a capacity drift that will eventually cause index load failures.
How does the repair module prevent data loss during index rebuilds?
The module implements extraction safety guards in check_extraction_safety() that compare the number of successfully extracted drawers against the ground-truth count from sqlite_drawer_count(). If SQLite reports more rows than were extracted (strong guard) or if extraction hit the 10,000-item default limit without verification (weak guard), the system raises TruncationDetected and aborts before overwriting the existing index.
What is the difference between scan_palace and status checks?
status() performs a fast metadata comparison by reading file headers and SQLite row counts without fetching actual vectors, making it suitable for frequent monitoring. scan_palace() performs a deep integrity test by attempting to fetch every drawer ID through the collection layer, identifying specific IDs that have become corrupted or un-fetchable despite existing in the database—operations that require significantly more time but pinpoint exact data loss locations.
Why does the repair module read SQLite directly instead of using the ChromaDB client?
Direct SQLite access via sqlite_drawer_count() bypasses the ChromaDB client to avoid connection overhead, client-side caching issues, and potential HNSW locking problems. This approach allows the repair module to obtain ground-truth row counts even when the HNSW index is corrupted or the ChromaDB client cannot initialize, ensuring that consistency checks remain possible during failure scenarios.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →