Using the MemPalace Sweeper for Per-Message Verbatim Recall
The MemPalace sweeper processes Claude-Code JSONL transcripts at message-level granularity, ensuring every user and assistant exchange is stored verbatim with resume-safe, idempotent upserts.
MemPalace is architected to store every word a user says verbatim. While the primary miners (miner.py, convo_miner.py) operate on file-level chunks, the Sweeper provides message-level granularity that guarantees no single exchange is ever lost, making it the definitive safety net for complete conversation archiving.
How the Sweeper Achieves Verbatim Recall
The implementation in mempalace/sweeper.py treats each message as an individual record, bypassing the size limitations of chunk-based processing.
Parsing Claude-Code JSONL Transcripts
The sweeper begins by normalizing heterogeneous Claude message payloads. The helper _flatten_content (lines 56–84) extracts text and tool calls into a unified string format while preserving verbatim content. The parse_claude_jsonl function then iterates over transcripts, yielding dictionaries containing:
session_id– the transcript’s session identifieruuid– a per-message UUIDtimestamp– ISO-8601 string for lexical sortingrole–userorassistantcontent– the flattened text
This parsing occurs between lines 88–101 in mempalace/sweeper.py, ensuring structured access to every individual message.
Resume-Safe Cursor Management
For each session_id, the sweeper queries the storage backend via get_palace_cursor (lines 47–62) to retrieve the maximum timestamp of already-ingested drawers. This cursor-based approach allows the process to resume safely after crashes, as any message with a timestamp lexically less than the cursor is known to be present. Because timestamps are compared as ISO-8601 strings, no date parsing is required, keeping the operation fast and reliable.
Idempotent Drawer Generation
To prevent duplicate storage, the sweeper generates deterministic drawer IDs using _drawer_id_for_message (lines 83–90). This function constructs IDs from the combination of session ID and message UUID, ensuring that running the sweeper twice on the same transcript produces a "no-op" upsert rather than redundant data.
Batch Upserts with Deduplication
Messages newer than the cursor are collected in batches (default size 64) and upserted via collection.upsert. Before each batch, the code checks which IDs already exist in the backend (lines 123–138) to maintain accurate metrics for drawers_added and drawers_already_present. This batching strategy optimizes throughput while preserving the guarantee of per-message storage—each drawer typically holds 1–5 KB representing a single exchange.
Directory-Wide Processing
The sweep_directory function (lines 302–324) recursively walks a directory, calling sweep on every *.jsonl file and aggregating results into a summary report. Errors are logged but still counted in files_attempted to ensure failures remain visible in the output metrics.
CLI Integration
The sweeper is exposed to end-users through mempalace/cli.py, where the sweep and sweep_directory commands are registered around line 617. This interface accepts file or directory paths and palace locations, printing the same structured results available via the Python API.
Code Examples
Sweep a Single Transcript
from mempalace.sweeper import sweep
result = sweep(
jsonl_path="/tmp/session.jsonl",
palace_path="/data/mempalace",
source_label="session.jsonl"
)
print(result)
# → {
# "drawers_added": 42,
# "drawers_already_present": 0,
# "drawers_upserted": 42,
# "drawers_skipped": 0,
# "cursor_by_session": {"abc123": "2024-05-01T12:34:56Z"}
# }
Sweep an Entire Directory
from mempalace.sweeper import sweep_directory
summary = sweep_directory(
dir_path="/tmp/transcripts",
palace_path="/data/mempalace"
)
print(summary)
# → {
# "files_attempted": 12,
# "files_succeeded": 12,
# "drawers_added": 317,
# "drawers_already_present": 5,
# "drawers_skipped": 3,
# "per_file": [...],
# "failures": []
# }
Command-Line Usage
$ mempalace sweep ./my_session.jsonl ./palace
# prints the same dict as the Python API
Summary
- Per-message granularity: Unlike file-level miners, the sweeper stores each exchange as an individual drawer in
mempalace/sweeper.py. - Resume-safe operation: The cursor mechanism using
get_palace_cursorallows interrupted sweeps to continue without reprocessing existing data. - Idempotent storage: Deterministic drawer IDs generated by
_drawer_id_for_messageprevent duplicates across multiple runs. - No size caps: Each drawer stores approximately 1–5 KB, ensuring verbatim recall without chunking limitations.
- Lexical timestamp sorting: ISO-8601 string comparison eliminates date parsing overhead while maintaining chronological accuracy.
Frequently Asked Questions
What is the difference between MemPalace miners and the sweeper?
The primary miners (miner.py, convo_miner.py) process conversations at the file level, creating chunks that may span multiple messages. The sweeper, implemented in mempalace/sweeper.py, operates at message granularity, ensuring that every individual user and assistant exchange is stored verbatim as a separate record.
How does the sweeper handle crashes or interruptions?
The sweeper uses a cursor-based resumption system. By calling get_palace_cursor to retrieve the maximum timestamp for each session_id, the sweeper skips any messages with timestamps earlier than the cursor. This design makes the process inherently resume-safe without requiring transaction logs.
What makes the storage idempotent?
Each message receives a deterministic drawer ID constructed from the session ID and message UUID via _drawer_id_for_message. Because these IDs are consistent across runs, upserting the same transcript twice results in the second run reporting those messages as drawers_already_present rather than creating duplicates.
What file formats does the MemPalace sweeper support?
The sweeper specifically processes Claude-Code JSONL transcripts. The parse_claude_jsonl function handles the specific schema produced by Claude-Code sessions, normalizing heterogeneous message payloads (including tool calls) into flat text content suitable for verbatim storage.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →