# Using the MemPalace Sweeper for Per-Message Verbatim Recall

> Master per-message verbatim recall with the MemPalace sweeper. Efficiently process Claude-Code JSONL transcripts for complete, resume-safe message storage.

- Repository: [MemPalace/mempalace](https://github.com/MemPalace/mempalace)
- Tags: how-to-guide
- Published: 2026-06-07

---

**The MemPalace sweeper processes Claude-Code JSONL transcripts at message-level granularity, ensuring every user and assistant exchange is stored verbatim with resume-safe, idempotent upserts.**

MemPalace is architected to store every word a user says verbatim. While the primary miners ([`miner.py`](https://github.com/MemPalace/mempalace/blob/main/miner.py), [`convo_miner.py`](https://github.com/MemPalace/mempalace/blob/main/convo_miner.py)) operate on *file-level* chunks, the **Sweeper** provides *message-level* granularity that guarantees no single exchange is ever lost, making it the definitive safety net for complete conversation archiving.

## How the Sweeper Achieves Verbatim Recall

The implementation in [`mempalace/sweeper.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/sweeper.py) treats each message as an individual record, bypassing the size limitations of chunk-based processing.

### Parsing Claude-Code JSONL Transcripts

The sweeper begins by normalizing heterogeneous Claude message payloads. The helper `_flatten_content` (lines 56–84) extracts text and tool calls into a unified string format while preserving verbatim content. The `parse_claude_jsonl` function then iterates over transcripts, yielding dictionaries containing:

- `session_id` – the transcript’s session identifier
- `uuid` – a per-message UUID
- `timestamp` – ISO-8601 string for lexical sorting
- `role` – `user` or `assistant`
- `content` – the flattened text

This parsing occurs between lines 88–101 in [`mempalace/sweeper.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/sweeper.py), ensuring structured access to every individual message.

### Resume-Safe Cursor Management

For each `session_id`, the sweeper queries the storage backend via `get_palace_cursor` (lines 47–62) to retrieve the **maximum timestamp** of already-ingested drawers. This cursor-based approach allows the process to resume safely after crashes, as any message with a timestamp lexically less than the cursor is known to be present. Because timestamps are compared as ISO-8601 strings, no date parsing is required, keeping the operation fast and reliable.

### Idempotent Drawer Generation

To prevent duplicate storage, the sweeper generates deterministic drawer IDs using `_drawer_id_for_message` (lines 83–90). This function constructs IDs from the combination of session ID and message UUID, ensuring that running the sweeper twice on the same transcript produces a "no-op" upsert rather than redundant data.

### Batch Upserts with Deduplication

Messages newer than the cursor are collected in batches (default size 64) and upserted via `collection.upsert`. Before each batch, the code checks which IDs already exist in the backend (lines 123–138) to maintain accurate metrics for `drawers_added` and `drawers_already_present`. This batching strategy optimizes throughput while preserving the guarantee of per-message storage—each drawer typically holds 1–5 KB representing a single exchange.

### Directory-Wide Processing

The `sweep_directory` function (lines 302–324) recursively walks a directory, calling `sweep` on every `*.jsonl` file and aggregating results into a summary report. Errors are logged but still counted in `files_attempted` to ensure failures remain visible in the output metrics.

## CLI Integration

The sweeper is exposed to end-users through [`mempalace/cli.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/cli.py), where the `sweep` and `sweep_directory` commands are registered around line 617. This interface accepts file or directory paths and palace locations, printing the same structured results available via the Python API.

## Code Examples

### Sweep a Single Transcript

```python
from mempalace.sweeper import sweep

result = sweep(
    jsonl_path="/tmp/session.jsonl",
    palace_path="/data/mempalace",
    source_label="session.jsonl"
)

print(result)

# → {

#   "drawers_added": 42,

#   "drawers_already_present": 0,

#   "drawers_upserted": 42,

#   "drawers_skipped": 0,

#   "cursor_by_session": {"abc123": "2024-05-01T12:34:56Z"}

# }

```

### Sweep an Entire Directory

```python
from mempalace.sweeper import sweep_directory

summary = sweep_directory(
    dir_path="/tmp/transcripts",
    palace_path="/data/mempalace"
)

print(summary)

# → {

#   "files_attempted": 12,

#   "files_succeeded": 12,

#   "drawers_added": 317,

#   "drawers_already_present": 5,

#   "drawers_skipped": 3,

#   "per_file": [...],

#   "failures": []

# }

```

### Command-Line Usage

```bash
$ mempalace sweep ./my_session.jsonl ./palace

# prints the same dict as the Python API

```

## Summary

- **Per-message granularity**: Unlike file-level miners, the sweeper stores each exchange as an individual drawer in [`mempalace/sweeper.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/sweeper.py).
- **Resume-safe operation**: The cursor mechanism using `get_palace_cursor` allows interrupted sweeps to continue without reprocessing existing data.
- **Idempotent storage**: Deterministic drawer IDs generated by `_drawer_id_for_message` prevent duplicates across multiple runs.
- **No size caps**: Each drawer stores approximately 1–5 KB, ensuring verbatim recall without chunking limitations.
- **Lexical timestamp sorting**: ISO-8601 string comparison eliminates date parsing overhead while maintaining chronological accuracy.

## Frequently Asked Questions

### What is the difference between MemPalace miners and the sweeper?

The primary miners ([`miner.py`](https://github.com/MemPalace/mempalace/blob/main/miner.py), [`convo_miner.py`](https://github.com/MemPalace/mempalace/blob/main/convo_miner.py)) process conversations at the file level, creating chunks that may span multiple messages. The **sweeper**, implemented in [`mempalace/sweeper.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/sweeper.py), operates at message granularity, ensuring that every individual user and assistant exchange is stored verbatim as a separate record.

### How does the sweeper handle crashes or interruptions?

The sweeper uses a cursor-based resumption system. By calling `get_palace_cursor` to retrieve the maximum timestamp for each `session_id`, the sweeper skips any messages with timestamps earlier than the cursor. This design makes the process inherently resume-safe without requiring transaction logs.

### What makes the storage idempotent?

Each message receives a deterministic drawer ID constructed from the session ID and message UUID via `_drawer_id_for_message`. Because these IDs are consistent across runs, upserting the same transcript twice results in the second run reporting those messages as `drawers_already_present` rather than creating duplicates.

### What file formats does the MemPalace sweeper support?

The sweeper specifically processes Claude-Code **JSONL** transcripts. The `parse_claude_jsonl` function handles the specific schema produced by Claude-Code sessions, normalizing heterogeneous message payloads (including tool calls) into flat text content suitable for verbatim storage.