How the MemPalace Normalize Module Converts Chat Exports to Standard Transcripts
The normalize module in mempalace/normalize.py executes a multi-stage pipeline that detects input formats, dispatches to specialized parsers for Claude Code, ChatGPT, OpenAI Codex, Gemini CLI, and Slack exports, strips UI noise and system metadata using line-anchored regexes, and outputs a canonical transcript where user turns are prefixed with > and assistant turns appear as plain text.
The mempalace/normalize.py module serves as the universal adapter for the MemPalace system, transforming the disparate export formats generated by modern AI assistants into a single, queryable structure. Understanding this conversion pipeline is essential for developers extending the library or debugging import issues, as it handles everything from 500 MiB safety limits to recursive tree traversal of ChatGPT conversation mappings.
File Intake and Format Detection
The normalization process begins with the normalize() function, which implements strict safety checks before attempting any parsing. According to the source code at lines 18-27, the module first validates that the input file is under 500 MiB to prevent memory exhaustion, then reads the entire content as UTF-8 while preserving Unicode characters and handling optional UTF-8-SIG encodings.
Early-Out Optimization for Plain Transcripts
Before invoking heavy JSON parsers, the module checks if the file is already in the target MemPalace format. As implemented at lines 33-37, if the file contains at least three lines beginning with the > character, normalize() returns the content unchanged. This optimization prevents re-processing of previously normalized transcripts.
JSON Detection and Router Dispatch
When the early-out fails, the module identifies JSON-based exports by checking file extensions (.json or .jsonl) or examining whether the first non-whitespace character is { or [ (lines 41-45). The content is then passed to _try_normalize_json(), which acts as a router to format-specific handlers implemented between lines 50-71.
Parsing AI-Specific Export Formats
The normalize module implements dedicated parsers for each supported chat platform, extracting (role, text) tuples that serve as the intermediate representation before final transcript generation.
Claude Code JSONL Processing
For Claude Code exports, the _try_claude_code_jsonl() handler processes each line as independent JSON records. The parser distinguishes human or user entries as user turns and assistant entries as assistant turns, collecting tool-use blocks into a tool_use_map to associate subsequent tool_result entries. After extracting text via _extract_content(), the parser applies strip_noise() per-message to prevent cross-turn contamination. Tool invocations are rendered into human-readable summaries like [Bash ...] or [Read ...] via internal formatting functions (lines 78-134).
OpenAI Codex and Gemini CLI Handling
The Codex parser filters for event_msg entries while ignoring synthetic response_item records, mapping payload.type values of user_message and agent_message to their respective roles. Conversely, the Gemini CLI handler waits for a session_metadata sentinel before consuming alternating user and gemini records, concatenating multiple content blocks into single turn strings.
Claude AI and ChatGPT JSON Structures
Claude AI exports are normalized through _collect_claude_messages(), which handles both flat messages arrays and privacy-export structures where conversations nest their own chat_messages. The function normalizes role names—converting user or human to the canonical user role, and assistant or ai to the assistant role—while pulling content from content or fallback text fields.
For ChatGPT's conversations.json, the parser traverses the mapping tree starting from the root node identified by parent=None. It follows the first child of each node recursively, extracting the author role and concatenating the parts of message content to produce a linear conversation flow (lines 104-123).
Slack Export Support
Slack exports require special handling for multi-party conversations. The parser alternates roles between user and assistant while preserving original speaker identifiers inside brackets (e.g., [U12345]). A provenance footer is appended to mark the Slack origin, ensuring traceability within the MemPalace system (lines 146-162).
Content Cleaning and Transcript Assembly
The Noise Stripping Layer
Chat exports frequently contain UI chrome, system tags, and hook output that must be removed before storage. The strip_noise() function (lines 93-110) employs line-anchored regex patterns defined in _NOISE_TAG_PATTERNS, _NOISE_LINE_PATTERNS, _HOOK_LINE_RE, and _COLLAPSED_LINES_RE to surgically remove metadata without affecting message content. This line-anchored approach guarantees that stray tags cannot accidentally consume content from neighboring messages.
Generating the Standard Transcript
After parsing yields a list of (role, text) tuples, _messages_to_transcript() (lines 134-162) constructs the final output. User turns are prefixed with > , assistant turns are written verbatim, and blank lines separate each turn to ensure readability. The module optionally invokes mempalace.spellcheck for text correction before returning the complete transcript string to the caller.
Practical Usage Examples
The normalize module exposes both a Python API and a command-line interface for standalone operation.
# Normalize a Claude Code JSONL export programmatically
from mempalace.normalize import normalize
transcript = normalize("/path/to/claude_code.jsonl")
print(transcript[:500]) # Preview first 500 characters
# Process a ChatGPT export via the CLI entry point
$ python -m mempalace.normalize ~/Downloads/chatgpt_conversations.json
File: conversations.json
Normalized: 18423 chars | 42 user turns detected
--- Preview (first 20 lines) ---
> How can I list all files in a directory?
Sure! Here's a Python snippet:
import os
...
# Integration with the MemPalace ingestion pipeline
from mempalace.normalize import normalize
from mempalace.palace import ingest_transcript
raw = normalize("slack_export.json")
ingest_transcript(raw) # Stores in appropriate wing/room
Summary
- The
normalize()function enforces a 500 MiB file size limit and UTF-8 encoding before processing begins. - Plain transcripts containing three or more
>-prefixed lines bypass parsing and return unchanged for performance. - Format-specific parsers handle Claude Code, OpenAI Codex, Gemini CLI, Claude AI, ChatGPT, and Slack exports, each extracting
(role, text)tuples. - Noise stripping uses line-anchored regexes in
strip_noise()to remove UI chrome without cross-turn bleed. - The final assembly in
_messages_to_transcript()prefixes user turns with>and assistant turns with plain text, creating the canonical MemPalace format.
Frequently Asked Questions
What chat export formats does the normalize module support?
The module supports plain text files with > markers, Claude AI JSON, ChatGPT conversations.json, Claude Code JSONL, OpenAI Codex JSONL, Gemini CLI JSONL, and Slack JSON exports. Each format has a dedicated _try_... parser function that handles its specific schema.
How does the module handle tool calls in Claude Code conversations?
The Claude Code parser collects tool invocations into a tool_use_map and associates subsequent tool_result blocks with their originating calls. These are rendered into concise, human-readable one-liners like [Bash ls -la] or [Read file.txt] using _format_tool_use() and _format_tool_result() before the text enters the transcript.
Why does the normalize module check for the > character before parsing JSON?
This early-out optimization at lines 33-37 prevents unnecessary computational overhead. If a file already contains at least three lines starting with >, the module assumes it is a pre-formatted MemPalace transcript and returns it unchanged, avoiding expensive JSON parsing and noise stripping operations.
What happens if an input file exceeds the size limit?
The normalize() function raises an error if the file size exceeds 500 MiB (as implemented at lines 18-27). This safety guard prevents memory exhaustion when processing extremely large chat exports, ensuring stable operation on resource-constrained systems.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →