# How the MemPalace Normalize Module Converts Chat Exports to Standard Transcripts

> Discover how the MemPalace normalize module converts diverse chat exports like ChatGPT, Slack, and Gemini into clean, standard transcripts. Learn about its multi-stage parsing pipeline.

- Repository: [MemPalace/mempalace](https://github.com/MemPalace/mempalace)
- Tags: internals
- Published: 2026-06-06

---

**The `normalize` module in [`mempalace/normalize.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/normalize.py) executes a multi-stage pipeline that detects input formats, dispatches to specialized parsers for Claude Code, ChatGPT, OpenAI Codex, Gemini CLI, and Slack exports, strips UI noise and system metadata using line-anchored regexes, and outputs a canonical transcript where user turns are prefixed with `>` and assistant turns appear as plain text.**

The [`mempalace/normalize.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/normalize.py) module serves as the universal adapter for the MemPalace system, transforming the disparate export formats generated by modern AI assistants into a single, queryable structure. Understanding this conversion pipeline is essential for developers extending the library or debugging import issues, as it handles everything from 500 MiB safety limits to recursive tree traversal of ChatGPT conversation mappings.

## File Intake and Format Detection

The normalization process begins with the `normalize()` function, which implements strict safety checks before attempting any parsing. According to the source code at lines 18-27, the module first validates that the input file is under 500 MiB to prevent memory exhaustion, then reads the entire content as UTF-8 while preserving Unicode characters and handling optional UTF-8-SIG encodings.

### Early-Out Optimization for Plain Transcripts

Before invoking heavy JSON parsers, the module checks if the file is already in the target MemPalace format. As implemented at lines 33-37, if the file contains at least three lines beginning with the `>` character, `normalize()` returns the content unchanged. This optimization prevents re-processing of previously normalized transcripts.

### JSON Detection and Router Dispatch

When the early-out fails, the module identifies JSON-based exports by checking file extensions (`.json` or `.jsonl`) or examining whether the first non-whitespace character is `{` or `[` (lines 41-45). The content is then passed to `_try_normalize_json()`, which acts as a router to format-specific handlers implemented between lines 50-71.

## Parsing AI-Specific Export Formats

The normalize module implements dedicated parsers for each supported chat platform, extracting `(role, text)` tuples that serve as the intermediate representation before final transcript generation.

### Claude Code JSONL Processing

For Claude Code exports, the `_try_claude_code_jsonl()` handler processes each line as independent JSON records. The parser distinguishes `human` or `user` entries as user turns and `assistant` entries as assistant turns, collecting tool-use blocks into a `tool_use_map` to associate subsequent `tool_result` entries. After extracting text via `_extract_content()`, the parser applies `strip_noise()` per-message to prevent cross-turn contamination. Tool invocations are rendered into human-readable summaries like `[Bash ...]` or `[Read ...]` via internal formatting functions (lines 78-134).

### OpenAI Codex and Gemini CLI Handling

The Codex parser filters for `event_msg` entries while ignoring synthetic `response_item` records, mapping `payload.type` values of `user_message` and `agent_message` to their respective roles. Conversely, the Gemini CLI handler waits for a `session_metadata` sentinel before consuming alternating `user` and `gemini` records, concatenating multiple `content` blocks into single turn strings.

### Claude AI and ChatGPT JSON Structures

Claude AI exports are normalized through `_collect_claude_messages()`, which handles both flat `messages` arrays and privacy-export structures where conversations nest their own `chat_messages`. The function normalizes role names—converting `user` or `human` to the canonical user role, and `assistant` or `ai` to the assistant role—while pulling content from `content` or fallback `text` fields.

For ChatGPT's [`conversations.json`](https://github.com/MemPalace/mempalace/blob/main/conversations.json), the parser traverses the `mapping` tree starting from the root node identified by `parent=None`. It follows the first child of each node recursively, extracting the author role and concatenating the `parts` of message content to produce a linear conversation flow (lines 104-123).

### Slack Export Support

Slack exports require special handling for multi-party conversations. The parser alternates roles between `user` and `assistant` while preserving original speaker identifiers inside brackets (e.g., `[U12345]`). A provenance footer is appended to mark the Slack origin, ensuring traceability within the MemPalace system (lines 146-162).

## Content Cleaning and Transcript Assembly

### The Noise Stripping Layer

Chat exports frequently contain UI chrome, system tags, and hook output that must be removed before storage. The `strip_noise()` function (lines 93-110) employs line-anchored regex patterns defined in `_NOISE_TAG_PATTERNS`, `_NOISE_LINE_PATTERNS`, `_HOOK_LINE_RE`, and `_COLLAPSED_LINES_RE` to surgically remove metadata without affecting message content. This line-anchored approach guarantees that stray tags cannot accidentally consume content from neighboring messages.

### Generating the Standard Transcript

After parsing yields a list of `(role, text)` tuples, `_messages_to_transcript()` (lines 134-162) constructs the final output. User turns are prefixed with `> `, assistant turns are written verbatim, and blank lines separate each turn to ensure readability. The module optionally invokes `mempalace.spellcheck` for text correction before returning the complete transcript string to the caller.

## Practical Usage Examples

The normalize module exposes both a Python API and a command-line interface for standalone operation.

```python

# Normalize a Claude Code JSONL export programmatically

from mempalace.normalize import normalize

transcript = normalize("/path/to/claude_code.jsonl")
print(transcript[:500])  # Preview first 500 characters

```

```bash

# Process a ChatGPT export via the CLI entry point

$ python -m mempalace.normalize ~/Downloads/chatgpt_conversations.json
File: conversations.json
Normalized: 18423 chars | 42 user turns detected

--- Preview (first 20 lines) ---
> How can I list all files in a directory?
Sure! Here's a Python snippet:
import os
...

```

```python

# Integration with the MemPalace ingestion pipeline

from mempalace.normalize import normalize
from mempalace.palace import ingest_transcript

raw = normalize("slack_export.json")
ingest_transcript(raw)  # Stores in appropriate wing/room

```

## Summary

- **The `normalize()` function** enforces a 500 MiB file size limit and UTF-8 encoding before processing begins.
- **Plain transcripts** containing three or more `>`-prefixed lines bypass parsing and return unchanged for performance.
- **Format-specific parsers** handle Claude Code, OpenAI Codex, Gemini CLI, Claude AI, ChatGPT, and Slack exports, each extracting `(role, text)` tuples.
- **Noise stripping** uses line-anchored regexes in `strip_noise()` to remove UI chrome without cross-turn bleed.
- **The final assembly** in `_messages_to_transcript()` prefixes user turns with `> ` and assistant turns with plain text, creating the canonical MemPalace format.

## Frequently Asked Questions

### What chat export formats does the normalize module support?

The module supports plain text files with `>` markers, Claude AI JSON, ChatGPT [`conversations.json`](https://github.com/MemPalace/mempalace/blob/main/conversations.json), Claude Code JSONL, OpenAI Codex JSONL, Gemini CLI JSONL, and Slack JSON exports. Each format has a dedicated `_try_...` parser function that handles its specific schema.

### How does the module handle tool calls in Claude Code conversations?

The Claude Code parser collects tool invocations into a `tool_use_map` and associates subsequent `tool_result` blocks with their originating calls. These are rendered into concise, human-readable one-liners like `[Bash ls -la]` or `[Read file.txt]` using `_format_tool_use()` and `_format_tool_result()` before the text enters the transcript.

### Why does the normalize module check for the `>` character before parsing JSON?

This early-out optimization at lines 33-37 prevents unnecessary computational overhead. If a file already contains at least three lines starting with `>`, the module assumes it is a pre-formatted MemPalace transcript and returns it unchanged, avoiding expensive JSON parsing and noise stripping operations.

### What happens if an input file exceeds the size limit?

The `normalize()` function raises an error if the file size exceeds 500 MiB (as implemented at lines 18-27). This safety guard prevents memory exhaustion when processing extremely large chat exports, ensuring stable operation on resource-constrained systems.