# Claude‑Code Transcript Processing Flow: Extracting Messages, Thinking Blocks, and Tool Usage

> Discover the Claude-Code transcript processing flow. Extract user messages, thinking blocks, and tool usage from JSONL files. Format data into XML payloads for agent consumption with letta-ai/claude-subconscious.

- Repository: [Letta/claude-subconscious](https://github.com/letta-ai/claude-subconscious)
- Tags: how-to-guide
- Published: 2026-03-26

---

**The transcript processing flow in letta-ai/claude-subconscious reads JSONL transcript files line‑by‑line, extracts user messages, assistant thinking blocks, and tool interactions through specialized parser functions, then formats them into Letta-compatible XML payloads for agent consumption.**

When a Claude‑Code session ends, the *send‑messages‑to‑Letta* hook processes the conversation history to build a compact, context‑rich payload. This pipeline handles plain text, internal reasoning, and tool execution traces according to the source code in [`scripts/transcript_utils.ts`](https://github.com/letta-ai/claude-subconscious/blob/main/scripts/transcript_utils.ts) and [`scripts/send_messages_to_letta.ts`](https://github.com/letta-ai/claude-subconscious/blob/main/scripts/send_messages_to_letta.ts).

## How the Transcript Processing Pipeline Works

### Step 1: Loading the JSONL Transcript

The `readTranscript` function in [`scripts/transcript_utils.ts`](https://github.com/letta-ai/claude-subconscious/blob/main/scripts/transcript_utils.ts) (lines 63‑87) streams the transcript file line‑by‑line. It parses each line with `JSON.parse` and returns an array of `TranscriptMessage` objects ready for content extraction.

```typescript
import { readTranscript } from './transcript_utils.js';

const messages = await readTranscript('path/to/transcript.jsonl');
// Returns: TranscriptMessage[]

```

### Step 2: Extracting Content Blocks with extractAllContent

The `extractAllContent` function (lines 92‑141) inspects the `content` field of each transcript entry. This field can be a plain string or an array of typed blocks. The function discriminates four block types:

- **`type: "text"`** – Plain user or assistant text.
- **`type: "thinking"`** – Internal reasoning generated by Claude.
- **`type: "tool_use"`** – Tool invocations with name and input parameters.
- **`type: "tool_result"`** – Outputs or errors returned from tool execution.

Each fragment is stored in an `ExtractedContent` object containing `textParts`, `thinkingParts`, `toolUses`, and `toolResults` arrays.

```typescript
import { extractAllContent } from './transcript_utils.js';

const extracted = extractAllContent(singleMessage);
console.log(extracted.text);        // Plain text content
console.log(extracted.thinking);    // Array of thinking blocks
console.log(extracted.toolUses);    // [{name: 'Read', input: {file_path: 'src/app.ts'}}]
console.log(extracted.toolResults); // [{tool_use_id: '...', content: '...', is_error: false}]

```

### Step 3: Mapping Tool Usage IDs

While iterating through the transcript, the pipeline builds a `toolNameMap` (lines 155‑162) that associates each `tool_use_id` with its corresponding tool name. This mapping allows later `tool_result` blocks to be rendered with the original tool name even when the result appears in a separate message.

### Step 4: Formatting Letta‑Ready Messages

The `formatMessagesForLetta` function (lines 155‑282) processes entries starting from the last synchronized index. It transforms extracted content into role‑based messages:

- **User messages** – Emitted as `{role: "user", text: ...}`.
- **Assistant thinking** – Wrapped as `[Thinking]: …`, truncated to 500 characters, and emitted as `{role: "assistant", text: ...}`.
- **Assistant tool calls** – Rendered as `[Tool: <name>] <summary>` with concise input summaries (e.g., file paths for `Read`, commands for `Bash`), emitted as `{role: "assistant", text: ...}`.
- **Tool results** – Rendered as `[Tool Result]: <toolName>]\n<content>` or `[Tool Error]` if `is_error` is true, emitted as `{role: "system", text: ...}` with truncated payloads.

```typescript
import { formatMessagesForLetta } from './transcript_utils.js';
import { loadSyncState } from './conversation_utils.js';

const state = loadSyncState(process.cwd(), sessionId);
const lettaMessages = formatMessagesForLetta(messages, state.lastProcessedIndex);

/*
[
  { role: 'user', text: 'Can you refactor this?' },
  { role: 'assistant', text: '[Thinking]: I will split the file...' },
  { role: 'assistant', text: '[Tool: Read] src/utils.ts' },
  { role: 'system', text: '[Tool Result]: Read]\nconst foo = …' }
]
*/

```

### Step 5: Converting to Letta XML

The hook in [`scripts/send_messages_to_letta.ts`](https://github.com/letta-ai/claude-subconscious/blob/main/scripts/send_messages_to_letta.ts) (lines 84‑90) converts the formatted message array into XML. Each entry becomes `<message role="…">…</message>` with XML‑escaped text before being handed to the Letta SDK worker.

```typescript
const xmlPayload = lettaMessages
  .map(m => `<message role="${m.role}">${escapeXml(m.text)}</message>`)
  .join('');

```

## Key Implementation Files

| File | Purpose |
|------|---------|
| [`scripts/transcript_utils.ts`](https://github.com/letta-ai/claude-subconscious/blob/main/scripts/transcript_utils.ts) | Core utilities for reading JSONL transcripts, extracting content blocks, and formatting Letta‑ready messages. |
| [`scripts/send_messages_to_letta.ts`](https://github.com/letta-ai/claude-subconscious/blob/main/scripts/send_messages_to_letta.ts) | Hook entry point that orchestrates the flow, builds the XML payload, and launches the background SDK worker. |
| [`scripts/conversation_utils.ts`](https://github.com/letta-ai/claude-subconscious/blob/main/scripts/conversation_utils.ts) | Manages sync state (last processed index) to prevent duplicate message transmission. |
| [`scripts/agent_config.ts`](https://github.com/letta-ai/claude-subconscious/blob/main/scripts/agent_config.ts) | Resolves the target Letta agent ID via environment variables, config files, or auto‑import logic. |

## Summary

- The **transcript processing flow** begins with `readTranscript` streaming JSONL files and parsing each line into structured objects.
- **Content extraction** happens in `extractAllContent`, which separates text, thinking blocks, tool calls, and tool results into distinct arrays.
- A **tool name mapping** preserves the relationship between tool usage IDs and their names for accurate result labeling.
- **Message formatting** assigns specific roles (user, assistant, system) and prefixes (e.g., `[Thinking]:`, `[Tool: ...]`) to each content type, applying truncation where necessary.
- The final output is an **XML payload** compatible with the Letta SDK, enabling Claude‑Code sessions to populate Letta agent context.

## Frequently Asked Questions

### What file format does the transcript processing flow expect?

The pipeline expects a **JSONL** (newline‑delimited JSON) file where each line represents a single `TranscriptMessage` object. The `readTranscript` function in [`scripts/transcript_utils.ts`](https://github.com/letta-ai/claude-subconscious/blob/main/scripts/transcript_utils.ts) parses these lines sequentially using `JSON.parse`.

### How does the system distinguish between thinking blocks and regular text?

Inside `extractAllContent`, the code checks the `type` field of each content block. Blocks with `type: "thinking"` are pushed to the `thinkingParts` array, while `type: "text"` blocks populate `textParts`. When formatting for Letta, thinking content is wrapped with the `[Thinking]:` prefix and truncated to 500 characters before being assigned the `assistant` role.

### What happens to tool results in the transcript processing flow?

Tool results (blocks with `type: "tool_result"`) are mapped to their original tool names using the `toolNameMap`, then formatted as `[Tool Result]: <name>]\n<content>` or `[Tool Error]` if the `is_error` flag is true. These are emitted with `role: "system"` to distinguish them from assistant-generated content.

### Where is the sync state stored to prevent duplicate processing?

The sync state—containing the `lastProcessedIndex`—is managed by functions in [`scripts/conversation_utils.ts`](https://github.com/letta-ai/claude-subconscious/blob/main/scripts/conversation_utils.ts). This state tracks which transcript entries have already been sent to Letta, allowing `formatMessagesForLetta` to process only new messages in subsequent hook executions.