Claude‑Code Transcript Processing Flow: Extracting Messages, Thinking Blocks, and Tool Usage
The transcript processing flow in letta-ai/claude-subconscious reads JSONL transcript files line‑by‑line, extracts user messages, assistant thinking blocks, and tool interactions through specialized parser functions, then formats them into Letta-compatible XML payloads for agent consumption.
When a Claude‑Code session ends, the send‑messages‑to‑Letta hook processes the conversation history to build a compact, context‑rich payload. This pipeline handles plain text, internal reasoning, and tool execution traces according to the source code in scripts/transcript_utils.ts and scripts/send_messages_to_letta.ts.
How the Transcript Processing Pipeline Works
Step 1: Loading the JSONL Transcript
The readTranscript function in scripts/transcript_utils.ts (lines 63‑87) streams the transcript file line‑by‑line. It parses each line with JSON.parse and returns an array of TranscriptMessage objects ready for content extraction.
import { readTranscript } from './transcript_utils.js';
const messages = await readTranscript('path/to/transcript.jsonl');
// Returns: TranscriptMessage[]
Step 2: Extracting Content Blocks with extractAllContent
The extractAllContent function (lines 92‑141) inspects the content field of each transcript entry. This field can be a plain string or an array of typed blocks. The function discriminates four block types:
type: "text"– Plain user or assistant text.type: "thinking"– Internal reasoning generated by Claude.type: "tool_use"– Tool invocations with name and input parameters.type: "tool_result"– Outputs or errors returned from tool execution.
Each fragment is stored in an ExtractedContent object containing textParts, thinkingParts, toolUses, and toolResults arrays.
import { extractAllContent } from './transcript_utils.js';
const extracted = extractAllContent(singleMessage);
console.log(extracted.text); // Plain text content
console.log(extracted.thinking); // Array of thinking blocks
console.log(extracted.toolUses); // [{name: 'Read', input: {file_path: 'src/app.ts'}}]
console.log(extracted.toolResults); // [{tool_use_id: '...', content: '...', is_error: false}]
Step 3: Mapping Tool Usage IDs
While iterating through the transcript, the pipeline builds a toolNameMap (lines 155‑162) that associates each tool_use_id with its corresponding tool name. This mapping allows later tool_result blocks to be rendered with the original tool name even when the result appears in a separate message.
Step 4: Formatting Letta‑Ready Messages
The formatMessagesForLetta function (lines 155‑282) processes entries starting from the last synchronized index. It transforms extracted content into role‑based messages:
- User messages – Emitted as
{role: "user", text: ...}. - Assistant thinking – Wrapped as
[Thinking]: …, truncated to 500 characters, and emitted as{role: "assistant", text: ...}. - Assistant tool calls – Rendered as
[Tool: <name>] <summary>with concise input summaries (e.g., file paths forRead, commands forBash), emitted as{role: "assistant", text: ...}. - Tool results – Rendered as
[Tool Result]: <toolName>]\n<content>or[Tool Error]ifis_erroris true, emitted as{role: "system", text: ...}with truncated payloads.
import { formatMessagesForLetta } from './transcript_utils.js';
import { loadSyncState } from './conversation_utils.js';
const state = loadSyncState(process.cwd(), sessionId);
const lettaMessages = formatMessagesForLetta(messages, state.lastProcessedIndex);
/*
[
{ role: 'user', text: 'Can you refactor this?' },
{ role: 'assistant', text: '[Thinking]: I will split the file...' },
{ role: 'assistant', text: '[Tool: Read] src/utils.ts' },
{ role: 'system', text: '[Tool Result]: Read]\nconst foo = …' }
]
*/
Step 5: Converting to Letta XML
The hook in scripts/send_messages_to_letta.ts (lines 84‑90) converts the formatted message array into XML. Each entry becomes <message role="…">…</message> with XML‑escaped text before being handed to the Letta SDK worker.
const xmlPayload = lettaMessages
.map(m => `<message role="${m.role}">${escapeXml(m.text)}</message>`)
.join('');
Key Implementation Files
| File | Purpose |
|---|---|
scripts/transcript_utils.ts |
Core utilities for reading JSONL transcripts, extracting content blocks, and formatting Letta‑ready messages. |
scripts/send_messages_to_letta.ts |
Hook entry point that orchestrates the flow, builds the XML payload, and launches the background SDK worker. |
scripts/conversation_utils.ts |
Manages sync state (last processed index) to prevent duplicate message transmission. |
scripts/agent_config.ts |
Resolves the target Letta agent ID via environment variables, config files, or auto‑import logic. |
Summary
- The transcript processing flow begins with
readTranscriptstreaming JSONL files and parsing each line into structured objects. - Content extraction happens in
extractAllContent, which separates text, thinking blocks, tool calls, and tool results into distinct arrays. - A tool name mapping preserves the relationship between tool usage IDs and their names for accurate result labeling.
- Message formatting assigns specific roles (user, assistant, system) and prefixes (e.g.,
[Thinking]:,[Tool: ...]) to each content type, applying truncation where necessary. - The final output is an XML payload compatible with the Letta SDK, enabling Claude‑Code sessions to populate Letta agent context.
Frequently Asked Questions
What file format does the transcript processing flow expect?
The pipeline expects a JSONL (newline‑delimited JSON) file where each line represents a single TranscriptMessage object. The readTranscript function in scripts/transcript_utils.ts parses these lines sequentially using JSON.parse.
How does the system distinguish between thinking blocks and regular text?
Inside extractAllContent, the code checks the type field of each content block. Blocks with type: "thinking" are pushed to the thinkingParts array, while type: "text" blocks populate textParts. When formatting for Letta, thinking content is wrapped with the [Thinking]: prefix and truncated to 500 characters before being assigned the assistant role.
What happens to tool results in the transcript processing flow?
Tool results (blocks with type: "tool_result") are mapped to their original tool names using the toolNameMap, then formatted as [Tool Result]: <name>]\n<content> or [Tool Error] if the is_error flag is true. These are emitted with role: "system" to distinguish them from assistant-generated content.
Where is the sync state stored to prevent duplicate processing?
The sync state—containing the lastProcessedIndex—is managed by functions in scripts/conversation_utils.ts. This state tracks which transcript entries have already been sent to Letta, allowing formatMessagesForLetta to process only new messages in subsequent hook executions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →