How Session Persistence and Automatic Compaction Maintain Context Across Conversation Turns in Prime Agent

Prime Agent stores every conversation turn as JSONL entries on disk and automatically compacts older messages into summaries when approaching token limits, ensuring continuous context across unlimited conversation turns without exceeding LLM context windows.

Prime Agent, developed by PrimeIntellect-ai, implements a two-layer architecture that combines durable session storage with intelligent summarization. This system ensures that session persistence and automatic compaction work together to maintain logical conversation continuity without manual intervention. By persisting raw message history to ~/.pi/agent/sessions/ and algorithmically compressing older turns via the compaction engine, the agent handles extended multi-turn workflows that would otherwise exceed model context constraints.

How Session Persistence Works

Session persistence guarantees that conversation state survives crashes, restarts, and arbitrary execution lengths. The system appends every exchange to a durable JSONL file that reconstructs the exact conversation tree on reload.

Session Directory and File Structure

When initializing a session, the SessionManager resolves a dedicated directory under the user's home folder. The getDefaultSessionDir() function constructs a path at ~/.pi/agent/sessions/<encoded-cwd>/, encoding the current working directory to isolate projects by location.

Inside this directory, session data lives in a JSON Lines file (*.jsonl). Each line represents a discrete entry in the conversation history. According to packages/coding-agent/src/core/session-manager.ts (lines 125-130), the system recognizes multiple entry types: user, assistant, tool, label, artifact, and critically, the compaction entry type that stores summaries replacing collapsed message blocks.

Live Updates and Durability

An AgentSession subscribes to persistence events via session.subscribe(() => {}). On every new message, SessionManager appends a JSON line and flushes the file immediately. This mechanism, implemented in packages/coding-agent/src/core/agent-session.ts, guarantees that even if the process crashes mid-conversation, the session file contains every turn up to the last flush.

Session Loading and Reconstruction

At startup, the agent calls SessionManager.create() (or inMemory() for testing), which invokes loadEntriesFromFile to read all JSONL entries. The manager reconstructs the in-memory tree and exposes buildSessionContext() to the rest of the system, used by both Interactive and Daemon modes to resume exactly where the conversation left off.

Automatic Compaction Logic

While persistence stores history, automatic compaction prevents token limit exhaustion by summarizing older conversation segments into dense semantic representations.

Token Threshold Triggers

The compaction system triggers based on configurable thresholds defined in packages/coding-agent/src/core/settings-manager.ts. The autoCompaction configuration includes enabled, reserveTokens, and keepRecentTokens parameters.

After each turn, AgentSession._getThresholdContextTokens() (lines 2060-2065 in agent-session.ts) calculates projected token usage. When usage plus the reserved margin exceeds the model's context window, the session automatically initiates compaction.

The Compaction Execution Flow

The compaction process follows a strict RPC-mediated pipeline:

  1. Initiation: AgentSession.compact() sends a compact command via the RPC client (packages/coding-agent/src/modes/rpc/rpc-client.ts).
  2. Server Processing: The RPC mode handler (packages/coding-agent/src/modes/rpc/rpc-mode.ts) invokes connection.compact(customInstructions).
  3. LLM Summarization: The compaction engine (packages/coding-agent/src/core/compaction/compaction.ts) extracts entries outside the "keep recent" window, generates a compact prompt using utilities from core/compaction/utils.ts, and calls the LLM.
  4. Result Persistence: The LLM returns a CompactionResult containing summary text, token count, and optional instructions. SessionManager writes this as a compaction entry (type "compaction") and removes the original messages from the in-memory tree.

Preserving Critical State Across Summaries

Not all history can be summarized away. The compaction engine explicitly preserves "sticky" entries including artifact paths, tool results, and IPython variables. The preserveFileOperations logic (lines 34-38 in compaction.ts) copies file operations from prior compactions and current tool calls into the new summary context. This ensures that after compaction, the agent retains access to persisted files and computational state despite the compressed dialogue history.

User Feedback During Compaction

The TUI (packages/tui/src/tui.ts, around line 2655) listens for compaction_start and compaction_end events, displaying status lines like "Auto-compacting…". After completion, a CompactionSummaryMessageComponent renders the summary and token savings, providing visibility into context window management.

Integrating Persistence and Compaction

Session persistence and automatic compaction operate as complementary layers that maintain context across conversation turns:

  1. Turn N: User input triggers AgentSession.appendMessage, which writes raw JSONL to disk immediately.
  2. Turn N+1: After the model responds, the session checks token usage against thresholds.
  3. Compaction Trigger: If limits are exceeded, the system summarizes all messages before the keepRecentTokens window, writes a compaction entry, and prunes replaced messages from memory.
  4. Subsequent Context: The model receives the most recent raw messages plus compactionSummary entries acting as placeholders for collapsed history. Since summaries persist in the JSONL file, restarts reconstruct the identical compacted view.

This architecture guarantees loss-less logical continuity while enforcing hard context window constraints.

Configuration and Implementation Examples

Create or resume a persisted session for the current working directory:

import { SessionManager } from "prime-agent/packages/coding-agent/src/core/session-manager.js";

const sessionMgr = SessionManager.create();      // uses ~/.pi/agent/sessions/<cwd>
const session = await sessionMgr.newSession();   // returns an AgentSession
session.subscribe(() => console.log("session persisted"));

Configure automatic compaction thresholds:

import { SettingsManager } from "prime-agent/packages/coding-agent/src/core/settings-manager.js";

SettingsManager.global().update({
  compaction: {
    enabled: true,          // turn on auto-compaction
    reserveTokens: 16384,   // keep this many tokens free for the next turn
    keepRecentTokens: 20000 // never compact the most recent N tokens
  }
});

Manually trigger compaction with custom instructions:

await session.compact("Summarize the discussion about file-system helpers.");

Inspect the persisted session file directly:

import { readFileSync } from "fs";
const raw = readFileSync(`${process.env.HOME}/.pi/agent/sessions/${encodeURIComponent(process.cwd())}/session.jsonl`, "utf8");
console.log(raw.split("\n").slice(-5).join("\n")); // shows latest entries, including any "compaction"

Summary

  • Prime Agent stores conversation history in JSONL files at ~/.pi/agent/sessions/<encoded-cwd>/, with each turn written immediately via SessionManager.
  • Entry types include user, assistant, tool, and compaction, allowing the system to interleave raw messages with summarization records.
  • Automatic compaction triggers when AgentSession._getThresholdContextTokens() detects imminent context limit breaches, configurable via SettingsManager (reserveTokens, keepRecentTokens).
  • The compaction pipeline spans RPC client calls (rpc-client.ts), server-side processing (rpc-mode.ts), and the LLM-powered summarization engine (compaction.ts).
  • State preservation ensures file operations, artifacts, and tool results survive compaction via preserveFileOperations logic.
  • Context reconstruction loads both recent raw messages and compaction entries, maintaining logical continuity across restarts and unlimited conversation turns.

Frequently Asked Questions

How does Prime Agent handle session recovery after a crash?

Prime Agent recovers sessions by reloading the JSONL file from ~/.pi/agent/sessions/<encoded-cwd>/. The SessionManager.create() method calls loadEntriesFromFile, which reconstructs the conversation tree including any previous compaction entries. Because SessionManager flushes every entry to disk immediately upon receiving new messages via the subscription mechanism in AgentSession, the replay resumes from the exact last persisted turn.

What happens to file references and tool outputs when messages get compacted?

The compaction engine explicitly preserves these through preserveFileOperations logic (lines 34-38 in packages/coding-agent/src/core/compaction/compaction.ts). When summarizing older dialogue, the system copies file operations from prior compactions and current tool calls into the new summary context. This ensures that artifact paths, IPython variables, and tool results remain accessible to the agent even after their original conversation turns are replaced by compact summaries.

Can users manually trigger compaction or adjust when it happens automatically?

Yes. Users can manually invoke session.compact(customInstructions) to summarize conversation history on demand with specific guidance. For automatic behavior, the SettingsManager exposes autoCompaction configuration including enabled, reserveTokens (buffer reserved for upcoming turns), and keepRecentTokens (recent messages exempt from compaction). These settings control when AgentSession._getThresholdContextTokens() triggers automatic compaction after each turn.

How does the compaction process affect the context window during active conversations?

During compaction, the system maintains continuity by retaining the most recent keepRecentTokens worth of raw messages while replacing older history with a dense summary. The compaction entry written to the session file acts as a placeholder containing the semantic content of collapsed turns. When building context for the LLM, buildSessionContext() interleaves these summaries with recent raw entries, ensuring the model receives relevant historical information without exceeding token limits. The TUI displays compaction_start and compaction_end events to indicate when this process occurs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →