What Is the Database Schema Used by pi-web? A Complete Guide to Its JSON-Lines Architecture
pi-web does not use a traditional relational database; instead, it persists all session data in JSON-Lines (.jsonl) files under ~/.pi/agent/sessions/, using a strict schema enforced by TypeScript types in lib/pi-types.ts and parsing logic in lib/session-reader.ts.
The agegr/pi-web repository implements a file-based persistence layer where every chat session is a standalone .jsonl file. This database schema used by pi-web is defined by the pi-SDK specification and documented in the project's AGENTS.md file, ensuring that each line in a session file represents a self-contained, schema-validated entry.
Core JSON-Lines Schema Structure
The schema supports six distinct entry types, each serving a specific function in the conversation lifecycle. Every entry shares common fields—type, id, and parentId—while type-specific fields carry the payload.
Session Header Entry
The session entry acts as a file header and must be the first line in any .jsonl file. It establishes the session identity and version.
// From lib/pi-types.ts
interface SessionHeader {
type: "session";
version: 3; // Current schema version
id: string; // UUID
timestamp: string; // ISO 8601
cwd: string; // Working directory
parentSession: string; // Path to parent file or empty string
}
The parentSession field enables session forking by linking to an ancestor file, allowing pi-web to maintain conversation lineage without duplicating history.
Message Entries
Message entries store the actual conversation content using an OpenAI-compatible schema. The message object contains role (user, assistant, or system) and content fields.
{"type":"message","id":"a1b2c3d4","parentId":"previous","message":{"role":"user","content":"Explain the database schema"}}
Tool results are encoded as messages with role: "toolResult" and include a toolCallId that references the originating tool invocation, enabling the reconstruction of tool call chains.
Model Changes and Configuration
The model_change entry records LLM provider switches during a session. This captures the provider name (e.g., "anthropic") and modelId (e.g., "claude-sonnet-4-6") along with a timestamp.
interface ModelChangeEntry {
type: "model_change";
id: string;
parentId: string | null;
provider: string;
modelId: string;
timestamp: string;
}
The session_info entry provides optional metadata, allowing users to assign human-readable names to sessions via the name field.
Compaction Records
Long-running sessions use compaction entries to manage file size. These entries summarize truncated history and mark the transition point for partial file reconstruction.
interface CompactionEntry {
type: "compaction";
id: string;
parentId: string;
summary: string; // Summarized content
firstKeptEntryId: string; // ID of first surviving entry
tokensBefore: number; // Token count prior to compaction
}
As implemented in lib/compaction-summary.ts, this mechanism allows pi-web to load recent conversation context without parsing entire file histories.
How the Schema Is Enforced in Code
Type safety and validation are enforced through a three-layer architecture defined in the repository's source code.
Type Definitions: The lib/pi-types.ts file contains the TypeScript interfaces (SessionHeader, MessageEntry, ModelChangeEntry, CompactionEntry) that guarantee compile-time correctness for all database operations.
Parsing and Normalization: The lib/session-reader.ts module reads .jsonl files line-by-line, parsing each JSON object and applying normalizeToolCalls() from lib/normalize.ts. This reconciliation step handles any mismatches between the on-disk format and the in-memory ToolCallContent type.
Write Operations: When creating new sessions, lib/rpc-manager.ts serializes the SessionHeader interface and writes it to ~/.pi/agent/sessions/<uuid>.jsonl. Subsequent entries are appended atomically using the same serialization logic.
Working with the Schema: Practical Examples
Creating a New Session File
The following pattern from lib/rpc-manager.ts demonstrates initializing a session with the required header:
import { SessionHeader } from "./lib/pi-types";
import { writeFileSync } from "fs";
const header: SessionHeader = {
type: "session",
version: 3,
id: crypto.randomUUID(),
timestamp: new Date().toISOString(),
cwd: process.cwd(),
parentSession: "", // Empty for root sessions
};
writeFileSync(sessionPath, JSON.stringify(header) + "\n");
Appending Conversation Messages
Messages are appended as single JSON lines without rewriting the entire file:
import { MessageEntry } from "./lib/pi-types";
import { appendFileSync } from "fs";
const entry: MessageEntry = {
type: "message",
id: crypto.randomUUID().slice(0, 8),
parentId: previousEntryId,
message: {
role: "assistant",
content: "The schema uses append-only JSON-Lines."
}
};
appendFileSync(sessionPath, JSON.stringify(entry) + "\n");
Recording Model Switches
To persist a provider change during a conversation:
import { ModelChangeEntry } from "./lib/pi-types";
const change: ModelChangeEntry = {
type: "model_change",
id: crypto.randomUUID().slice(0, 8),
parentId: lastEntryId,
provider: "openai",
modelId: "gpt-4",
timestamp: new Date().toISOString()
};
appendFileSync(sessionPath, JSON.stringify(change) + "\n");
Summary
- pi-web uses JSON-Lines files, not SQL databases, storing data in
~/.pi/agent/sessions/with one file per conversation. - Six entry types define the schema:
session,message,model_change,toolResult(embedded in messages),compaction, andsession_info. - Type safety is enforced through TypeScript interfaces in
lib/pi-types.tsand runtime parsing inlib/session-reader.ts. - Compaction entries prevent unbounded file growth by summarizing and truncating older history while maintaining conversation continuity via
firstKeptEntryId. - Session forking is supported through the
parentSessionfield in headers, enabling non-destructive branching of conversations.
Frequently Asked Questions
Does pi-web use SQLite or PostgreSQL for storage?
No. According to the agegr/pi-web source code, the system explicitly avoids traditional relational databases. All persistent state is stored in local .jsonl files using the JSON-Lines format, making the architecture entirely file-based and portable across systems without database server dependencies.
Where are the session files physically located?
Session files are stored in the user's home directory under ~/.pi/agent/sessions/. Each session receives a unique UUID filename with the .jsonl extension. The cwd field in the session header records the working directory where the session was initiated, but the files themselves reside in this centralized location.
How does pi-web handle conversation history that grows too large?
The schema implements compaction through entries written by lib/compaction-summary.ts. When a session file exceeds size thresholds, older entries are summarized into a compaction record containing a summary field and firstKeptEntryId. The reader in lib/session-reader.ts uses this marker to reconstruct context without loading the entire file history.
What is the purpose of the parentId field in entries?
The parentId field establishes a linked list structure within the JSON-Lines file, where each entry references its predecessor. This append-only design allows pi-web to reconstruct conversation state sequentially and handle branching logic when tool calls or compactions insert entries mid-stream. For session headers, the parentSession field serves a similar purpose by referencing external parent files during session forks.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →