How Session Parsers Extract Data from Various Agent Formats in AgentsView

Session parsers extract data from various agent formats by streaming JSONL files line-by-line, extracting specific fields with gjson, and applying agent-specific transformations to populate unified ParsedSession and ParsedMessage structs.

The agentsview project normalizes conversation logs from dozens of AI coding agents into a queryable SQLite database. At the heart of this pipeline are specialized session parsers that handle the unique JSONL structures of each agent while conforming to a common extraction interface defined in internal/parser/types.go.

The Parser Architecture: Registry and Provider Interface

The system uses a registry-based architecture to manage disparate agent formats. In internal/parser/types.go, the Registry slice defines an AgentDef for each supported agent, specifying metadata and extraction behavior.

var Registry = []AgentDef{
    {Type: AgentClaude, DisplayName: "Claude Code", FileBased: true},
    {Type: AgentCopilot, DisplayName: "Copilot", FileBased: true},
    // ... additional agents
}

The FileBased flag signals the sync engine to invoke a file-system parser. Each agent implements a provider that exposes the ParseUploadedTranscript method, which serves as the uniform entry point for extraction. For example, internal/parser/claude_provider.go wraps the Claude-specific logic:

func (p *claudeProvider) ParseUploadedTranscript(
    path, project, machine string,
) ([]ParseResult, error) {
    return claudeParseWithExclusions(path, project, machine)
}

Similarly, internal/parser/copilot_provider.go delegates to parseCopilot, ensuring every agent follows the same contract while allowing format-specific implementations.

Line-by-Line Ingestion: The lineReader Utility

To handle arbitrarily large session files without exhausting memory, parsers use a reusable lineReader defined in internal/parser/linereader.go. This utility buffers input and safely accumulates lines that exceed the internal buffer size (up to 64 MiB), preventing crashes from malformed or extreme JSONL entries.

lr := newLineReader(f, maxLineSize)
for {
    line, ok := lr.next()
    if !ok { break }
    // Process raw JSON line
}

This streaming approach ensures that multi-gigabyte session files from long-running agent conversations can be processed with constant memory overhead.

JSON Field Extraction with gjson

Rather than unmarshaling entire JSON objects into Go structs, parsers extract only the required fields using the gjson library. This optimization avoids the overhead of parsing irrelevant agent-specific metadata.

In internal/parser/claude.go, extraction looks like this:

entryType := gjson.Get(line, "type").Str
if ts := extractTimestamp(line); !ts.IsZero() {
    // Process timestamp
}

Parsers pull out universal schema fields such as id, type, timestamp, content, tool_use_id, model, and token_usage, leaving the raw JSON structure otherwise intact.

Agent-Specific Parsing Strategies

While the core extraction pattern is consistent, each agent requires specialized handling for its unique conversation model.

Claude Code: DAG Construction and Tool Result Handling

The Claude parser in internal/parser/claude.go handles complex conversation trees. It builds a directed acyclic graph (DAG) from uuid and parentUuid fields to detect conversation forks and branches.

The implementation also manages queued commands—attachments that arrive while a tool call is running—and resolves file references using persistedToolResultPathRe. The function claudeParseWithExclusions orchestrates this logic, returning fully resolved ParseResult slices that include tool output inlined into the message stream.

GitHub Copilot: Tool Usage and Ordinal Sorting

In internal/parser/copilot.go, the parseCopilot function handles Copilot's JSONL format, which emits separate records for user messages, assistant responses, and tool invocations. The parser aggregates usage events emitted by the Copilot extension and sorts messages by the ordinal field to reconstruct the correct chronological order regardless of file write sequence.

OpenAI Codex: Session Index Resolution

The Codex parser demonstrates handling external metadata files. In internal/parser/codex.go, the ParseCodexSessionIndexTitles function reads a companion session_index.jsonl file to map raw session IDs to human-readable titles, enriching the ParsedSession before persistence.

From Raw Lines to Unified Schema: The ParseResult Pipeline

After processing all lines, parsers construct a ParseResult containing a ParsedSession, a slice of ParsedMessage objects, and optional ParsedUsageEvent entries. The sync engine then persists these to SQLite.

For token usage, parsers may emit explicit fields, but when absent, the generic helper InferTokenPresence (in internal/parser/types.go) inspects the raw token_usage JSON to set boolean flags like HasContextTokens and HasOutputTokens.

Implementation Example: Parsing Claude and Copilot Sessions

The following example demonstrates how to invoke the parsers directly for Claude and Copilot session files:

package main

import (
    "fmt"
    "github.com/kenn-io/agentsview/internal/parser"
)

func main() {
    // Parse a Claude session file
    claude := &parser.ClaudeProvider{}
    results, err := claude.ParseUploadedTranscript(
        "/home/user/.claude/projects/myproject/2023-09-01-abc123.jsonl",
        "myproject",
        "my-machine",
    )
    if err != nil {
        panic(err)
    }
    
    for _, r := range results {
        fmt.Printf("Session %s – %d messages, %d tokens total\n",
            r.Session.ID, len(r.Messages), r.Session.TotalOutputTokens)
    }

    // Parse a Copilot session file
    copilot := &parser.CopilotProvider{}
    cpResults, _ := copilot.ParseUploadedTranscript(
        "/home/user/.copilot/session-state/2024-01-15-xyz.jsonl",
        "my-project",
        "my-machine",
    )
    fmt.Println("Copilot parsed", len(cpResults), "sessions")
}

Summary

  • Registry-driven architecture: internal/parser/types.go defines agents via AgentDef structs, enabling a plugin-like system where FileBased flags trigger filesystem parsing.
  • Streaming ingestion: The lineReader in internal/parser/linereader.go processes JSONL files line-by-line with memory safety for large inputs.
  • Selective extraction: Parsers use gjson to extract only necessary fields from each JSON line, avoiding expensive unmarshaling of full objects.
  • Agent-specific logic: Claude handles DAGs and tool results (claude.go), Copilot manages ordinals and usage events (copilot.go), and Codex resolves external index files (codex.go).
  • Unified output: All parsers return ParseResult structs containing ParsedSession and ParsedMessage data, normalized for SQLite storage.

Frequently Asked Questions

What is the role of the Provider interface in the parser architecture?

The Provider interface defines the contract that every agent parser must implement, specifically the ParseUploadedTranscript method. This abstraction allows the sync engine in internal/parser/discovery.go to treat all agents uniformly—whether parsing Claude, Copilot, or Codex files—by calling the same method signature regardless of the underlying format.

How does AgentsView handle extremely large session files without running out of memory?

The parsers use a custom lineReader utility located in internal/parser/linereader.go. This reader buffers input and accumulates oversized lines in a string builder rather than loading the entire file into memory, enabling safe processing of multi-gigabyte JSONL files with constant memory overhead.

Why does the Claude parser use gjson instead of standard JSON unmarshaling?

The Claude parser uses gjson to perform selective field extraction because Claude session files contain complex nested structures including conversation DAGs and tool attachments. Extracting only the uuid, parentUuid, type, and timestamp fields with gjson.Get is significantly faster than unmarshaling the entire JSON object into a Go struct, especially when processing millions of lines.

What distinguishes the Copilot parser's handling of conversation ordering?

The Copilot parser explicitly sorts messages by the ordinal field after extraction. Because Copilot writes conversation events asynchronously to the JSONL file, the physical line order may not match the logical conversation flow. The parser in internal/parser/copilot.go aggregates these events and sorts them by ordinal before constructing the final ParsedMessage slice.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →