AgentEvent Schema and Unified Telemetry Format in ADR: Complete Guide

The AgentEvent schema is ADR's standardized data model that normalizes telemetry from diverse AI-coding agents into a single, self-describing JSON format.

AI-coding assistants like Claude Code, Cursor, Cline, Codex, and Warp produce wildly different log structures, making security monitoring and observability a challenge. The ADR (Agent Detection & Response) repository solves this through a unified AgentEvent schema—defined in Sensor/adr_sensor/schemas/agent_event_schema.py—that every parser must emit. This article breaks down the schema structure, its core components, and how to work with it programmatically.

AgentEvent Schema Structure

The AgentEvent class serves as the canonical data model for all telemetry flowing through ADR. Every concrete parser in Sensor/adr_sensor/parsers/ produces a list of these objects, enabling downstream components to process agent activity without caring about the original log format.

Core Identification Fields

These fields answer when, where, and who for every event:

  • timestamp – datetime in UTC when the event occurred
  • source – String identifier of the originating agent (e.g., "claude", "cursor", "cline")
  • session_id – Unique identifier grouping related interactions into a user session

Chat History (ChatMessage)

The chat_history field contains a List[ChatMessage] representing the full conversation turn sequence. Each ChatMessage includes:

Field Type Description
role str "user" or "assistant"
content str Message text content
tools Optional[List[ToolUsage]] Tool calls made by the agent during this turn
sequence_id Optional[int] Ordering identifier for complex multi-turn sequences

The ToolUsage nested object captures function calls with tool_name, tool_type, arguments, result, and status fields—critical for security analysis of what actions agents actually execute.

Metadata Fields

Optional but highly valuable for attribution and debugging:

  • user_id – Identity of the human operator
  • project_path – Filesystem location where the agent is running
  • model – Specific model version (e.g., "claude-2", "gpt-4")
  • hostname – Machine where the session occurred
  • username – OS-level user account
  • raw_log_path – Reference to original source log for audit trails

Session Context

The session_context field is an Optional[Dict] that preserves agent-specific state. Parsers for IDE-integrated tools like Cursor or Warp use this to capture editor-specific metadata that doesn't fit the standard schema.

Chunking Support

Large sessions are split across multiple AgentEvent records using these fields:

  • is_chunked – Boolean indicating if this is part of a chunked session
  • total_chunks – Total number of chunks in the complete session
  • chunk_sequence – Zero-based index of this chunk
  • is_truncated – Flag if content was cut due to size limits

This design ensures arbitrarily large agent sessions can be processed without memory pressure, while preserving chronological reconstruction.

UUID Generation and Content Hashing

Every AgentEvent receives a deterministic UUID generated via SHA-256 hash of:

  1. hostname
  2. username
  3. timestamp (ISO format)
  4. source
  5. session_id
  6. Chat history length
  7. First 100 characters of each message's content

This yields a 64-character hexadecimal string that is:

  • Deterministic – Same inputs always produce same UUID
  • Collision-resistant – SHA-256 provides strong uniqueness guarantees
  • Traceable – Derived from session content, not random assignment

The get_content_hash() method exposes this directly; uuid is the auto-generated property.

Utility Methods for Downstream Processing

The schema includes built-in analysis helpers that security tools and observers rely on:

Method Purpose
has_meaningful_content() Returns True if event contains non-empty chat history or tool usage—filters noise from empty keepalive events
get_non_null_fields() Returns dict of populated fields for efficient storage
get_summary() Generates human-readable one-line summary with source, timestamp, session, turn count, and tool count
get_content_hash() Exposes raw SHA-256 hash computation
to_dict() / to_json() Serialization to Python dict or JSON string
has_tool_usage() Quick check if any message contains tool calls

Working with AgentEvent: Code Examples

Creating an AgentEvent from Parsed Data

from datetime import datetime
from adr_sensor.schemas.agent_event_schema import AgentEvent, ChatMessage, ToolUsage

# Build conversation turns

user_msg = ChatMessage(
    role="user",
    content="Refactor this authentication function to use async/await",
)

assistant_msg = ChatMessage(
    role="assistant",
    content="I'll refactor `verify_token()` to be async. Here's the patch:",
    tools=[
        ToolUsage(
            tool_name="apply_patch",
            tool_type="function_call",
            arguments={"file": "auth.py", "patch": "--- a/auth.py\n+++ b/auth.py..."},
            result="Successfully applied patch to auth.py",
            status="success",
        )
    ],
)

# Construct the unified event

event = AgentEvent(
    timestamp=datetime.utcnow(),
    source="claude",
    session_id="session_abc123",
    chat_history=[user_msg, assistant_msg],
    user_id="engineer@company.com",
    project_path="/home/engineer/api-service",
    model="claude-3-opus-20240229",
    hostname="dev-workstation-07",
    username="engineer",
)

# Ready for storage, transmission, or analysis

print(event.to_json(indent=2))

Filtering Noise with Content Checks


# In a batch processing pipeline

from adr_sensor.schemas.agent_event_schema import AgentEvent

def process_events(events: list[AgentEvent]) -> list[AgentEvent]:
    meaningful = []
    for event in events:
        if event.has_meaningful_content():
            meaningful.append(event)
        else:
            # Log skipped events for audit

            logger.debug(f"Skipped empty event: {event.uuid}")
    return meaningful

Generating Summaries for SIEM Dashboards


# Quick visibility into agent activity

for event in recent_events:
    summary = event.get_summary()
    # Output: "[CLAUDE] 2024-01-15 09:23:17 | Session: session_abc123... | Turns: 4 | Tools: 2"

    siem.send_event(summary, severity=calculate_risk(event))

Handling Chunked Sessions


# Reconstructing large sessions from chunks

chunks = fetch_chunks_for_session(session_id="session_xyz789")
chunks.sort(key=lambda e: e.chunk_sequence)

full_history = []
for chunk in chunks:
    full_history.extend(chunk.chat_history)
    if chunk.is_truncated:
        logger.warning(f"Session {session_id} was truncated at chunk {chunk.chunk_sequence}")

Integration with the Parser Architecture

The unified format is enforced through the abstract base parser in Sensor/adr_sensor/parsers/base_parser.py. Every concrete parser (for Claude, Cursor, Cline, etc.) implements:

from abc import ABC, abstractmethod
from typing import List
from adr_sensor.schemas.agent_event_schema import AgentEvent

class BaseParser(ABC):
    @abstractmethod
    def parse(self, raw_log_path: str) -> List[AgentEvent]:
        """Transform agent-specific logs into unified AgentEvent objects."""
        pass
    
    @property
    @abstractmethod
    def supported_sources(self) -> List[str]:
        """Return list of source identifiers this parser handles."""
        pass

This contract ensures the Observer (Sensor/adr_sensor/observer.py) can orchestrate multiple parsers without source-specific logic—each returns List[AgentEvent], which the observer aggregates, deduplicates, and routes to storage or security rules.

Security and Observability Benefits

The AgentEvent unified telemetry format enables:

  • Cross-agent detection rules – Write once, apply to Claude, Cursor, Codex, etc.
  • Consistent audit trails – Every action has structured provenance regardless of source tool
  • Tool call analysis – Centralized visibility into what functions agents actually invoke
  • Session reconstruction – Complete conversation history for incident investigation
  • Chunked scalability – Handle enterprise-scale agent deployments without data loss

Summary

  • AgentEvent in Sensor/adr_sensor/schemas/agent_event_schema.py is ADR's canonical schema for normalizing all AI-coding agent telemetry
  • Core fields (timestamp, source, session_id) provide universal identification; chat_history captures full conversation context with optional ToolUsage records
  • Deterministic UUIDs derived from session content enable reliable deduplication and traceability
  • Chunking fields support arbitrarily large sessions without memory constraints
  • Utility methods (has_meaningful_content, get_summary, to_json) streamline downstream processing
  • Base parser contract enforces that every source-specific parser emits List[AgentEvent], creating true format unification

Frequently Asked Questions

What agents does the AgentEvent schema support?

The schema itself is agent-agnostic; it supports any AI-coding tool with a corresponding parser. Current implementations exist for Claude Code, Cursor, Cline, Codex CLI, Warp, and Codeium. New agents require only a new parser class that inherits from BaseParser and transforms native logs into AgentEvent objects.

How does ADR handle sessions that exceed size limits?

Large sessions use chunking fields (is_chunked, total_chunks, chunk_sequence, is_truncated) to split across multiple AgentEvent records. Each chunk preserves the same session_id and shares deterministic UUID components, enabling chronological reconstruction. The is_truncated flag alerts analysts when content was cut.

Can I extend AgentEvent with custom fields?

The session_context dictionary accepts arbitrary key-value pairs for parser-specific data that doesn't fit standard fields. For schema-wide extensions, modify Sensor/adr_sensor/schemas/agent_event_schema.py—the dataclass structure ensures type safety while JSON serialization handles nested objects automatically.

Where is the UUID actually generated in the source code?

The UUID is computed automatically during AgentEvent instantiation via the get_content_hash() method, which concatenates hostname, username, timestamp, source, session ID, and message content summaries before SHA-256 hashing. This occurs in Sensor/adr_sensor/schemas/agent_event_schema.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →