AgentEvent Schema and Unified Telemetry Format in ADR: Complete Guide
The AgentEvent schema is ADR's standardized data model that normalizes telemetry from diverse AI-coding agents into a single, self-describing JSON format.
AI-coding assistants like Claude Code, Cursor, Cline, Codex, and Warp produce wildly different log structures, making security monitoring and observability a challenge. The ADR (Agent Detection & Response) repository solves this through a unified AgentEvent schema—defined in Sensor/adr_sensor/schemas/agent_event_schema.py—that every parser must emit. This article breaks down the schema structure, its core components, and how to work with it programmatically.
AgentEvent Schema Structure
The AgentEvent class serves as the canonical data model for all telemetry flowing through ADR. Every concrete parser in Sensor/adr_sensor/parsers/ produces a list of these objects, enabling downstream components to process agent activity without caring about the original log format.
Core Identification Fields
These fields answer when, where, and who for every event:
timestamp–datetimein UTC when the event occurredsource– String identifier of the originating agent (e.g.,"claude","cursor","cline")session_id– Unique identifier grouping related interactions into a user session
Chat History (ChatMessage)
The chat_history field contains a List[ChatMessage] representing the full conversation turn sequence. Each ChatMessage includes:
| Field | Type | Description |
|---|---|---|
role |
str |
"user" or "assistant" |
content |
str |
Message text content |
tools |
Optional[List[ToolUsage]] |
Tool calls made by the agent during this turn |
sequence_id |
Optional[int] |
Ordering identifier for complex multi-turn sequences |
The ToolUsage nested object captures function calls with tool_name, tool_type, arguments, result, and status fields—critical for security analysis of what actions agents actually execute.
Metadata Fields
Optional but highly valuable for attribution and debugging:
user_id– Identity of the human operatorproject_path– Filesystem location where the agent is runningmodel– Specific model version (e.g.,"claude-2","gpt-4")hostname– Machine where the session occurredusername– OS-level user accountraw_log_path– Reference to original source log for audit trails
Session Context
The session_context field is an Optional[Dict] that preserves agent-specific state. Parsers for IDE-integrated tools like Cursor or Warp use this to capture editor-specific metadata that doesn't fit the standard schema.
Chunking Support
Large sessions are split across multiple AgentEvent records using these fields:
is_chunked– Boolean indicating if this is part of a chunked sessiontotal_chunks– Total number of chunks in the complete sessionchunk_sequence– Zero-based index of this chunkis_truncated– Flag if content was cut due to size limits
This design ensures arbitrarily large agent sessions can be processed without memory pressure, while preserving chronological reconstruction.
UUID Generation and Content Hashing
Every AgentEvent receives a deterministic UUID generated via SHA-256 hash of:
hostnameusernametimestamp(ISO format)sourcesession_id- Chat history length
- First 100 characters of each message's content
This yields a 64-character hexadecimal string that is:
- Deterministic – Same inputs always produce same UUID
- Collision-resistant – SHA-256 provides strong uniqueness guarantees
- Traceable – Derived from session content, not random assignment
The get_content_hash() method exposes this directly; uuid is the auto-generated property.
Utility Methods for Downstream Processing
The schema includes built-in analysis helpers that security tools and observers rely on:
| Method | Purpose |
|---|---|
has_meaningful_content() |
Returns True if event contains non-empty chat history or tool usage—filters noise from empty keepalive events |
get_non_null_fields() |
Returns dict of populated fields for efficient storage |
get_summary() |
Generates human-readable one-line summary with source, timestamp, session, turn count, and tool count |
get_content_hash() |
Exposes raw SHA-256 hash computation |
to_dict() / to_json() |
Serialization to Python dict or JSON string |
has_tool_usage() |
Quick check if any message contains tool calls |
Working with AgentEvent: Code Examples
Creating an AgentEvent from Parsed Data
from datetime import datetime
from adr_sensor.schemas.agent_event_schema import AgentEvent, ChatMessage, ToolUsage
# Build conversation turns
user_msg = ChatMessage(
role="user",
content="Refactor this authentication function to use async/await",
)
assistant_msg = ChatMessage(
role="assistant",
content="I'll refactor `verify_token()` to be async. Here's the patch:",
tools=[
ToolUsage(
tool_name="apply_patch",
tool_type="function_call",
arguments={"file": "auth.py", "patch": "--- a/auth.py\n+++ b/auth.py..."},
result="Successfully applied patch to auth.py",
status="success",
)
],
)
# Construct the unified event
event = AgentEvent(
timestamp=datetime.utcnow(),
source="claude",
session_id="session_abc123",
chat_history=[user_msg, assistant_msg],
user_id="engineer@company.com",
project_path="/home/engineer/api-service",
model="claude-3-opus-20240229",
hostname="dev-workstation-07",
username="engineer",
)
# Ready for storage, transmission, or analysis
print(event.to_json(indent=2))
Filtering Noise with Content Checks
# In a batch processing pipeline
from adr_sensor.schemas.agent_event_schema import AgentEvent
def process_events(events: list[AgentEvent]) -> list[AgentEvent]:
meaningful = []
for event in events:
if event.has_meaningful_content():
meaningful.append(event)
else:
# Log skipped events for audit
logger.debug(f"Skipped empty event: {event.uuid}")
return meaningful
Generating Summaries for SIEM Dashboards
# Quick visibility into agent activity
for event in recent_events:
summary = event.get_summary()
# Output: "[CLAUDE] 2024-01-15 09:23:17 | Session: session_abc123... | Turns: 4 | Tools: 2"
siem.send_event(summary, severity=calculate_risk(event))
Handling Chunked Sessions
# Reconstructing large sessions from chunks
chunks = fetch_chunks_for_session(session_id="session_xyz789")
chunks.sort(key=lambda e: e.chunk_sequence)
full_history = []
for chunk in chunks:
full_history.extend(chunk.chat_history)
if chunk.is_truncated:
logger.warning(f"Session {session_id} was truncated at chunk {chunk.chunk_sequence}")
Integration with the Parser Architecture
The unified format is enforced through the abstract base parser in Sensor/adr_sensor/parsers/base_parser.py. Every concrete parser (for Claude, Cursor, Cline, etc.) implements:
from abc import ABC, abstractmethod
from typing import List
from adr_sensor.schemas.agent_event_schema import AgentEvent
class BaseParser(ABC):
@abstractmethod
def parse(self, raw_log_path: str) -> List[AgentEvent]:
"""Transform agent-specific logs into unified AgentEvent objects."""
pass
@property
@abstractmethod
def supported_sources(self) -> List[str]:
"""Return list of source identifiers this parser handles."""
pass
This contract ensures the Observer (Sensor/adr_sensor/observer.py) can orchestrate multiple parsers without source-specific logic—each returns List[AgentEvent], which the observer aggregates, deduplicates, and routes to storage or security rules.
Security and Observability Benefits
The AgentEvent unified telemetry format enables:
- Cross-agent detection rules – Write once, apply to Claude, Cursor, Codex, etc.
- Consistent audit trails – Every action has structured provenance regardless of source tool
- Tool call analysis – Centralized visibility into what functions agents actually invoke
- Session reconstruction – Complete conversation history for incident investigation
- Chunked scalability – Handle enterprise-scale agent deployments without data loss
Summary
AgentEventinSensor/adr_sensor/schemas/agent_event_schema.pyis ADR's canonical schema for normalizing all AI-coding agent telemetry- Core fields (
timestamp,source,session_id) provide universal identification;chat_historycaptures full conversation context with optionalToolUsagerecords - Deterministic UUIDs derived from session content enable reliable deduplication and traceability
- Chunking fields support arbitrarily large sessions without memory constraints
- Utility methods (
has_meaningful_content,get_summary,to_json) streamline downstream processing - Base parser contract enforces that every source-specific parser emits
List[AgentEvent], creating true format unification
Frequently Asked Questions
What agents does the AgentEvent schema support?
The schema itself is agent-agnostic; it supports any AI-coding tool with a corresponding parser. Current implementations exist for Claude Code, Cursor, Cline, Codex CLI, Warp, and Codeium. New agents require only a new parser class that inherits from BaseParser and transforms native logs into AgentEvent objects.
How does ADR handle sessions that exceed size limits?
Large sessions use chunking fields (is_chunked, total_chunks, chunk_sequence, is_truncated) to split across multiple AgentEvent records. Each chunk preserves the same session_id and shares deterministic UUID components, enabling chronological reconstruction. The is_truncated flag alerts analysts when content was cut.
Can I extend AgentEvent with custom fields?
The session_context dictionary accepts arbitrary key-value pairs for parser-specific data that doesn't fit standard fields. For schema-wide extensions, modify Sensor/adr_sensor/schemas/agent_event_schema.py—the dataclass structure ensures type safety while JSON serialization handles nested objects automatically.
Where is the UUID actually generated in the source code?
The UUID is computed automatically during AgentEvent instantiation via the get_content_hash() method, which concatenates hostname, username, timestamp, source, session ID, and message content summaries before SHA-256 hashing. This occurs in Sensor/adr_sensor/schemas/agent_event_schema.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →