# AgentEvent Schema and Unified Telemetry Format in ADR: Complete Guide

> Understand the AgentEvent schema and unified telemetry format in ADR. This guide explains how ADR normalizes AI agent telemetry into a single, self-describing JSON format for comprehensive insights.

- Repository: [Uber Open Source/ADR](https://github.com/uber/ADR)
- Tags: deep-dive
- Published: 2026-08-07

---

**The `AgentEvent` schema is ADR's standardized data model that normalizes telemetry from diverse AI-coding agents into a single, self-describing JSON format.**

AI-coding assistants like Claude Code, Cursor, Cline, Codex, and Warp produce wildly different log structures, making security monitoring and observability a challenge. The **ADR** (Agent Detection & Response) repository solves this through a unified `AgentEvent` schema—defined in [`Sensor/adr_sensor/schemas/agent_event_schema.py`](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/schemas/agent_event_schema.py)—that every parser must emit. This article breaks down the schema structure, its core components, and how to work with it programmatically.

## AgentEvent Schema Structure

The `AgentEvent` class serves as the **canonical data model** for all telemetry flowing through ADR. Every concrete parser in `Sensor/adr_sensor/parsers/` produces a list of these objects, enabling downstream components to process agent activity without caring about the original log format.

### Core Identification Fields

These fields answer **when**, **where**, and **who** for every event:

- **`timestamp`** – `datetime` in UTC when the event occurred
- **`source`** – String identifier of the originating agent (e.g., `"claude"`, `"cursor"`, `"cline"`)
- **`session_id`** – Unique identifier grouping related interactions into a user session

### Chat History (ChatMessage)

The **`chat_history`** field contains a `List[ChatMessage]` representing the full conversation turn sequence. Each `ChatMessage` includes:

| Field | Type | Description |
|-------|------|-------------|
| `role` | `str` | `"user"` or `"assistant"` |
| `content` | `str` | Message text content |
| `tools` | `Optional[List[ToolUsage]]` | Tool calls made by the agent during this turn |
| `sequence_id` | `Optional[int]` | Ordering identifier for complex multi-turn sequences |

The **`ToolUsage`** nested object captures function calls with `tool_name`, `tool_type`, `arguments`, `result`, and `status` fields—critical for security analysis of what actions agents actually execute.

### Metadata Fields

Optional but **highly valuable** for attribution and debugging:

- **`user_id`** – Identity of the human operator
- **`project_path`** – Filesystem location where the agent is running
- **`model`** – Specific model version (e.g., `"claude-2"`, `"gpt-4"`)
- **`hostname`** – Machine where the session occurred
- **`username`** – OS-level user account
- **`raw_log_path`** – Reference to original source log for audit trails

### Session Context

The **`session_context`** field is an `Optional[Dict]` that preserves agent-specific state. Parsers for IDE-integrated tools like Cursor or Warp use this to capture editor-specific metadata that doesn't fit the standard schema.

### Chunking Support

Large sessions are split across multiple `AgentEvent` records using these fields:

- **`is_chunked`** – Boolean indicating if this is part of a chunked session
- **`total_chunks`** – Total number of chunks in the complete session
- **`chunk_sequence`** – Zero-based index of this chunk
- **`is_truncated`** – Flag if content was cut due to size limits

This design ensures **arbitrarily large agent sessions** can be processed without memory pressure, while preserving chronological reconstruction.

## UUID Generation and Content Hashing

Every `AgentEvent` receives a **deterministic UUID** generated via SHA-256 hash of:

1. `hostname`
2. `username`
3. `timestamp` (ISO format)
4. `source`
5. `session_id`
6. Chat history length
7. First 100 characters of each message's content

This yields a **64-character hexadecimal string** that is:
- **Deterministic** – Same inputs always produce same UUID
- **Collision-resistant** – SHA-256 provides strong uniqueness guarantees
- **Traceable** – Derived from session content, not random assignment

The `get_content_hash()` method exposes this directly; `uuid` is the auto-generated property.

## Utility Methods for Downstream Processing

The schema includes **built-in analysis helpers** that security tools and observers rely on:

| Method | Purpose |
|--------|---------|
| `has_meaningful_content()` | Returns `True` if event contains non-empty chat history or tool usage—filters noise from empty keepalive events |
| `get_non_null_fields()` | Returns dict of populated fields for efficient storage |
| `get_summary()` | Generates human-readable one-line summary with source, timestamp, session, turn count, and tool count |
| `get_content_hash()` | Exposes raw SHA-256 hash computation |
| `to_dict()` / `to_json()` | Serialization to Python dict or JSON string |
| `has_tool_usage()` | Quick check if any message contains tool calls |

## Working with AgentEvent: Code Examples

### Creating an AgentEvent from Parsed Data

```python
from datetime import datetime
from adr_sensor.schemas.agent_event_schema import AgentEvent, ChatMessage, ToolUsage

# Build conversation turns

user_msg = ChatMessage(
    role="user",
    content="Refactor this authentication function to use async/await",
)

assistant_msg = ChatMessage(
    role="assistant",
    content="I'll refactor `verify_token()` to be async. Here's the patch:",
    tools=[
        ToolUsage(
            tool_name="apply_patch",
            tool_type="function_call",
            arguments={"file": "auth.py", "patch": "--- a/auth.py\n+++ b/auth.py..."},
            result="Successfully applied patch to auth.py",
            status="success",
        )
    ],
)

# Construct the unified event

event = AgentEvent(
    timestamp=datetime.utcnow(),
    source="claude",
    session_id="session_abc123",
    chat_history=[user_msg, assistant_msg],
    user_id="engineer@company.com",
    project_path="/home/engineer/api-service",
    model="claude-3-opus-20240229",
    hostname="dev-workstation-07",
    username="engineer",
)

# Ready for storage, transmission, or analysis

print(event.to_json(indent=2))

```

### Filtering Noise with Content Checks

```python

# In a batch processing pipeline

from adr_sensor.schemas.agent_event_schema import AgentEvent

def process_events(events: list[AgentEvent]) -> list[AgentEvent]:
    meaningful = []
    for event in events:
        if event.has_meaningful_content():
            meaningful.append(event)
        else:
            # Log skipped events for audit

            logger.debug(f"Skipped empty event: {event.uuid}")
    return meaningful

```

### Generating Summaries for SIEM Dashboards

```python

# Quick visibility into agent activity

for event in recent_events:
    summary = event.get_summary()
    # Output: "[CLAUDE] 2024-01-15 09:23:17 | Session: session_abc123... | Turns: 4 | Tools: 2"

    siem.send_event(summary, severity=calculate_risk(event))

```

### Handling Chunked Sessions

```python

# Reconstructing large sessions from chunks

chunks = fetch_chunks_for_session(session_id="session_xyz789")
chunks.sort(key=lambda e: e.chunk_sequence)

full_history = []
for chunk in chunks:
    full_history.extend(chunk.chat_history)
    if chunk.is_truncated:
        logger.warning(f"Session {session_id} was truncated at chunk {chunk.chunk_sequence}")

```

## Integration with the Parser Architecture

The unified format is enforced through the **abstract base parser** in [`Sensor/adr_sensor/parsers/base_parser.py`](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/parsers/base_parser.py). Every concrete parser (for Claude, Cursor, Cline, etc.) implements:

```python
from abc import ABC, abstractmethod
from typing import List
from adr_sensor.schemas.agent_event_schema import AgentEvent

class BaseParser(ABC):
    @abstractmethod
    def parse(self, raw_log_path: str) -> List[AgentEvent]:
        """Transform agent-specific logs into unified AgentEvent objects."""
        pass
    
    @property
    @abstractmethod
    def supported_sources(self) -> List[str]:
        """Return list of source identifiers this parser handles."""
        pass

```

This contract ensures the **Observer** ([`Sensor/adr_sensor/observer.py`](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/observer.py)) can orchestrate multiple parsers without source-specific logic—each returns `List[AgentEvent]`, which the observer aggregates, deduplicates, and routes to storage or security rules.

## Security and Observability Benefits

The `AgentEvent` unified telemetry format enables:

- **Cross-agent detection rules** – Write once, apply to Claude, Cursor, Codex, etc.
- **Consistent audit trails** – Every action has structured provenance regardless of source tool
- **Tool call analysis** – Centralized visibility into what functions agents actually invoke
- **Session reconstruction** – Complete conversation history for incident investigation
- **Chunked scalability** – Handle enterprise-scale agent deployments without data loss

## Summary

- **`AgentEvent`** in [`Sensor/adr_sensor/schemas/agent_event_schema.py`](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/schemas/agent_event_schema.py) is ADR's canonical schema for normalizing all AI-coding agent telemetry
- **Core fields** (`timestamp`, `source`, `session_id`) provide universal identification; **`chat_history`** captures full conversation context with optional **`ToolUsage`** records
- **Deterministic UUIDs** derived from session content enable reliable deduplication and traceability
- **Chunking fields** support arbitrarily large sessions without memory constraints
- **Utility methods** (`has_meaningful_content`, `get_summary`, `to_json`) streamline downstream processing
- **Base parser contract** enforces that every source-specific parser emits `List[AgentEvent]`, creating true format unification

## Frequently Asked Questions

### What agents does the AgentEvent schema support?

The schema itself is **agent-agnostic**; it supports any AI-coding tool with a corresponding parser. Current implementations exist for Claude Code, Cursor, Cline, Codex CLI, Warp, and Codeium. New agents require only a new parser class that inherits from `BaseParser` and transforms native logs into `AgentEvent` objects.

### How does ADR handle sessions that exceed size limits?

Large sessions use **chunking fields** (`is_chunked`, `total_chunks`, `chunk_sequence`, `is_truncated`) to split across multiple `AgentEvent` records. Each chunk preserves the same `session_id` and shares deterministic UUID components, enabling chronological reconstruction. The `is_truncated` flag alerts analysts when content was cut.

### Can I extend AgentEvent with custom fields?

The **`session_context`** dictionary accepts arbitrary key-value pairs for parser-specific data that doesn't fit standard fields. For schema-wide extensions, modify [`Sensor/adr_sensor/schemas/agent_event_schema.py`](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/schemas/agent_event_schema.py)—the dataclass structure ensures type safety while JSON serialization handles nested objects automatically.

### Where is the UUID actually generated in the source code?

The UUID is computed automatically during `AgentEvent` instantiation via the `get_content_hash()` method, which concatenates hostname, username, timestamp, source, session ID, and message content summaries before SHA-256 hashing. This occurs in [`Sensor/adr_sensor/schemas/agent_event_schema.py`](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/schemas/agent_event_schema.py).