# Best Practices for Analyzing AI Agent session_id and Timestamp Patterns in ADR

> Master AI agent session analysis in ADR. Normalize timestamps to UTC, prefix session IDs, and use deterministic filenames for reliable tracking and change detection. Learn best practices now.

- Repository: [Uber Open Source/ADR](https://github.com/uber/ADR)
- Tags: best-practices
- Published: 2026-08-07

---

**Always normalize timestamps to UTC using `normalize_timestamp()`, prefix `session_id` values to identify AI agent sources, and leverage the deterministic filename format `adr.{session_id}.{timestamp}.json` for reliable session tracking and change detection.**

The Uber ADR repository provides a complete telemetry pipeline for capturing AI agent interactions. Understanding how `session_id` and timestamp metadata flow through the system is essential for building accurate analytics, detecting session updates, and maintaining data consistency across multiple AI tools.

## Core Data Model: AgentEvent and Session Metadata

Every captured interaction in ADR is stored as an **`AgentEvent`**. This dataclass, defined in [`Sensor/adr_sensor/schemas/agent_event_schema.py`](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/schemas/agent_event_schema.py), includes two critical fields for temporal analysis.

### session_id: Multi-Source Session Identification

The `session_id` field (line 49 in [`agent_event_schema.py`](https://github.com/uber/ADR/blob/main/agent_event_schema.py)) uniquely identifies a conversation across the entire pipeline. Each parser generates IDs with a **source-specific prefix**:

- `claude_parser` → `claude_{uuid}`
- `warp_parser` → `warp_{uuid}`
- `cursor_parser` → `cursor_{uuid}`

This prefix design enables immediate source attribution without additional metadata lookups.

### timestamp: UTC-Normalized Temporal Anchors

All timestamps are normalized to UTC to eliminate timezone drift. The implementation in [`Sensor/adr_sensor/utils/timestamp_utils.py`](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/utils/timestamp_utils.py) provides three essential utilities:

| Function | Purpose | Lines |
|----------|---------|-------|
| `normalize_timestamp` | Converts ISO strings, Unix seconds, or Unix milliseconds to UTC-aware `datetime` | 9–28 |
| `format_timestamp_for_filename` | Generates `YYYYMMDD_HHMMSS` strings for deterministic filenames | 53–64 |
| `parse_timestamp_from_filename` | Extracts timestamps from `adr.{session_id}.{timestamp}.json` files | 67–83 |

## Essential Utilities for Timestamp Analysis

### Normalizing Raw Timestamps

Use `normalize_timestamp` to handle heterogeneous timestamp formats from different AI agent APIs:

```python
from adr_sensor.utils.timestamp_utils import normalize_timestamp

# Handles ISO 8601, Unix seconds, and Unix milliseconds automatically

utc_dt = normalize_timestamp("2024-10-12T15:45:30Z")
utc_dt = normalize_timestamp(1728726807)      # Unix seconds

utc_dt = normalize_timestamp(1728726807000)   # Unix milliseconds

```

All outputs are timezone-aware UTC `datetime` objects suitable for reliable comparison and storage.

### Generating Deterministic Filenames

The `format_timestamp_for_filename` function creates filesystem-safe timestamps for session persistence:

```python
from adr_sensor.utils.timestamp_utils import format_timestamp_for_filename

filename_ts = format_timestamp_for_filename(event.timestamp)

# Returns: "20241012_154530"

```

### Parsing Timestamps from Existing Files

For pipeline idempotency, extract timestamps from previously persisted sessions:

```python
from adr_sensor.utils.timestamp_utils import parse_timestamp_from_filename

ts = parse_timestamp_from_filename("adr.claude_abc.20241012_154530.json")

# Returns: datetime(2024, 10, 12, 15, 45, 30, tzinfo=UTC)

```

## The AgentObserver Orchestration Layer

The **`AgentObserver`** class in [`Sensor/adr_sensor/observer.py`](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/observer.py) manages the complete lifecycle of session data. Understanding its internal methods is crucial for custom analysis pipelines.

### Session ID Sanitization

Before persistence, session IDs are cleaned via `_clean_filename` (lines 11–31):

```python
clean_id = observer._clean_filename(event.session_id)

# Removes illegal filesystem characters, limits to 200 characters

```

### Building Session Timelines

The `_get_existing_session_files` method (lines 82–108) scans the session directory and returns a mapping of session IDs to their stored timestamps:

```python
existing = observer._get_existing_session_files()

# Returns: {"claude_abc": {"timestamp": datetime(...), "path": Path(...)}}

```

### Detecting Changed Sessions

Compare incoming events against stored state using normalized, truncated timestamps:

```python
from adr_sensor.observer import AgentObserver
from adr_sensor.utils.timestamp_utils import normalize_timestamp

observer = AgentObserver()
new_events = observer.ingest_all()[0]
existing = observer._get_existing_session_files()

changed = []
for ev in new_events:
    clean_id = observer._clean_filename(ev.session_id)
    
    if clean_id not in existing:
        changed.append(ev)  # New session

        continue
    
    # Compare with second-level precision

    stored_ts = normalize_timestamp(existing[clean_id]["timestamp"]).replace(microsecond=0)
    ev_ts = normalize_timestamp(ev.timestamp).replace(microsecond=0)
    
    if ev_ts > stored_ts:
        changed.append(ev)  # Updated session

```

## Recommended Analysis Practices

### 1. Always Normalize Before Comparison

Different parsers emit timestamps in varying formats. `normalize_timestamp` ensures uniform UTC comparison:

```python

# Defensive timestamp handling

norm_ts = normalize_timestamp(raw_ts).replace(microsecond=0)

```

### 2. Extract Source from session_id Prefix

Split the `session_id` to group analytics by AI agent tool:

```python
source = event.session_id.split("_")[0]  # "claude", "warp", or "cursor"

```

### 3. Leverage Content Hashing for Subtle Changes

When timestamps match but content may differ, use `get_content_hash()` from [`agent_event_schema.py`](https://github.com/uber/ADR/blob/main/agent_event_schema.py):

```python
content_fingerprint = event.get_content_hash()  # SHA-256 of chat history

```

### 4. Batch Process with Existing File Mapping

For large-scale analysis, build the session index once, then stream only changed events:

```python
existing = observer._get_existing_session_files()  # Index

# Process new_events filtered against existing timestamps

```

## Practical Code Examples

### Example 1: Build Source-Timeline Index from Session Files

```python
import json
from pathlib import Path
from adr_sensor.utils.timestamp_utils import parse_timestamp_from_filename

def build_source_timeline(session_dir: Path) -> dict:
    """Create {source: [(timestamp, session_id)]} mapping."""
    timeline = {}
    
    for file_path in session_dir.glob("adr.*.json"):
        # Parse filename: adr.{session_id}.{timestamp}.json

        stem = file_path.stem  # "adr.claude_abc.20241012_154530"

        parts = stem.split(".", 1)[1].rsplit(".", 1)  # ["claude_abc", "20241012_154530"]

        
        if len(parts) != 2:
            continue
            
        raw_session_id, timestamp_str = parts
        ts = parse_timestamp_from_filename(file_path.name)
        
        if not ts:
            continue
            
        source = raw_session_id.split("_")[0]
        timeline.setdefault(source, []).append((ts, raw_session_id))
    
    # Sort chronologically per source

    for source in timeline:
        timeline[source].sort(key=lambda x: x[0])
    
    return timeline

# Usage

session_dir = Path.home() / ".cache" / "adr_sensor"
timeline = build_source_timeline(session_dir)

for src, entries in timeline.items():
    print(f"{src.upper()}: {len(entries)} sessions")
    for ts, sess in entries[:3]:
        print(f"  {ts.isoformat()}  {sess}")

```

### Example 2: Visualize Session Activity by Source

```python
import matplotlib.pyplot as plt

def plot_session_distribution(timeline: dict):
    sources = list(timeline.keys())
    counts = [len(entries) for entries in timeline.values()]
    
    plt.figure(figsize=(10, 6))
    plt.barh(sources, counts, color=["#F97316", "#3B82F6", "#10B981"])
    plt.xlabel("Session Count")
    plt.title("AI Agent Session Distribution")
    plt.tight_layout()
    plt.savefig("adr_session_distribution.png")

plot_session_distribution(timeline)

```

### Example 3: Detect Stale Sessions Beyond Threshold

```python
from datetime import datetime, timedelta

def find_stale_sessions(observer: AgentObserver, days: int = 7) -> list:
    """Return session IDs not updated within specified days."""
    cutoff = datetime.now(UTC) - timedelta(days=days)
    existing = observer._get_existing_session_files()
    
    stale = []
    for session_id, info in existing.items():
        if normalize_timestamp(info["timestamp"]) < cutoff:
            stale.append((session_id, info["timestamp"], info["path"]))
    
    return stale

```

## Key Source Files Reference

| File | Purpose | Critical Lines |
|------|---------|--------------|
| [`Sensor/adr_sensor/schemas/agent_event_schema.py`](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/schemas/agent_event_schema.py) | `AgentEvent` dataclass, `session_id` field, `get_content_hash()` | 49 |
| [`Sensor/adr_sensor/utils/timestamp_utils.py`](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/utils/timestamp_utils.py) | `normalize_timestamp()`, `format_timestamp_for_filename()`, `parse_timestamp_from_filename()` | 9–28, 53–64, 67–83 |
| [`Sensor/adr_sensor/observer.py`](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/observer.py) | `AgentObserver`, `_clean_filename()`, `_get_existing_session_files()` | 11–31, 23–48, 82–108 |
| [`Sensor/adr_sensor/parsers/claude_parser.py`](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/parsers/claude_parser.py) | `claude_` prefix generation | — |
| [`Sensor/adr_sensor/parsers/warp_parser.py`](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/parsers/warp_parser.py) | `warp_` prefix generation | — |
| [`Sensor/adr_sensor/parsers/cursor_parser.py`](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/parsers/cursor_parser.py) | `cursor_` prefix generation | — |

## Summary

- **Normalize all timestamps** to UTC using `normalize_timestamp()` before any comparison or storage operation
- **Clean session IDs** with `observer._clean_filename()` to prevent filesystem errors and ensure dictionary key consistency
- **Parse source tools** from the `session_id` prefix (`claude_`, `warp_`, `cursor_`) for grouped analytics
- **Use deterministic filenames** (`adr.{session_id}.{timestamp}.json`) for idempotent, reversible pipelines
- **Compare at second precision** by truncating microseconds to avoid false positives from sub-second jitter
- **Leverage content hashing** via `get_content_hash()` when timestamps alone cannot detect meaningful changes
- **Batch with existing file maps** using `_get_existing_session_files()` for efficient incremental processing

## Frequently Asked Questions

### How do I handle timezone inconsistencies from different AI agent APIs?

All timestamps must pass through `normalize_timestamp()` in [`Sensor/adr_sensor/utils/timestamp_utils.py`](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/utils/timestamp_utils.py). This utility automatically detects ISO 8601 strings, Unix seconds, and Unix milliseconds, converting each to a timezone-aware UTC `datetime`. No manual timezone conversion is required.

### Why does my session detection show false positives for unchanged sessions?

Sub-second precision in timestamps often causes spurious differences. Truncate microseconds before comparison: `normalize_timestamp(ts).replace(microsecond=0)`. This aligns with the second-level granularity used in ADR's filename format.

### Can I identify which AI tool generated a session without parsing the full event?

Yes. The `session_id` prefix encodes the source directly. Split on underscore: `event.session_id.split("_")[0]` returns `claude`, `warp`, `cursor`, or another parser-specific identifier. This design eliminates the need for additional metadata lookups.

### What is the maximum safe length for a session_id in filenames?

The `_clean_filename` method enforces a 200-character limit after removing illegal filesystem characters. Exceeding this truncates the ID, so avoid embedding large payloads in session identifiers.