Best Practices for Analyzing AI Agent session_id and Timestamp Patterns in ADR

Always normalize timestamps to UTC using normalize_timestamp(), prefix session_id values to identify AI agent sources, and leverage the deterministic filename format adr.{session_id}.{timestamp}.json for reliable session tracking and change detection.

The Uber ADR repository provides a complete telemetry pipeline for capturing AI agent interactions. Understanding how session_id and timestamp metadata flow through the system is essential for building accurate analytics, detecting session updates, and maintaining data consistency across multiple AI tools.

Core Data Model: AgentEvent and Session Metadata

Every captured interaction in ADR is stored as an AgentEvent. This dataclass, defined in Sensor/adr_sensor/schemas/agent_event_schema.py, includes two critical fields for temporal analysis.

session_id: Multi-Source Session Identification

The session_id field (line 49 in agent_event_schema.py) uniquely identifies a conversation across the entire pipeline. Each parser generates IDs with a source-specific prefix:

  • claude_parser → claude_{uuid}
  • warp_parser → warp_{uuid}
  • cursor_parser → cursor_{uuid}

This prefix design enables immediate source attribution without additional metadata lookups.

timestamp: UTC-Normalized Temporal Anchors

All timestamps are normalized to UTC to eliminate timezone drift. The implementation in Sensor/adr_sensor/utils/timestamp_utils.py provides three essential utilities:

Function Purpose Lines
normalize_timestamp Converts ISO strings, Unix seconds, or Unix milliseconds to UTC-aware datetime 9–28
format_timestamp_for_filename Generates YYYYMMDD_HHMMSS strings for deterministic filenames 53–64
parse_timestamp_from_filename Extracts timestamps from adr.{session_id}.{timestamp}.json files 67–83

Essential Utilities for Timestamp Analysis

Normalizing Raw Timestamps

Use normalize_timestamp to handle heterogeneous timestamp formats from different AI agent APIs:

from adr_sensor.utils.timestamp_utils import normalize_timestamp

# Handles ISO 8601, Unix seconds, and Unix milliseconds automatically

utc_dt = normalize_timestamp("2024-10-12T15:45:30Z")
utc_dt = normalize_timestamp(1728726807)      # Unix seconds

utc_dt = normalize_timestamp(1728726807000)   # Unix milliseconds

All outputs are timezone-aware UTC datetime objects suitable for reliable comparison and storage.

Generating Deterministic Filenames

The format_timestamp_for_filename function creates filesystem-safe timestamps for session persistence:

from adr_sensor.utils.timestamp_utils import format_timestamp_for_filename

filename_ts = format_timestamp_for_filename(event.timestamp)

# Returns: "20241012_154530"

Parsing Timestamps from Existing Files

For pipeline idempotency, extract timestamps from previously persisted sessions:

from adr_sensor.utils.timestamp_utils import parse_timestamp_from_filename

ts = parse_timestamp_from_filename("adr.claude_abc.20241012_154530.json")

# Returns: datetime(2024, 10, 12, 15, 45, 30, tzinfo=UTC)

The AgentObserver Orchestration Layer

The AgentObserver class in Sensor/adr_sensor/observer.py manages the complete lifecycle of session data. Understanding its internal methods is crucial for custom analysis pipelines.

Session ID Sanitization

Before persistence, session IDs are cleaned via _clean_filename (lines 11–31):

clean_id = observer._clean_filename(event.session_id)

# Removes illegal filesystem characters, limits to 200 characters

Building Session Timelines

The _get_existing_session_files method (lines 82–108) scans the session directory and returns a mapping of session IDs to their stored timestamps:

existing = observer._get_existing_session_files()

# Returns: {"claude_abc": {"timestamp": datetime(...), "path": Path(...)}}

Detecting Changed Sessions

Compare incoming events against stored state using normalized, truncated timestamps:

from adr_sensor.observer import AgentObserver
from adr_sensor.utils.timestamp_utils import normalize_timestamp

observer = AgentObserver()
new_events = observer.ingest_all()[0]
existing = observer._get_existing_session_files()

changed = []
for ev in new_events:
    clean_id = observer._clean_filename(ev.session_id)
    
    if clean_id not in existing:
        changed.append(ev)  # New session

        continue
    
    # Compare with second-level precision

    stored_ts = normalize_timestamp(existing[clean_id]["timestamp"]).replace(microsecond=0)
    ev_ts = normalize_timestamp(ev.timestamp).replace(microsecond=0)
    
    if ev_ts > stored_ts:
        changed.append(ev)  # Updated session

1. Always Normalize Before Comparison

Different parsers emit timestamps in varying formats. normalize_timestamp ensures uniform UTC comparison:


# Defensive timestamp handling

norm_ts = normalize_timestamp(raw_ts).replace(microsecond=0)

2. Extract Source from session_id Prefix

Split the session_id to group analytics by AI agent tool:

source = event.session_id.split("_")[0]  # "claude", "warp", or "cursor"

3. Leverage Content Hashing for Subtle Changes

When timestamps match but content may differ, use get_content_hash() from agent_event_schema.py:

content_fingerprint = event.get_content_hash()  # SHA-256 of chat history

4. Batch Process with Existing File Mapping

For large-scale analysis, build the session index once, then stream only changed events:

existing = observer._get_existing_session_files()  # Index

# Process new_events filtered against existing timestamps

Practical Code Examples

Example 1: Build Source-Timeline Index from Session Files

import json
from pathlib import Path
from adr_sensor.utils.timestamp_utils import parse_timestamp_from_filename

def build_source_timeline(session_dir: Path) -> dict:
    """Create {source: [(timestamp, session_id)]} mapping."""
    timeline = {}
    
    for file_path in session_dir.glob("adr.*.json"):
        # Parse filename: adr.{session_id}.{timestamp}.json

        stem = file_path.stem  # "adr.claude_abc.20241012_154530"

        parts = stem.split(".", 1)[1].rsplit(".", 1)  # ["claude_abc", "20241012_154530"]

        
        if len(parts) != 2:
            continue
            
        raw_session_id, timestamp_str = parts
        ts = parse_timestamp_from_filename(file_path.name)
        
        if not ts:
            continue
            
        source = raw_session_id.split("_")[0]
        timeline.setdefault(source, []).append((ts, raw_session_id))
    
    # Sort chronologically per source

    for source in timeline:
        timeline[source].sort(key=lambda x: x[0])
    
    return timeline

# Usage

session_dir = Path.home() / ".cache" / "adr_sensor"
timeline = build_source_timeline(session_dir)

for src, entries in timeline.items():
    print(f"{src.upper()}: {len(entries)} sessions")
    for ts, sess in entries[:3]:
        print(f"  {ts.isoformat()}  {sess}")

Example 2: Visualize Session Activity by Source

import matplotlib.pyplot as plt

def plot_session_distribution(timeline: dict):
    sources = list(timeline.keys())
    counts = [len(entries) for entries in timeline.values()]
    
    plt.figure(figsize=(10, 6))
    plt.barh(sources, counts, color=["#F97316", "#3B82F6", "#10B981"])
    plt.xlabel("Session Count")
    plt.title("AI Agent Session Distribution")
    plt.tight_layout()
    plt.savefig("adr_session_distribution.png")

plot_session_distribution(timeline)

Example 3: Detect Stale Sessions Beyond Threshold

from datetime import datetime, timedelta

def find_stale_sessions(observer: AgentObserver, days: int = 7) -> list:
    """Return session IDs not updated within specified days."""
    cutoff = datetime.now(UTC) - timedelta(days=days)
    existing = observer._get_existing_session_files()
    
    stale = []
    for session_id, info in existing.items():
        if normalize_timestamp(info["timestamp"]) < cutoff:
            stale.append((session_id, info["timestamp"], info["path"]))
    
    return stale

Key Source Files Reference

File Purpose Critical Lines
Sensor/adr_sensor/schemas/agent_event_schema.py AgentEvent dataclass, session_id field, get_content_hash() 49
Sensor/adr_sensor/utils/timestamp_utils.py normalize_timestamp(), format_timestamp_for_filename(), parse_timestamp_from_filename() 9–28, 53–64, 67–83
Sensor/adr_sensor/observer.py AgentObserver, _clean_filename(), _get_existing_session_files() 11–31, 23–48, 82–108
Sensor/adr_sensor/parsers/claude_parser.py claude_ prefix generation —
Sensor/adr_sensor/parsers/warp_parser.py warp_ prefix generation —
Sensor/adr_sensor/parsers/cursor_parser.py cursor_ prefix generation —

Summary

  • Normalize all timestamps to UTC using normalize_timestamp() before any comparison or storage operation
  • Clean session IDs with observer._clean_filename() to prevent filesystem errors and ensure dictionary key consistency
  • Parse source tools from the session_id prefix (claude_, warp_, cursor_) for grouped analytics
  • Use deterministic filenames (adr.{session_id}.{timestamp}.json) for idempotent, reversible pipelines
  • Compare at second precision by truncating microseconds to avoid false positives from sub-second jitter
  • Leverage content hashing via get_content_hash() when timestamps alone cannot detect meaningful changes
  • Batch with existing file maps using _get_existing_session_files() for efficient incremental processing

Frequently Asked Questions

How do I handle timezone inconsistencies from different AI agent APIs?

All timestamps must pass through normalize_timestamp() in Sensor/adr_sensor/utils/timestamp_utils.py. This utility automatically detects ISO 8601 strings, Unix seconds, and Unix milliseconds, converting each to a timezone-aware UTC datetime. No manual timezone conversion is required.

Why does my session detection show false positives for unchanged sessions?

Sub-second precision in timestamps often causes spurious differences. Truncate microseconds before comparison: normalize_timestamp(ts).replace(microsecond=0). This aligns with the second-level granularity used in ADR's filename format.

Can I identify which AI tool generated a session without parsing the full event?

Yes. The session_id prefix encodes the source directly. Split on underscore: event.session_id.split("_")[0] returns claude, warp, cursor, or another parser-specific identifier. This design eliminates the need for additional metadata lookups.

What is the maximum safe length for a session_id in filenames?

The _clean_filename method enforces a 200-character limit after removing illegal filesystem characters. Exceeding this truncates the ID, so avoid embedding large payloads in session identifiers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →