Best Practices for Analyzing AI Agent session_id and Timestamp Patterns in ADR
Always normalize timestamps to UTC using normalize_timestamp(), prefix session_id values to identify AI agent sources, and leverage the deterministic filename format adr.{session_id}.{timestamp}.json for reliable session tracking and change detection.
The Uber ADR repository provides a complete telemetry pipeline for capturing AI agent interactions. Understanding how session_id and timestamp metadata flow through the system is essential for building accurate analytics, detecting session updates, and maintaining data consistency across multiple AI tools.
Core Data Model: AgentEvent and Session Metadata
Every captured interaction in ADR is stored as an AgentEvent. This dataclass, defined in Sensor/adr_sensor/schemas/agent_event_schema.py, includes two critical fields for temporal analysis.
session_id: Multi-Source Session Identification
The session_id field (line 49 in agent_event_schema.py) uniquely identifies a conversation across the entire pipeline. Each parser generates IDs with a source-specific prefix:
claude_parser→claude_{uuid}warp_parser→warp_{uuid}cursor_parser→cursor_{uuid}
This prefix design enables immediate source attribution without additional metadata lookups.
timestamp: UTC-Normalized Temporal Anchors
All timestamps are normalized to UTC to eliminate timezone drift. The implementation in Sensor/adr_sensor/utils/timestamp_utils.py provides three essential utilities:
| Function | Purpose | Lines |
|---|---|---|
normalize_timestamp |
Converts ISO strings, Unix seconds, or Unix milliseconds to UTC-aware datetime |
9–28 |
format_timestamp_for_filename |
Generates YYYYMMDD_HHMMSS strings for deterministic filenames |
53–64 |
parse_timestamp_from_filename |
Extracts timestamps from adr.{session_id}.{timestamp}.json files |
67–83 |
Essential Utilities for Timestamp Analysis
Normalizing Raw Timestamps
Use normalize_timestamp to handle heterogeneous timestamp formats from different AI agent APIs:
from adr_sensor.utils.timestamp_utils import normalize_timestamp
# Handles ISO 8601, Unix seconds, and Unix milliseconds automatically
utc_dt = normalize_timestamp("2024-10-12T15:45:30Z")
utc_dt = normalize_timestamp(1728726807) # Unix seconds
utc_dt = normalize_timestamp(1728726807000) # Unix milliseconds
All outputs are timezone-aware UTC datetime objects suitable for reliable comparison and storage.
Generating Deterministic Filenames
The format_timestamp_for_filename function creates filesystem-safe timestamps for session persistence:
from adr_sensor.utils.timestamp_utils import format_timestamp_for_filename
filename_ts = format_timestamp_for_filename(event.timestamp)
# Returns: "20241012_154530"
Parsing Timestamps from Existing Files
For pipeline idempotency, extract timestamps from previously persisted sessions:
from adr_sensor.utils.timestamp_utils import parse_timestamp_from_filename
ts = parse_timestamp_from_filename("adr.claude_abc.20241012_154530.json")
# Returns: datetime(2024, 10, 12, 15, 45, 30, tzinfo=UTC)
The AgentObserver Orchestration Layer
The AgentObserver class in Sensor/adr_sensor/observer.py manages the complete lifecycle of session data. Understanding its internal methods is crucial for custom analysis pipelines.
Session ID Sanitization
Before persistence, session IDs are cleaned via _clean_filename (lines 11–31):
clean_id = observer._clean_filename(event.session_id)
# Removes illegal filesystem characters, limits to 200 characters
Building Session Timelines
The _get_existing_session_files method (lines 82–108) scans the session directory and returns a mapping of session IDs to their stored timestamps:
existing = observer._get_existing_session_files()
# Returns: {"claude_abc": {"timestamp": datetime(...), "path": Path(...)}}
Detecting Changed Sessions
Compare incoming events against stored state using normalized, truncated timestamps:
from adr_sensor.observer import AgentObserver
from adr_sensor.utils.timestamp_utils import normalize_timestamp
observer = AgentObserver()
new_events = observer.ingest_all()[0]
existing = observer._get_existing_session_files()
changed = []
for ev in new_events:
clean_id = observer._clean_filename(ev.session_id)
if clean_id not in existing:
changed.append(ev) # New session
continue
# Compare with second-level precision
stored_ts = normalize_timestamp(existing[clean_id]["timestamp"]).replace(microsecond=0)
ev_ts = normalize_timestamp(ev.timestamp).replace(microsecond=0)
if ev_ts > stored_ts:
changed.append(ev) # Updated session
Recommended Analysis Practices
1. Always Normalize Before Comparison
Different parsers emit timestamps in varying formats. normalize_timestamp ensures uniform UTC comparison:
# Defensive timestamp handling
norm_ts = normalize_timestamp(raw_ts).replace(microsecond=0)
2. Extract Source from session_id Prefix
Split the session_id to group analytics by AI agent tool:
source = event.session_id.split("_")[0] # "claude", "warp", or "cursor"
3. Leverage Content Hashing for Subtle Changes
When timestamps match but content may differ, use get_content_hash() from agent_event_schema.py:
content_fingerprint = event.get_content_hash() # SHA-256 of chat history
4. Batch Process with Existing File Mapping
For large-scale analysis, build the session index once, then stream only changed events:
existing = observer._get_existing_session_files() # Index
# Process new_events filtered against existing timestamps
Practical Code Examples
Example 1: Build Source-Timeline Index from Session Files
import json
from pathlib import Path
from adr_sensor.utils.timestamp_utils import parse_timestamp_from_filename
def build_source_timeline(session_dir: Path) -> dict:
"""Create {source: [(timestamp, session_id)]} mapping."""
timeline = {}
for file_path in session_dir.glob("adr.*.json"):
# Parse filename: adr.{session_id}.{timestamp}.json
stem = file_path.stem # "adr.claude_abc.20241012_154530"
parts = stem.split(".", 1)[1].rsplit(".", 1) # ["claude_abc", "20241012_154530"]
if len(parts) != 2:
continue
raw_session_id, timestamp_str = parts
ts = parse_timestamp_from_filename(file_path.name)
if not ts:
continue
source = raw_session_id.split("_")[0]
timeline.setdefault(source, []).append((ts, raw_session_id))
# Sort chronologically per source
for source in timeline:
timeline[source].sort(key=lambda x: x[0])
return timeline
# Usage
session_dir = Path.home() / ".cache" / "adr_sensor"
timeline = build_source_timeline(session_dir)
for src, entries in timeline.items():
print(f"{src.upper()}: {len(entries)} sessions")
for ts, sess in entries[:3]:
print(f" {ts.isoformat()} {sess}")
Example 2: Visualize Session Activity by Source
import matplotlib.pyplot as plt
def plot_session_distribution(timeline: dict):
sources = list(timeline.keys())
counts = [len(entries) for entries in timeline.values()]
plt.figure(figsize=(10, 6))
plt.barh(sources, counts, color=["#F97316", "#3B82F6", "#10B981"])
plt.xlabel("Session Count")
plt.title("AI Agent Session Distribution")
plt.tight_layout()
plt.savefig("adr_session_distribution.png")
plot_session_distribution(timeline)
Example 3: Detect Stale Sessions Beyond Threshold
from datetime import datetime, timedelta
def find_stale_sessions(observer: AgentObserver, days: int = 7) -> list:
"""Return session IDs not updated within specified days."""
cutoff = datetime.now(UTC) - timedelta(days=days)
existing = observer._get_existing_session_files()
stale = []
for session_id, info in existing.items():
if normalize_timestamp(info["timestamp"]) < cutoff:
stale.append((session_id, info["timestamp"], info["path"]))
return stale
Key Source Files Reference
| File | Purpose | Critical Lines |
|---|---|---|
Sensor/adr_sensor/schemas/agent_event_schema.py |
AgentEvent dataclass, session_id field, get_content_hash() |
49 |
Sensor/adr_sensor/utils/timestamp_utils.py |
normalize_timestamp(), format_timestamp_for_filename(), parse_timestamp_from_filename() |
9–28, 53–64, 67–83 |
Sensor/adr_sensor/observer.py |
AgentObserver, _clean_filename(), _get_existing_session_files() |
11–31, 23–48, 82–108 |
Sensor/adr_sensor/parsers/claude_parser.py |
claude_ prefix generation |
— |
Sensor/adr_sensor/parsers/warp_parser.py |
warp_ prefix generation |
— |
Sensor/adr_sensor/parsers/cursor_parser.py |
cursor_ prefix generation |
— |
Summary
- Normalize all timestamps to UTC using
normalize_timestamp()before any comparison or storage operation - Clean session IDs with
observer._clean_filename()to prevent filesystem errors and ensure dictionary key consistency - Parse source tools from the
session_idprefix (claude_,warp_,cursor_) for grouped analytics - Use deterministic filenames (
adr.{session_id}.{timestamp}.json) for idempotent, reversible pipelines - Compare at second precision by truncating microseconds to avoid false positives from sub-second jitter
- Leverage content hashing via
get_content_hash()when timestamps alone cannot detect meaningful changes - Batch with existing file maps using
_get_existing_session_files()for efficient incremental processing
Frequently Asked Questions
How do I handle timezone inconsistencies from different AI agent APIs?
All timestamps must pass through normalize_timestamp() in Sensor/adr_sensor/utils/timestamp_utils.py. This utility automatically detects ISO 8601 strings, Unix seconds, and Unix milliseconds, converting each to a timezone-aware UTC datetime. No manual timezone conversion is required.
Why does my session detection show false positives for unchanged sessions?
Sub-second precision in timestamps often causes spurious differences. Truncate microseconds before comparison: normalize_timestamp(ts).replace(microsecond=0). This aligns with the second-level granularity used in ADR's filename format.
Can I identify which AI tool generated a session without parsing the full event?
Yes. The session_id prefix encodes the source directly. Split on underscore: event.session_id.split("_")[0] returns claude, warp, cursor, or another parser-specific identifier. This design eliminates the need for additional metadata lookups.
What is the maximum safe length for a session_id in filenames?
The _clean_filename method enforces a 200-character limit after removing illegal filesystem characters. Exceeding this truncates the ID, so avoid embedding large payloads in session identifiers.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →