How to Integrate ADR Sensor Output with SIEM and Security Pipelines
Stream structured JSON telemetry from the ADR Sensor to any SIEM—Splunk, Elastic Stack, Sentinel, or custom endpoints—by configuring the sensor's export APIs and pointing your log forwarder at the output directory.
The ADR Sensor, developed by Uber, captures and records AI-coding agent activities from tools like Claude, Cursor, Cline, Warp, and Codex. This telemetry provides security teams with critical visibility into AI-driven code generation. Integrating ADR Sensor output with your SIEM enables real-time monitoring, threat detection, and compliance auditing without modifying the sensor's core code.
ADR Sensor Export APIs for SIEM Integration
The sensor exposes two primary export methods in [adr_sensor/observer.py](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/observer.py). Both write structured JSON that any modern SIEM can ingest.
| Export API | Purpose | Location |
|---|---|---|
AgentObserver.save_to_file() |
Writes a single aggregate JSON (or JSON Lines) file with all captured events and optional system configuration | observer.py#L59 |
AgentObserver.save_sessions_to_individual_files() |
Writes one JSON file per agent session (adr.<session_id>.<timestamp>.json) for incremental processing |
observer.py#L23 |
Both methods accept an output directory parameter. Point this to a path monitored by your log forwarder for seamless pipeline integration.
Step-by-Step SIEM Integration Workflow
1. Configure the Sensor Output Directory
Initialize AgentObserver with a directory your forwarder already watches:
from pathlib import Path
from adr_sensor import AgentObserver
observer = AgentObserver(output_dir=Path("/var/log/adr_sensor"))
2. Capture and Export Telemetry
Ingest events from supported agents, then choose your export strategy:
events, configs = observer.ingest_all() # Pull logs from all agents
# Option A: Single aggregated file (batch processing)
observer.save_to_file(events, configs)
# Option B: Per-session files (real-time, incremental pipelines)
observer.save_sessions_to_individual_files(events)
3. Forward to Your SIEM Platform
Configure your log forwarder to monitor /var/log/adr_sensor:
- Elastic Stack: Configure Filebeat with
json.keys_under_root: trueto flatten nested JSON structures - Splunk: Point Universal Forwarder at the directory; set sourcetype to
JSON - Azure Sentinel: Use the Log Analytics agent with custom JSON parsing
- Custom HTTP endpoint: Deploy a lightweight poller (see Python example below)
4. Map ADR Schema Fields to SIEM Attributes
The AgentEvent schema ([schemas/agent_event_schema.py](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/schemas/agent_event_schema.py)) and SystemConfiguration ([schemas/system_config_schema.py](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/schemas/system_config_schema.py)) provide rich telemetry. Typical field mappings:
| ADR Sensor Field | SIEM Field Example |
|---|---|
event.timestamp |
@timestamp |
event.session_id |
sessionId |
event.agent_name |
sourceAgent |
chat_message.role |
userRole |
tool_usage.name |
toolName |
system_configuration.host_os |
hostOS |
system_configuration.python_version |
pythonVersion |
Implement mappings via Filebeat processors, Splunk field extractions, or a transformation layer.
5. Enrich and Correlate Events
The SystemConfiguration payload attaches host metadata to every session. Correlate ADR events with:
- Process creation logs
- Network flow data
- Auditd/systemd logs
- Cloud trail events
Custom HTTP Endpoint Integration Example
For security pipelines requiring direct API ingestion, poll the output directory and POST to your SIEM:
import json
import time
from pathlib import Path
import requests
WATCH_DIR = Path("/var/log/adr_sensor")
SIEM_ENDPOINT = "https://siem.example.com/api/events"
def post_event(event: dict, auth_token: str) -> None:
headers = {
"Authorization": f"Bearer {auth_token}",
"Content-Type": "application/json"
}
resp = requests.post(SIEM_ENDPOINT, headers=headers, json=event, timeout=10)
resp.raise_for_status()
def stream_new_files(auth_token: str) -> None:
seen = set()
while True:
for file_path in WATCH_DIR.glob("adr.*.json"):
if file_path in seen:
continue
seen.add(file_path)
with file_path.open() as f:
event = json.load(f)
post_event(event, auth_token)
time.sleep(5) # Replace with inotify/watchdog for production
if __name__ == "__main__":
import os
token = os.environ.get("SIEM_API_TOKEN") # Never hardcode secrets
stream_new_files(token)
Security note: Store authentication credentials in environment variables or a secrets manager. Never commit tokens to source control.
Key Source Files for Integration Development
| File | Purpose |
|---|---|
[adr_sensor/observer.py](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/observer.py) |
Core orchestrator; save_to_file() and save_sessions_to_individual_files() |
[adr_sensor/cli.py](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/cli.py) |
Command-line wrapper for ad-hoc runs |
[adr_sensor/schemas/agent_event_schema.py](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/schemas/agent_event_schema.py) |
AgentEvent JSON structure definition |
[adr_sensor/schemas/system_config_schema.py](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/schemas/system_config_schema.py) |
Host configuration schema |
[adr_sensor/utils/timestamp_utils.py](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/utils/timestamp_utils.py) |
Sortable timestamp generation for session filenames |
Summary
- ADR Sensor output is native JSON—no parsing layer required for SIEM integration
- Two export modes: single aggregated file (
save_to_file) or per-session files (save_sessions_to_individual_files) - Zero code changes needed: Configure the output directory to match your log forwarder's watch path
- Schema includes full context: Agent events plus system configuration enable rich correlation
- Production tip: Use
save_sessions_to_individual_filesfor real-time pipelines; aggregate for batch analytics
Frequently Asked Questions
What SIEM platforms are compatible with ADR Sensor output?
Any platform that ingests JSON logs. Verified patterns exist for Splunk (Universal Forwarder with JSON sourcetype), Elastic Stack (Filebeat with json.keys_under_root), Azure Sentinel (Log Analytics custom JSON), and generic HTTP endpoints. The sensor's standard JSON structure eliminates custom parser development.
Does ADR Sensor require code modifications for SIEM integration?
No. The sensor's export APIs accept an output directory parameter. Point this to a location your forwarder monitors—no changes to [observer.py](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/observer.py) or other core files are necessary. For custom pipelines, the provided Python example shows minimal wrapper code.
How do I handle real-time versus batch ingestion?
Use save_sessions_to_individual_files() for real-time streams—each new session generates a discrete file you can tail or watch with inotify. Use save_to_file() for scheduled batch jobs that aggregate all events into a single JSON document. Both approaches use identical schemas, so downstream parsers work unchanged.
What security considerations apply to ADR Sensor telemetry?
Treat exported JSON as sensitive security data—it contains complete chat histories and tool invocations. Secure the output directory with appropriate filesystem permissions, encrypt data in transit to your SIEM, and store API tokens in secrets managers. The sensor itself does not encrypt at rest; implement this at the storage or transport layer.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →