Understanding the BaseParser Abstract Class for Developing Custom ADR Parsers
The BaseParser abstract class defines a minimal interface—requiring only the implementation of parse_all() returning List[AgentEvent]—that enables the Uber ADR (AI-Driven Reconnaissance) sensor to ingest and normalize logs from any AI agent without modifying downstream analysis components.
The Uber ADR repository provides a pluggable architecture for monitoring AI-agent activity across your infrastructure. At the heart of this extensibility sits the BaseParser class in the Sensor component, which serves as the single mandated extension point for adding support for new agent log formats. By inheriting from this abstract base, developers can integrate custom telemetry sources while guaranteeing that the Detection pipeline receives uniformly structured data.
BaseParser Location and Core Contract
The BaseParser class resides at Sensor/adr_sensor/parsers/base_parser.py. It establishes a strict contract that every concrete parser must satisfy: the implementation of a single abstract method.
The parse_all Method
Concrete subclasses must override the parse_all method with the following signature:
def parse_all(self) -> List[AgentEvent]:
...
This method is responsible for locating agent-specific log files, extracting relevant telemetry, transforming that data into the unified schema, and returning a list of AgentEvent objects defined in Sensor/adr_sensor/schemas/agent_event_schema.py. The Sensor Core iterates over all discovered BaseParser subclasses and invokes this method to collect events before passing them to the Detection component.
Design Rationale and Architecture
The BaseParser abstraction provides three architectural benefits that simplify maintenance and encourage contribution.
Decoupling Agent-Specific Logic
Concrete parsers handle only the idiosyncrasies of their specific log formats—whether JSONL, XML, or proprietary binary structures. Once transformed into AgentEvent objects, the data conforms to a standard schema, allowing the Detection pipeline in Detection/main_detector.py to consume telemetry without knowing the origin agent.
Extensibility Without Pipeline Changes
Adding support for a new AI agent requires creating a single file in Sensor/adr_sensor/parsers/ that subclasses BaseParser. The existing sensor pipeline automatically discovers and invokes the new parser through package imports, requiring zero modifications to core ingestion logic or visualization tools.
Isolated Testability
Because each parser operates independently and returns plain Python objects, developers can unit-test parsers in isolation by feeding synthetic log files and asserting on the produced AgentEvent list. This eliminates the need to spin up the entire ADR stack during development.
How BaseParser Integrates with the ADR System
The sensor discovers parsers through import statements within the parsers package. The workflow proceeds as follows:
- Discovery: The Sensor Core imports all modules in
Sensor/adr_sensor/parsers/, triggering registration of subclasses. - Execution: Each parser's
parse_allmethod enumerates relevant files (e.g.,~/.claude/projectsfor Claude logs), optionally filtering by age or pattern. - Normalization: Raw log entries are converted into
AgentEventobjects containingChatMessageandToolUsagesub-objects. - Consumption: The unified events flow into
Detection/main_detector.pyfor security analysis or persistence.
Implementing a Custom Parser
Follow these steps to create a valid extension that the sensor will automatically pick up.
Step 1: Create the Parser Module
Create a new Python file in Sensor/adr_sensor/parsers/, such as myagent_parser.py.
Step 2: Implement the Abstract Method
Subclass BaseParser and provide a concrete parse_all implementation that returns List[AgentEvent]:
# Sensor/adr_sensor/parsers/myagent_parser.py
from typing import List
from pathlib import Path
from ..schemas.agent_event_schema import AgentEvent, ChatMessage, ToolUsage
from .base_parser import BaseParser
class MyAgentParser(BaseParser):
"""Parser for MyAgent logs stored under ~/.myagent/logs."""
def __init__(self, log_dir: Path | None = None):
self.base_path = log_dir or Path.home() / ".myagent" / "logs"
def parse_all(self) -> List[AgentEvent]:
events: List[AgentEvent] = []
for log_file in self.base_path.rglob("*.jsonl"):
events.extend(self._parse_file(log_file))
return events
def _parse_file(self, file_path: Path) -> List[AgentEvent]:
# Extract timestamps, messages, and tool usage from native format
# Build AgentEvent objects using the shared schema
return []
Step 3: Register the Parser
Ensure the sensor imports your module by adding an import statement to Sensor/adr_sensor/parsers/__init__.py or by structuring the package so that the Sensor Core's discovery mechanism loads the file.
Error Handling and Performance Characteristics
Production parsers in the Uber ADR repository follow strict efficiency patterns that custom implementations should replicate.
Memory Efficiency: Parsers read log files line-by-line rather than loading entire JSON documents into memory. After processing, raw strings are discarded immediately to keep the heap footprint minimal.
Argument Truncation: Large tool arguments are truncated using utils/string_utils.truncate_middle before being stored in ToolUsage objects, preventing memory bloat from oversized LLM prompts or file contents.
Graceful Degradation: All exceptions during file reading or parsing are captured and logged without aborting the entire parse_all operation. This ensures that corrupted log files or permission errors in one agent directory do not block ingestion from other agents.
Stale Log Filtering: Concrete implementations like the ClaudeParser skip files older than a configurable threshold (default 14 days) to avoid processing obsolete telemetry, as implemented in Sensor/adr_sensor/parsers/claude_parser.py.
Complete Implementation Examples
Minimal Skeleton for a New Parser
# Sensor/adr_sensor/parsers/custom_parser.py
from typing import List
from pathlib import Path
from ..schemas.agent_event_schema import AgentEvent
from .base_parser import BaseParser
class CustomParser(BaseParser):
def __init__(self):
self.base_path = Path.home() / ".custom_agent" / "logs"
def parse_all(self) -> List[AgentEvent]:
events = []
for log_file in self.base_path.glob("*.log"):
# Transformation logic here
events.append(AgentEvent(...))
return events
Reference Implementation: ClaudeParser
The ClaudeParser at Sensor/adr_sensor/parsers/claude_parser.py demonstrates the canonical implementation pattern:
class ClaudeParser(BaseParser):
"""Parser for Claude Code JSONL log files."""
def __init__(self, max_age_days: int = 14):
self.base_path = Path.home() / ".claude/projects"
self.max_age_days = max_age_days
def parse_all(self) -> List[AgentEvent]:
# Locate recent JSONL files, filter by age via
# utils/timestamp_utils.normalize_timestamp, truncate
# oversized arguments, and construct AgentEvent objects.
...
This implementation streams JSONL files, normalizes timestamps, applies the truncation utility from Sensor/adr_sensor/utils/string_utils.py, and assembles the final event list.
Integrating Custom Parsers into the Detection Pipeline
Once implemented, instantiate your parser and pass the events directly to the detector:
from Sensor.adr_sensor.parsers.myagent_parser import MyAgentParser
from Detection.main_detector import MainDetector
parser = MyAgentParser()
events = parser.parse_all()
detector = MainDetector()
detector.process_events(events)
Because MyAgentParser adheres to the BaseParser contract, MainDetector treats its output identically to built-in parsers for Cursor, Warp, or Claude logs.
Summary
- Location:
BaseParserlives inSensor/adr_sensor/parsers/base_parser.pyand defines the mandatoryparse_all()interface. - Contract: Subclasses return
List[AgentEvent]using the shared schema fromSensor/adr_sensor/schemas/agent_event_schema.py. - Benefits: The abstraction decouples log-format specifics from analysis logic, enabling plug-and-play extensibility.
- Performance: Implementations should stream files line-by-line, truncate large arguments, and handle errors gracefully without crashing the sensor.
- Registration: Parsers are auto-discovered when imported into the
parserspackage namespace.
Frequently Asked Questions
What methods must I implement when subclassing BaseParser?
You must implement only the parse_all(self) -> List[AgentEvent] method. The constructor is flexible and may accept configuration parameters (such as log_dir or max_age_days), but the return type must strictly be a list of AgentEvent objects as defined in the shared schema.
How does the ADR sensor discover new parser implementations?
The Sensor Core discovers parsers by importing modules within the Sensor/adr_sensor/parsers/ package. When you create a new file containing a BaseParser subclass, ensure it is imported in __init__.py or loaded by the package's import chain. Once loaded, the class becomes available to the ingestion loop automatically.
What performance considerations should I keep in mind when implementing parse_all?
Follow the memory-efficient patterns used in ClaudeParser and CursorParser: read files line-by-line rather than loading them entirely, discard raw data after transforming it into AgentEvent objects, and truncate large strings using utils/string_utils.truncate_middle. Additionally, filter out stale logs based on age to prevent processingobsolete files.
How do I ensure my custom parser output is compatible with the Detection pipeline?
Import AgentEvent, ChatMessage, and ToolUsage from Sensor/adr_sensor/schemas/agent_event_schema.py and populate these dataclasses according to the schema definitions. As long as parse_all() returns a List[AgentEvent], the MainDetector in Detection/main_detector.py will consume the output without requiring format-specific knowledge.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →