Understanding the `has_meaningful_content()` Filtering Logic in Uber ADR
The has_meaningful_content() method in Uber's ADR (AI Detection and Response) system filters AI-agent telemetry by requiring either multi-turn conversations, tool usage, or substantive text content while rejecting empty events, very short messages, and known noise patterns like authentication errors.
In the Uber/ADR repository, this method serves as the primary gatekeeper for AI-agent observability data. Every parsed log entry from agents like Cursor, Claude, Warp, and Codex must pass this filter before downstream security analytics and benchmarking pipelines process it. The following sections break down exactly how this filtering works, where it applies across the codebase, and how to interpret its behavior when building on or extending the ADR sensor.
How has_meaningful_content() Works
The filtering logic resides in Sensor/adr_sensor/schemas/agent_event_schema.py at line 118, implemented as a method on the AgentEvent dataclass. The method evaluates five sequential criteria to determine whether an event merits retention.
Step 1: Reject Empty Chat History
if not self.chat_history:
return False
Events without any messages contain no analyzable context and are immediately discarded.
Step 2: Accept Multi-Turn Conversations
if len(self.chat_history) >= 2:
return True
Any interaction with two or more turns qualifies as meaningful automatically—there is demonstrable back-and-forth between user and agent.
Step 3: Accept Single-Turn Events with Tool Usage
if msg.tools:
return True
Even single messages are retained if the assistant invoked a tool. Tool usage (e.g., MCP tool calls, code generation functions) signals actionable agent behavior regardless of conversational depth.
Step 4: Filter Minimal Textual Content
content = msg.content.strip()
if len(content) <= 5:
return False
Strings of five characters or fewer—such as "OK", "…", or "yes"—provide no analytical value and are rejected.
Step 5: Filter Known Noise Patterns
NOISE_PATTERNS = (
"invalid api key",
"please run /login",
"authentication failed",
"unauthorized",
"session expired",
"rate limit",
"api error",
"warmup",
)
if any(pattern in content.lower() for pattern in NOISE_PATTERNS):
return False
System-generated errors and warm-up messages that do not reflect genuine user-driven activity are excluded. This list is case-insensitive for robust matching.
If none of these conditions return False, the method ultimately returns True, admitting the event for downstream processing.
Where the Filter Is Applied
Multiple parsers across the ADR sensor invoke has_meaningful_content() to prune raw telemetry:
warp_parser.py— Filters Warp agent logscursor_parser.py— Filters Cursor agent logscodex_parser.py— Filters Codex agent logsclaude_parser.py— Filters Claude agent logsclaude_desktop_parser.py— Filters Claude Desktop logscline_parser.py— Filters Cline agent logsobserver.py— Performs final aggregation filtering
In observer.py, the typical pattern aggregates entries from all parsers then applies the filter:
filtered = [e for e in entries if e.has_meaningful_content()]
This centralized pruning ensures only relevant telemetry persists to storage and analytics.
Code Examples
Multi-Turn Conversation (Retained)
from datetime import datetime
from adr_sensor.schemas.agent_event_schema import AgentEvent, ChatMessage
event = AgentEvent(
timestamp=datetime.utcnow(),
source="claude",
session_id="sess_01",
chat_history=[
ChatMessage(role="user", content="How do I sort a list?"),
ChatMessage(role="assistant", content="You can use the `sorted()` function.", tools=[]),
],
)
assert event.has_meaningful_content() # → True
Single-Turn with Tool Usage (Retained)
event = AgentEvent(
timestamp=datetime.utcnow(),
source="cursor",
session_id="sess_02",
chat_history=[
ChatMessage(
role="assistant",
content="",
tools=[ToolUsage(tool_name="search", tool_type="mcp_tool", arguments={"query": "grep"})],
),
],
)
assert event.has_meaningful_content() # → True
Short Error Message (Filtered)
event = AgentEvent(
timestamp=datetime.utcnow(),
source="codex",
session_id="sess_03",
chat_history=[ChatMessage(role="assistant", content="OK")],
)
assert not event.has_meaningful_content() # → False
Known Noise Pattern (Filtered)
event = AgentEvent(
timestamp=datetime.utcnow(),
source="warp",
session_id="sess_04",
chat_history=[ChatMessage(role="assistant", content="Rate limit exceeded. Please try later.")],
)
assert not event.has_meaningging_content() # → False
Key Source Files
| File | Purpose |
|---|---|
Sensor/adr_sensor/schemas/agent_event_schema.py |
Defines AgentEvent dataclass and has_meaningful_content() method |
Sensor/adr_sensor/parsers/warp_parser.py |
Applies filter to Warp agent logs |
Sensor/adr_sensor/parsers/cursor_parser.py |
Applies filter to Cursor agent logs |
Sensor/adr_sensor/parsers/codex_parser.py |
Applies filter to Codex agent logs |
Sensor/adr_sensor/parsers/claude_parser.py |
Applies filter to Claude agent logs |
Sensor/adr_sensor/parsers/claude_desktop_parser.py |
Applies filter to Claude Desktop logs |
Sensor/adr_sensor/parsers/cline_parser.py |
Applies filter to Cline agent logs |
Sensor/adr_sensor/observer.py |
Aggregates and final-filters all parser outputs |
Architectural Impact
Data hygiene — The schema guarantee that every stored AgentEvent contributes actionable insight eliminates storage of empty or irrelevant records.
Performance — Downstream components including benchmarking suites, security guardrails, and export pipelines operate on leaner datasets, reducing CPU and I/O overhead.
Security focus — Noise-free logs improve signal-to-noise ratio for threat-intelligence modules consuming these events, enabling more reliable anomaly detection.
Summary
has_meaningful_content()is defined inSensor/adr_sensor/schemas/agent_event_schema.pyas anAgentEventmethod- Multi-turn conversations (≥2 messages) are automatically retained
- Single-turn events are retained only with tool usage or substantial text (>5 characters)
- Empty events, very short messages, and known noise patterns (authentication errors, rate limits, warmups) are filtered
- The filter is invoked by all major parsers and the central
observer.pyaggregator
Frequently Asked Questions
What counts as "meaningful content" in Uber ADR?
According to the Uber/ADR source code, meaningful content includes: (1) any conversation with two or more turns, (2) single-turn events where the assistant used a tool, or (3) single-turn events with more than five characters of text that do not match known noise patterns. Everything else is discarded.
How do I add new noise patterns to the filter?
Modify the NOISE_PATTERNS tuple in Sensor/adr_sensor/schemas/agent_event_schema.py. Add lowercase strings to match against; the comparison is case-insensitive. Rebuild and redeploy the sensor for changes to take effect across all parsers.
Why does the filter accept empty content if tools are present?
Tool invocations indicate the AI agent performed an actionable operation—searching code, generating files, calling APIs—even when the conversational content is minimal. This captures functionally significant behavior that security and audit workflows need to track.
Where does filtering happen in the ADR pipeline?
Each parser (warp_parser.py, cursor_parser.py, etc.) can apply has_meaningful_content(), and the observer.py module performs a final filtering pass after aggregating entries from all sources. This dual-layer approach ensures both parser-specific and global data hygiene.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →