Understanding the `has_meaningful_content()` Filtering Logic in Uber ADR

The has_meaningful_content() method in Uber's ADR (AI Detection and Response) system filters AI-agent telemetry by requiring either multi-turn conversations, tool usage, or substantive text content while rejecting empty events, very short messages, and known noise patterns like authentication errors.

In the Uber/ADR repository, this method serves as the primary gatekeeper for AI-agent observability data. Every parsed log entry from agents like Cursor, Claude, Warp, and Codex must pass this filter before downstream security analytics and benchmarking pipelines process it. The following sections break down exactly how this filtering works, where it applies across the codebase, and how to interpret its behavior when building on or extending the ADR sensor.

How has_meaningful_content() Works

The filtering logic resides in Sensor/adr_sensor/schemas/agent_event_schema.py at line 118, implemented as a method on the AgentEvent dataclass. The method evaluates five sequential criteria to determine whether an event merits retention.

Step 1: Reject Empty Chat History

if not self.chat_history:
    return False

Events without any messages contain no analyzable context and are immediately discarded.

Step 2: Accept Multi-Turn Conversations

if len(self.chat_history) >= 2:
    return True

Any interaction with two or more turns qualifies as meaningful automatically—there is demonstrable back-and-forth between user and agent.

Step 3: Accept Single-Turn Events with Tool Usage

if msg.tools:
    return True

Even single messages are retained if the assistant invoked a tool. Tool usage (e.g., MCP tool calls, code generation functions) signals actionable agent behavior regardless of conversational depth.

Step 4: Filter Minimal Textual Content

content = msg.content.strip()
if len(content) <= 5:
    return False

Strings of five characters or fewer—such as "OK", "…", or "yes"—provide no analytical value and are rejected.

Step 5: Filter Known Noise Patterns

NOISE_PATTERNS = (
    "invalid api key",
    "please run /login",
    "authentication failed",
    "unauthorized",
    "session expired",
    "rate limit",
    "api error",
    "warmup",
)
if any(pattern in content.lower() for pattern in NOISE_PATTERNS):
    return False

System-generated errors and warm-up messages that do not reflect genuine user-driven activity are excluded. This list is case-insensitive for robust matching.

If none of these conditions return False, the method ultimately returns True, admitting the event for downstream processing.

Where the Filter Is Applied

Multiple parsers across the ADR sensor invoke has_meaningful_content() to prune raw telemetry:

In observer.py, the typical pattern aggregates entries from all parsers then applies the filter:

filtered = [e for e in entries if e.has_meaningful_content()]

This centralized pruning ensures only relevant telemetry persists to storage and analytics.

Code Examples

Multi-Turn Conversation (Retained)

from datetime import datetime
from adr_sensor.schemas.agent_event_schema import AgentEvent, ChatMessage

event = AgentEvent(
    timestamp=datetime.utcnow(),
    source="claude",
    session_id="sess_01",
    chat_history=[
        ChatMessage(role="user", content="How do I sort a list?"),
        ChatMessage(role="assistant", content="You can use the `sorted()` function.", tools=[]),
    ],
)

assert event.has_meaningful_content()   # → True

Single-Turn with Tool Usage (Retained)

event = AgentEvent(
    timestamp=datetime.utcnow(),
    source="cursor",
    session_id="sess_02",
    chat_history=[
        ChatMessage(
            role="assistant",
            content="",
            tools=[ToolUsage(tool_name="search", tool_type="mcp_tool", arguments={"query": "grep"})],
        ),
    ],
)

assert event.has_meaningful_content()   # → True

Short Error Message (Filtered)

event = AgentEvent(
    timestamp=datetime.utcnow(),
    source="codex",
    session_id="sess_03",
    chat_history=[ChatMessage(role="assistant", content="OK")],
)

assert not event.has_meaningful_content()   # → False

Known Noise Pattern (Filtered)

event = AgentEvent(
    timestamp=datetime.utcnow(),
    source="warp",
    session_id="sess_04",
    chat_history=[ChatMessage(role="assistant", content="Rate limit exceeded. Please try later.")],
)

assert not event.has_meaningging_content()   # → False

Key Source Files

File Purpose
Sensor/adr_sensor/schemas/agent_event_schema.py Defines AgentEvent dataclass and has_meaningful_content() method
Sensor/adr_sensor/parsers/warp_parser.py Applies filter to Warp agent logs
Sensor/adr_sensor/parsers/cursor_parser.py Applies filter to Cursor agent logs
Sensor/adr_sensor/parsers/codex_parser.py Applies filter to Codex agent logs
Sensor/adr_sensor/parsers/claude_parser.py Applies filter to Claude agent logs
Sensor/adr_sensor/parsers/claude_desktop_parser.py Applies filter to Claude Desktop logs
Sensor/adr_sensor/parsers/cline_parser.py Applies filter to Cline agent logs
Sensor/adr_sensor/observer.py Aggregates and final-filters all parser outputs

Architectural Impact

Data hygiene — The schema guarantee that every stored AgentEvent contributes actionable insight eliminates storage of empty or irrelevant records.

Performance — Downstream components including benchmarking suites, security guardrails, and export pipelines operate on leaner datasets, reducing CPU and I/O overhead.

Security focus — Noise-free logs improve signal-to-noise ratio for threat-intelligence modules consuming these events, enabling more reliable anomaly detection.

Summary

  • has_meaningful_content() is defined in Sensor/adr_sensor/schemas/agent_event_schema.py as an AgentEvent method
  • Multi-turn conversations (≥2 messages) are automatically retained
  • Single-turn events are retained only with tool usage or substantial text (>5 characters)
  • Empty events, very short messages, and known noise patterns (authentication errors, rate limits, warmups) are filtered
  • The filter is invoked by all major parsers and the central observer.py aggregator

Frequently Asked Questions

What counts as "meaningful content" in Uber ADR?

According to the Uber/ADR source code, meaningful content includes: (1) any conversation with two or more turns, (2) single-turn events where the assistant used a tool, or (3) single-turn events with more than five characters of text that do not match known noise patterns. Everything else is discarded.

How do I add new noise patterns to the filter?

Modify the NOISE_PATTERNS tuple in Sensor/adr_sensor/schemas/agent_event_schema.py. Add lowercase strings to match against; the comparison is case-insensitive. Rebuild and redeploy the sensor for changes to take effect across all parsers.

Why does the filter accept empty content if tools are present?

Tool invocations indicate the AI agent performed an actionable operation—searching code, generating files, calling APIs—even when the conversational content is minimal. This captures functionally significant behavior that security and audit workflows need to track.

Where does filtering happen in the ADR pipeline?

Each parser (warp_parser.py, cursor_parser.py, etc.) can apply has_meaningful_content(), and the observer.py module performs a final filtering pass after aggregating entries from all sources. This dual-layer approach ensures both parser-specific and global data hygiene.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →