How ADR Achieves Zero False Positives on Benign Enterprise Workflows

ADR eliminates false positives on normal enterprise activity through a two-tier detection architecture that combines permissive high-recall triage with deep agentic reasoning.

The ADR (Agentic AI Detection and Response) system from Uber is designed to detect malicious AI agent behavior without disrupting legitimate business operations. According to the uber/ADR source code, this is accomplished through a deliberately layered approach that filters benign workflows at the earliest stage possible.

The Two-Tier Detection Architecture

ADR's Detection/ module implements a dual-agent pipeline that balances speed, cost, and precision. The architecture is explicitly designed so that ordinary enterprise workflows never trigger alerts.

High-Recall Triage Layer

The first tier in Detection/triage/ applies lightweight heuristics that are intentionally over-inclusive. Rather than attempting precise classification, this layer uses simple, fast checks to identify sessions that might exhibit risky behavior. The README describes this as combining "high‑recall triage with deeper agentic reasoning" to ensure genuine benign workflows pass through untouched.

This permissive approach means:

  • Low latency for the vast majority of sessions
  • Minimal cost by avoiding expensive model calls
  • Zero friction on normal enterprise operations

Deep Agentic Reasoning Layer

Sessions flagged by the triage layer proceed to a context-rich analysis implemented in Detection/context_providers/. This second tier performs thorough reasoning about:

  • Intent behind tool invocations
  • Execution traces across multi-step workflows
  • Contextual patterns that indicate genuine threats

Because only a small fraction of sessions reach this stage, the deep model can afford comprehensive analysis without impacting overall system performance.

Empirical Validation: 0% False Positive Rate

The ADR-Bench evaluation demonstrates this architecture's effectiveness. Table 2 in the README reports a 0% false-positive rate on the full benchmark suite—meaning normal enterprise workflows never generated spurious alerts during evaluation.

Reproduction steps are documented in docs/REPRODUCIBILITY.md for independent verification of these results.

Running the Two-Tier Pipeline

Default Dual-Agent Detection

Execute the full pipeline with both triage and deep reasoning layers:

cd ADR/Detection
uv sync                     # install dependencies

export ANTHROPIC_API_KEY="…"
export OPENAI_API_KEY="…"

# Run two-tier detection on benchmark suite

adr run --benchmark ./benchmark/adr_bench.jsonl

Triage-Only Mode

For quick smoke tests that bypass the expensive deep layer, use the lightweight detector:

adr run --detector llamafirewall --benchmark ./benchmark/adr_bench.jsonl

The llamafirewall option demonstrates the triage layer in isolation—useful for debugging or cost-sensitive scenarios, though it sacrifices the precision guarantees of the full pipeline.

Key Implementation Files

Path Purpose
Detection/triage/ Lightweight heuristics for the first-tier screening
Detection/context_providers/ Contextual reasoning logic for the deep analysis layer
Detection/README.md Usage documentation for invoking detectors
docs/REPRODUCIBILITY.md Benchmark methodology for the zero false-positive claim

Summary

  • Two-tier architecture prevents benign workflows from ever reaching expensive analysis
  • High-recall triage in Detection/triage/ applies permissive heuristics that rarely flag normal activity
  • Deep agentic reasoning in Detection/context_providers/ provides precise threat assessment for suspicious sessions only
  • 0% false-positive rate empirically validated on ADR-Bench evaluation set

Frequently Asked Questions

How does ADR's triage layer avoid catching benign enterprise workflows?

The triage layer uses deliberately permissive heuristics designed for high recall rather than precision. It errs on the side of allowing sessions through, ensuring that normal business operations—file access, API calls, data processing—rarely trigger flags. Only sessions with ambiguous or potentially risky patterns advance to deeper analysis.

Why use two tiers instead of a single precise model?

A single precise model would require expensive inference on every session, creating unacceptable latency and cost for enterprise deployments. The two-tier design keeps costs low by filtering benign traffic early with cheap heuristics, reserving comprehensive reasoning for the small subset of sessions that actually need it.

What is the llamafirewall detector and when should it be used?

The llamafirewall detector runs only the triage layer without deep reasoning. Use it for quick smoke tests, debugging, or scenarios where cost matters more than precision. It does not provide the zero false-positive guarantee of the full dual-agent pipeline.

Where is the zero false-positive claim documented?

The 0% false-positive rate is reported in Table 2 of the main README (lines 635-639) and can be reproduced using the methodology in docs/REPRODUCIBILITY.md. The benchmark evaluates ADR against the full ADR-Bench dataset of enterprise agent sessions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →