How ADR Achieves Zero False Positives on Benign Enterprise Workflows
ADR eliminates false positives on normal enterprise activity through a two-tier detection architecture that combines permissive high-recall triage with deep agentic reasoning.
The ADR (Agentic AI Detection and Response) system from Uber is designed to detect malicious AI agent behavior without disrupting legitimate business operations. According to the uber/ADR source code, this is accomplished through a deliberately layered approach that filters benign workflows at the earliest stage possible.
The Two-Tier Detection Architecture
ADR's Detection/ module implements a dual-agent pipeline that balances speed, cost, and precision. The architecture is explicitly designed so that ordinary enterprise workflows never trigger alerts.
High-Recall Triage Layer
The first tier in Detection/triage/ applies lightweight heuristics that are intentionally over-inclusive. Rather than attempting precise classification, this layer uses simple, fast checks to identify sessions that might exhibit risky behavior. The README describes this as combining "high‑recall triage with deeper agentic reasoning" to ensure genuine benign workflows pass through untouched.
This permissive approach means:
- Low latency for the vast majority of sessions
- Minimal cost by avoiding expensive model calls
- Zero friction on normal enterprise operations
Deep Agentic Reasoning Layer
Sessions flagged by the triage layer proceed to a context-rich analysis implemented in Detection/context_providers/. This second tier performs thorough reasoning about:
- Intent behind tool invocations
- Execution traces across multi-step workflows
- Contextual patterns that indicate genuine threats
Because only a small fraction of sessions reach this stage, the deep model can afford comprehensive analysis without impacting overall system performance.
Empirical Validation: 0% False Positive Rate
The ADR-Bench evaluation demonstrates this architecture's effectiveness. Table 2 in the README reports a 0% false-positive rate on the full benchmark suite—meaning normal enterprise workflows never generated spurious alerts during evaluation.
Reproduction steps are documented in docs/REPRODUCIBILITY.md for independent verification of these results.
Running the Two-Tier Pipeline
Default Dual-Agent Detection
Execute the full pipeline with both triage and deep reasoning layers:
cd ADR/Detection
uv sync # install dependencies
export ANTHROPIC_API_KEY="…"
export OPENAI_API_KEY="…"
# Run two-tier detection on benchmark suite
adr run --benchmark ./benchmark/adr_bench.jsonl
Triage-Only Mode
For quick smoke tests that bypass the expensive deep layer, use the lightweight detector:
adr run --detector llamafirewall --benchmark ./benchmark/adr_bench.jsonl
The llamafirewall option demonstrates the triage layer in isolation—useful for debugging or cost-sensitive scenarios, though it sacrifices the precision guarantees of the full pipeline.
Key Implementation Files
| Path | Purpose |
|---|---|
Detection/triage/ |
Lightweight heuristics for the first-tier screening |
Detection/context_providers/ |
Contextual reasoning logic for the deep analysis layer |
Detection/README.md |
Usage documentation for invoking detectors |
docs/REPRODUCIBILITY.md |
Benchmark methodology for the zero false-positive claim |
Summary
- Two-tier architecture prevents benign workflows from ever reaching expensive analysis
- High-recall triage in
Detection/triage/applies permissive heuristics that rarely flag normal activity - Deep agentic reasoning in
Detection/context_providers/provides precise threat assessment for suspicious sessions only - 0% false-positive rate empirically validated on ADR-Bench evaluation set
Frequently Asked Questions
How does ADR's triage layer avoid catching benign enterprise workflows?
The triage layer uses deliberately permissive heuristics designed for high recall rather than precision. It errs on the side of allowing sessions through, ensuring that normal business operations—file access, API calls, data processing—rarely trigger flags. Only sessions with ambiguous or potentially risky patterns advance to deeper analysis.
Why use two tiers instead of a single precise model?
A single precise model would require expensive inference on every session, creating unacceptable latency and cost for enterprise deployments. The two-tier design keeps costs low by filtering benign traffic early with cheap heuristics, reserving comprehensive reasoning for the small subset of sessions that actually need it.
What is the llamafirewall detector and when should it be used?
The llamafirewall detector runs only the triage layer without deep reasoning. Use it for quick smoke tests, debugging, or scenarios where cost matters more than precision. It does not provide the zero false-positive guarantee of the full dual-agent pipeline.
Where is the zero false-positive claim documented?
The 0% false-positive rate is reported in Table 2 of the main README (lines 635-639) and can be reproduced using the methodology in docs/REPRODUCIBILITY.md. The benchmark evaluates ADR against the full ADR-Bench dataset of enterprise agent sessions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →