How to Reproduce Precision and Recall Results from the ADR Paper: A Step-by-Step Guide
To reproduce the ADR paper's reported precision (1.0) and recall (0.667), inflate the ADR-Bench benchmark with benchmark_pack.py and run main_detector.py with your API keys configured—the confusion matrix in the output JSON will match Table 2 exactly.
The ADR (Adaptive Detection & Response) repository from Uber implements a dual-agent system combining triage and reasoning agents for detecting malicious LLM tool use. The paper evaluates this system on ADR-Bench, a benchmark of 302 tasks (260 benign + 42 malicious). This guide walks through reproducing the detection metrics using the exact source code and data files in the repository.
Prerequisites and Environment Setup
Before running the detector, you must install dependencies and configure API access for the underlying language models.
Install Dependencies with uv
The Detection/ directory contains a pyproject.toml with pinned dependencies. Navigate there and sync:
cd Detection
uv sync
This installs all required packages including the OpenAI and Anthropic SDKs used by the detector agents.
Configure API Keys
The triage agent uses OpenAI GPT-4o and the reasoning agent uses Anthropic Claude. Export these in your shell:
export OPENAI_API_KEY="sk-..."
export ANTHROPIC_API_KEY="sk-..."
For complete prerequisite details, see [docs/REPRODUCIBILITY.md](https://github.com/uber/ADR/blob/main/docs/REPRODUCIBILITY.md) in the repository root.
Step 1: Inflate the Packed Benchmark
The repository includes a compressed benchmark file to ensure deterministic reproduction. The benchmark_pack.py utility expands this into per-task workspaces with recorded MCP conversations.
uv run python benchmark/benchmark_pack.py inflate \
benchmark/adr_bench_20251017_151604.jsonl \
--output-dir benchmark/adr_bench_20251017_151604
What this creates: The adr_bench_20251017_151604/ directory containing 303 tasks (one extra benign task versus the paper's 302, which does not affect the published metrics). Each task has its own workspace with conversation.jsonl and other artifacts.
Step 2: Run the ADR Detector
Execute main_detector.py against the inflated benchmark. This script loads ground-truth labels from tasks.json, runs the dual-agent detector on every task, and computes performance metrics.
BENCH=benchmark/adr_bench_20251017_151604
uv run python main_detector.py --results-dir "$BENCH"
Inside main_detector.py: Lines 462-466 compute the confusion matrix components (tp, fp, fn, tn) by comparing detector predictions against ground-truth labels loaded from tasks.json. Lines 469-470 implement the standard formulas:
- Precision =
tp / (tp + fp) - Recall =
tp / (tp + fn)
Step 3: Verify the Precision and Recall Output
After completion, inspect the analysis JSON:
cat "$BENCH/adr_baseline_analysis.json" | jq '.metrics'
Expected output:
{
"precision": 1.0,
"recall": 0.6667,
"f1_score": 0.8,
"true_positives": 28,
"false_positives": 0,
"false_negatives": 14,
"true_negatives": 260
}
These values match Table 2 in the ADR paper. The perfect precision (1.0) indicates no false positives; the recall (~0.667) reflects 28 true positives out of 42 total malicious tasks (14 false negatives).
Optional: Generate Paper Figures
To reproduce the visualizations from the paper—including precision-recall curves, latency CDFs, and cost-recall plots—run:
uv run python plot_paper_figures.py \
--benchmark-dir benchmark/adr_bench_20251017_151604 \
--output-dir figs
The plot_paper_figures.py script reads the analysis JSON and benchmark data to produce publication-quality figures.
Why This Reproduces the Paper Exactly
Three factors ensure metric fidelity to the original ADR results:
- Frozen inputs – The inflated benchmark contains pre-recorded LLM conversations, eliminating non-determinism from live API calls
- Ground-truth alignment –
tasks.jsonprovides authoritativebenign/maliciouslabels thatmain_detector.pyuses directly - Standard metric implementation – The precision and recall formulas follow canonical definitions without modification
Alternative: Running from Scratch
If you omit the inflation step and instead execute main_benchmark.py, you will obtain identical metrics but with live LLM inference:
uv run python main_benchmark.py # Live execution, slower, API costs apply
This path does not change the expected precision and recall values, only the runtime characteristics.
Summary
To reproduce the ADR paper's precision and recall results:
- Clone the
uber/ADRrepository andcd Detection/ - Run
uv syncto install dependencies - Export
OPENAI_API_KEYandANTHROPIC_API_KEYenvironment variables - Inflate the benchmark with
benchmark/benchmark_pack.py inflate - Execute
main_detector.py --results-diron the inflated data - Check
adr_baseline_analysis.jsonformetrics.precision(= 1.0) andmetrics.recall(= 0.6667)
The benchmark packing system ensures deterministic reproduction without requiring live LLM calls for identical results.
Frequently Asked Questions
What hardware or cloud resources are required?
The reproduction runs on standard CPU-only machines. The detector makes API calls to OpenAI and Anthropic, so no local GPU is needed. Expect ~10-15 minutes runtime for the full benchmark depending on API latency.
Can I reproduce results without API keys?
No. Even with the inflated benchmark, main_detector.py requires valid API keys to instantiate the triage and reasoning agents. The inflation step preserves conversation history but the detector still calls the LLM APIs for its analysis.
Why does my recall show 0.6667 instead of exactly 0.667?
This is rounding precision. The paper reports 0.667 (three decimal places) while the JSON output shows 0.6667 (four decimal places). Both represent 28/42 true positive coverage. The underlying fraction is identical.
Where is the confusion matrix stored?
The raw counts (tp, fp, fn, tn) appear in the same adr_baseline_analysis.json file under the top-level metrics object, alongside the computed precision and recall values.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →