How to Reproduce Precision and Recall Results from the ADR Paper: A Step-by-Step Guide

To reproduce the ADR paper's reported precision (1.0) and recall (0.667), inflate the ADR-Bench benchmark with benchmark_pack.py and run main_detector.py with your API keys configured—the confusion matrix in the output JSON will match Table 2 exactly.

The ADR (Adaptive Detection & Response) repository from Uber implements a dual-agent system combining triage and reasoning agents for detecting malicious LLM tool use. The paper evaluates this system on ADR-Bench, a benchmark of 302 tasks (260 benign + 42 malicious). This guide walks through reproducing the detection metrics using the exact source code and data files in the repository.

Prerequisites and Environment Setup

Before running the detector, you must install dependencies and configure API access for the underlying language models.

Install Dependencies with uv

The Detection/ directory contains a pyproject.toml with pinned dependencies. Navigate there and sync:

cd Detection
uv sync

This installs all required packages including the OpenAI and Anthropic SDKs used by the detector agents.

Configure API Keys

The triage agent uses OpenAI GPT-4o and the reasoning agent uses Anthropic Claude. Export these in your shell:

export OPENAI_API_KEY="sk-..."
export ANTHROPIC_API_KEY="sk-..."

For complete prerequisite details, see [docs/REPRODUCIBILITY.md](https://github.com/uber/ADR/blob/main/docs/REPRODUCIBILITY.md) in the repository root.

Step 1: Inflate the Packed Benchmark

The repository includes a compressed benchmark file to ensure deterministic reproduction. The benchmark_pack.py utility expands this into per-task workspaces with recorded MCP conversations.

uv run python benchmark/benchmark_pack.py inflate \
    benchmark/adr_bench_20251017_151604.jsonl \
    --output-dir benchmark/adr_bench_20251017_151604

What this creates: The adr_bench_20251017_151604/ directory containing 303 tasks (one extra benign task versus the paper's 302, which does not affect the published metrics). Each task has its own workspace with conversation.jsonl and other artifacts.

Step 2: Run the ADR Detector

Execute main_detector.py against the inflated benchmark. This script loads ground-truth labels from tasks.json, runs the dual-agent detector on every task, and computes performance metrics.

BENCH=benchmark/adr_bench_20251017_151604
uv run python main_detector.py --results-dir "$BENCH"

Inside main_detector.py: Lines 462-466 compute the confusion matrix components (tp, fp, fn, tn) by comparing detector predictions against ground-truth labels loaded from tasks.json. Lines 469-470 implement the standard formulas:

  • Precision = tp / (tp + fp)
  • Recall = tp / (tp + fn)

Step 3: Verify the Precision and Recall Output

After completion, inspect the analysis JSON:

cat "$BENCH/adr_baseline_analysis.json" | jq '.metrics'

Expected output:

{
  "precision": 1.0,
  "recall": 0.6667,
  "f1_score": 0.8,
  "true_positives": 28,
  "false_positives": 0,
  "false_negatives": 14,
  "true_negatives": 260
}

These values match Table 2 in the ADR paper. The perfect precision (1.0) indicates no false positives; the recall (~0.667) reflects 28 true positives out of 42 total malicious tasks (14 false negatives).

Optional: Generate Paper Figures

To reproduce the visualizations from the paper—including precision-recall curves, latency CDFs, and cost-recall plots—run:

uv run python plot_paper_figures.py \
    --benchmark-dir benchmark/adr_bench_20251017_151604 \
    --output-dir figs

The plot_paper_figures.py script reads the analysis JSON and benchmark data to produce publication-quality figures.

Why This Reproduces the Paper Exactly

Three factors ensure metric fidelity to the original ADR results:

  • Frozen inputs – The inflated benchmark contains pre-recorded LLM conversations, eliminating non-determinism from live API calls
  • Ground-truth alignment – tasks.json provides authoritative benign/malicious labels that main_detector.py uses directly
  • Standard metric implementation – The precision and recall formulas follow canonical definitions without modification

Alternative: Running from Scratch

If you omit the inflation step and instead execute main_benchmark.py, you will obtain identical metrics but with live LLM inference:

uv run python main_benchmark.py  # Live execution, slower, API costs apply

This path does not change the expected precision and recall values, only the runtime characteristics.

Summary

To reproduce the ADR paper's precision and recall results:

  • Clone the uber/ADR repository and cd Detection/
  • Run uv sync to install dependencies
  • Export OPENAI_API_KEY and ANTHROPIC_API_KEY environment variables
  • Inflate the benchmark with benchmark/benchmark_pack.py inflate
  • Execute main_detector.py --results-dir on the inflated data
  • Check adr_baseline_analysis.json for metrics.precision (= 1.0) and metrics.recall (= 0.6667)

The benchmark packing system ensures deterministic reproduction without requiring live LLM calls for identical results.

Frequently Asked Questions

What hardware or cloud resources are required?

The reproduction runs on standard CPU-only machines. The detector makes API calls to OpenAI and Anthropic, so no local GPU is needed. Expect ~10-15 minutes runtime for the full benchmark depending on API latency.

Can I reproduce results without API keys?

No. Even with the inflated benchmark, main_detector.py requires valid API keys to instantiate the triage and reasoning agents. The inflation step preserves conversation history but the detector still calls the LLM APIs for its analysis.

Why does my recall show 0.6667 instead of exactly 0.667?

This is rounding precision. The paper reports 0.667 (three decimal places) while the JSON output shows 0.6667 (four decimal places). Both represent 28/42 true positive coverage. The underlying fraction is identical.

Where is the confusion matrix stored?

The raw counts (tp, fp, fn, tn) appear in the same adr_baseline_analysis.json file under the top-level metrics object, alongside the computed precision and recall values.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →