# How to Reproduce Precision and Recall Results from the ADR Paper: A Step-by-Step Guide

> Reproduce ADR paper precision and recall results by inflating the ADR benchmark and running main_detector.py. Match Table 2 exactly with this step-by-step guide.

- Repository: [Uber Open Source/ADR](https://github.com/uber/ADR)
- Tags: how-to-guide
- Published: 2026-08-06

---

**To reproduce the ADR paper's reported precision (1.0) and recall (0.667), inflate the ADR-Bench benchmark with [`benchmark_pack.py`](https://github.com/uber/ADR/blob/main/benchmark_pack.py) and run [`main_detector.py`](https://github.com/uber/ADR/blob/main/main_detector.py) with your API keys configured—the confusion matrix in the output JSON will match Table 2 exactly.**

The **ADR (Adaptive Detection & Response)** repository from Uber implements a dual-agent system combining triage and reasoning agents for detecting malicious LLM tool use. The paper evaluates this system on **ADR-Bench**, a benchmark of 302 tasks (260 benign + 42 malicious). This guide walks through reproducing the detection metrics using the exact source code and data files in the repository.

## Prerequisites and Environment Setup

Before running the detector, you must install dependencies and configure API access for the underlying language models.

### Install Dependencies with `uv`

The `Detection/` directory contains a [`pyproject.toml`](https://github.com/uber/ADR/blob/main/pyproject.toml) with pinned dependencies. Navigate there and sync:

```bash
cd Detection
uv sync

```

This installs all required packages including the OpenAI and Anthropic SDKs used by the detector agents.

### Configure API Keys

The triage agent uses **OpenAI GPT-4o** and the reasoning agent uses **Anthropic Claude**. Export these in your shell:

```bash
export OPENAI_API_KEY="sk-..."
export ANTHROPIC_API_KEY="sk-..."

```

For complete prerequisite details, see [[`docs/REPRODUCIBILITY.md`](https://github.com/uber/ADR/blob/main/docs/REPRODUCIBILITY.md)](https://github.com/uber/ADR/blob/main/docs/REPRODUCIBILITY.md) in the repository root.

## Step 1: Inflate the Packed Benchmark

The repository includes a compressed benchmark file to ensure deterministic reproduction. The [`benchmark_pack.py`](https://github.com/uber/ADR/blob/main/benchmark_pack.py) utility expands this into per-task workspaces with recorded MCP conversations.

```bash
uv run python benchmark/benchmark_pack.py inflate \
    benchmark/adr_bench_20251017_151604.jsonl \
    --output-dir benchmark/adr_bench_20251017_151604

```

**What this creates:** The `adr_bench_20251017_151604/` directory containing 303 tasks (one extra benign task versus the paper's 302, which does not affect the published metrics). Each task has its own workspace with `conversation.jsonl` and other artifacts.

## Step 2: Run the ADR Detector

Execute [`main_detector.py`](https://github.com/uber/ADR/blob/main/main_detector.py) against the inflated benchmark. This script loads ground-truth labels from [`tasks.json`](https://github.com/uber/ADR/blob/main/tasks.json), runs the dual-agent detector on every task, and computes performance metrics.

```bash
BENCH=benchmark/adr_bench_20251017_151604
uv run python main_detector.py --results-dir "$BENCH"

```

**Inside [`main_detector.py`](https://github.com/uber/ADR/blob/main/main_detector.py):** Lines 462-466 compute the confusion matrix components (`tp`, `fp`, `fn`, `tn`) by comparing detector predictions against ground-truth labels loaded from [`tasks.json`](https://github.com/uber/ADR/blob/main/tasks.json). Lines 469-470 implement the standard formulas:

- **Precision** = `tp / (tp + fp)`
- **Recall** = `tp / (tp + fn)`

## Step 3: Verify the Precision and Recall Output

After completion, inspect the analysis JSON:

```bash
cat "$BENCH/adr_baseline_analysis.json" | jq '.metrics'

```

**Expected output:**

```json
{
  "precision": 1.0,
  "recall": 0.6667,
  "f1_score": 0.8,
  "true_positives": 28,
  "false_positives": 0,
  "false_negatives": 14,
  "true_negatives": 260
}

```

These values match **Table 2** in the ADR paper. The perfect precision (1.0) indicates no false positives; the recall (~0.667) reflects 28 true positives out of 42 total malicious tasks (14 false negatives).

## Optional: Generate Paper Figures

To reproduce the visualizations from the paper—including precision-recall curves, latency CDFs, and cost-recall plots—run:

```bash
uv run python plot_paper_figures.py \
    --benchmark-dir benchmark/adr_bench_20251017_151604 \
    --output-dir figs

```

The [`plot_paper_figures.py`](https://github.com/uber/ADR/blob/main/plot_paper_figures.py) script reads the analysis JSON and benchmark data to produce publication-quality figures.

## Why This Reproduces the Paper Exactly

Three factors ensure metric fidelity to the original ADR results:

- **Frozen inputs** – The inflated benchmark contains pre-recorded LLM conversations, eliminating non-determinism from live API calls
- **Ground-truth alignment** – [`tasks.json`](https://github.com/uber/ADR/blob/main/tasks.json) provides authoritative `benign`/`malicious` labels that [`main_detector.py`](https://github.com/uber/ADR/blob/main/main_detector.py) uses directly
- **Standard metric implementation** – The precision and recall formulas follow canonical definitions without modification

## Alternative: Running from Scratch

If you omit the inflation step and instead execute [`main_benchmark.py`](https://github.com/uber/ADR/blob/main/main_benchmark.py), you will obtain identical metrics but with live LLM inference:

```bash
uv run python main_benchmark.py  # Live execution, slower, API costs apply

```

This path does not change the expected precision and recall values, only the runtime characteristics.

## Summary

To reproduce the ADR paper's precision and recall results:

- Clone the `uber/ADR` repository and `cd Detection/`
- Run `uv sync` to install dependencies
- Export `OPENAI_API_KEY` and `ANTHROPIC_API_KEY` environment variables
- Inflate the benchmark with `benchmark/benchmark_pack.py inflate`
- Execute `main_detector.py --results-dir` on the inflated data
- Check [`adr_baseline_analysis.json`](https://github.com/uber/ADR/blob/main/adr_baseline_analysis.json) for `metrics.precision` (= 1.0) and `metrics.recall` (= 0.6667)

The benchmark packing system ensures deterministic reproduction without requiring live LLM calls for identical results.

## Frequently Asked Questions

### What hardware or cloud resources are required?

The reproduction runs on standard CPU-only machines. The detector makes API calls to OpenAI and Anthropic, so no local GPU is needed. Expect ~10-15 minutes runtime for the full benchmark depending on API latency.

### Can I reproduce results without API keys?

No. Even with the inflated benchmark, [`main_detector.py`](https://github.com/uber/ADR/blob/main/main_detector.py) requires valid API keys to instantiate the triage and reasoning agents. The inflation step preserves conversation history but the detector still calls the LLM APIs for its analysis.

### Why does my recall show 0.6667 instead of exactly 0.667?

This is rounding precision. The paper reports **0.667** (three decimal places) while the JSON output shows **0.6667** (four decimal places). Both represent 28/42 true positive coverage. The underlying fraction is identical.

### Where is the confusion matrix stored?

The raw counts (`tp`, `fp`, `fn`, `tn`) appear in the same [`adr_baseline_analysis.json`](https://github.com/uber/ADR/blob/main/adr_baseline_analysis.json) file under the top-level `metrics` object, alongside the computed precision and recall values.