# How ADR-Bench Benchmarks AI Agent Security with 300+ Tasks: A Complete Technical Breakdown

> ADR-Bench benchmarks AI agent security with 303 tasks. Discover its unique JSON catalogue, MCP server, and three-step workflow for comprehensive evaluation and visualization.

- Repository: [Uber Open Source/ADR](https://github.com/uber/ADR)
- Tags: deep-dive
- Published: 2026-08-07

---

**ADR-Bench evaluates AI agent security through 303 benchmark tasks (261 benign workflows + 42 malicious attack scenarios) using a unified JSON catalogue, MCP server fixtures, and a reproducible three-step workflow: inflate, detect, and visualize.**

ADR-Bench is the evaluation engine inside **ADR** (Agentic AI Detection and Response), Uber's open-source framework for securing AI agents. It measures how well security detectors distinguish legitimate business workflows from adversarial attacks in realistic enterprise environments. This article explains the architecture, key components, and practical usage based on the actual implementation in `uber/ADR`.

---

## Core Architecture of ADR-Bench

ADR-Bench operates on three interconnected concepts: a unified task catalogue, MCP-based tool sandboxing, and deterministic execution orchestration.

### Task Catalogue: 303 Scenarios in One JSON File

The [`Detection/tasks.json`](https://github.com/uber/ADR/blob/main/Detection/tasks.json) file contains **303 benchmark tasks**—**261 benign business-workflow tasks** and **42 malicious attack scenarios** covering all 17 agent-attack techniques from the ADR research paper.

Each task entry follows a consistent schema:

- `task_id`: Unique identifier
- `description`: Human-readable summary
- `user_prompt`: The actual input sent to the agent
- `required_mcp_servers`: List of MCP servers needed (e.g., `memory`, `arxiv_toolkit`)
- `expected_tool_calls`: Expected sequence of tool invocations
- `label`: Ground-truth classification (`benign` or `malicious`)

Unlike benchmarks such as AgentDojo that separate ground truth into a secondary file, ADR-Bench embeds labels directly in each task entry. This simplifies validation and eliminates synchronization issues between task definitions and labels.

### MCP Server Fixtures: Realistic Tool Invocation

The [`Detection/mcp_servers_registry.json`](https://github.com/uber/ADR/blob/main/Detection/mcp_servers_registry.json) maps logical server names to concrete tool implementations. When [`main_benchmark.py`](https://github.com/uber/ADR/blob/main/main_benchmark.py) executes a task, it automatically spins up the required **MCP (Multi-Channel Platform)** servers so the agent can invoke them via the **ADR Sensor** protocol.

This design ensures that:

- Tool-use behavior is captured faithfully for detection analysis
- Safety policies can be enforced at the MCP layer
- Benchmark tasks reflect realistic enterprise integrations

---

## The Benchmark Driver: Executing 300+ Tasks

The [`Detection/main_benchmark.py`](https://github.com/uber/ADR/blob/main/Detection/main_benchmark.py) file orchestrates the entire evaluation. It provides deterministic, configurable execution through the `Config` class and `TaskManager` utilities.

### Loading Tasks Programmatically

```python
from Detection.main_benchmark import TaskManager, Config

# Initialize configuration from config_benchmark.yaml

cfg = Config()

# Load all 303 tasks

tm = TaskManager(cfg)
tasks = tm.load_ads_tasks()  # Returns list of dicts from tasks.json

# Inspect first task

first_task = tasks[0]
print(f"Task {first_task['task_id']}: {first_task['description']}")
print(f"Label: {first_task['label']}")
print(f"Required MCP servers: {first_task['required_mcp_servers']}")

```

### Execution Flow and Concurrency

The driver:

1. Loads tasks via `TaskManager.load_ads_tasks()`
2. Launches specified MCP servers per task
3. Runs the target agent (Claude Code CLI by default) against each prompt
4. Records conversation logs to `benchmark/adr_bench_<timestamp>/task_<id>/`

Concurrency is capped by `max_concurrent_tasks` (default **10**) in [`Detection/config_benchmark.yaml`](https://github.com/uber/ADR/blob/main/Detection/config_benchmark.yaml), keeping resource usage predictable across the full 300+ task suite.

---

## Fast-Path Evaluation: Packed Benchmarks

Running 300+ agent-MCP interactions is expensive. ADR-Bench provides a **packed benchmark shortcut** to bypass re-execution.

### Inflating Pre-Generated Conversations

The [`benchmark/benchmark_pack.py`](https://github.com/uber/ADR/blob/main/benchmark/benchmark_pack.py) utility unpacks a pre-inflated JSONL file into the standard directory structure:

```bash

# Inflate the packed benchmark (no API calls required)

uv run python benchmark/benchmark_pack.py inflate \
    benchmark/adr_bench_20251017_151604.jsonl \
    --output-dir benchmark/adr_bench_20251017_151604

```

This creates identical `task_<id>/conversation.json` files as a live run, enabling rapid detector evaluation without additional LLM API costs.

---

## Detection and Evaluation Pipeline

After conversations are generated, [`Detection/main_detector.py`](https://github.com/uber/ADR/blob/main/Detection/main_detector.py) evaluates security detectors against the benchmark.

### Running Detectors on Benchmark Output

```bash

# Set benchmark directory

BENCH=benchmark/adr_bench_20251017_151604

# Run ADR detector (dual-agent architecture)

uv run python main_detector.py --results-dir "$BENCH"

# Output: adr_baseline_analysis.json with precision, recall, FPR metrics

```

The detector loader:

- Retrieves ground-truth labels from [`tasks.json`](https://github.com/uber/ADR/blob/main/tasks.json)
- Applies the chosen detector (default **ADR dual-agent**, or baselines like **LlamaFirewall**)
- Writes per-task analysis to `*_baseline_analysis.json`
- Aggregates metrics: **precision**, **recall**, **false-positive count**

### Baseline Detector Interface

All detectors implement the `BaseDetector` abstract class in [`Detection/guardrail/base_detector.py`](https://github.com/uber/ADR/blob/main/Detection/guardrail/base_detector.py). This extensible interface enables direct comparison between ADR and alternative approaches without modifying evaluation infrastructure.

### Reproducing Paper Figures

```bash

# Generate publication-ready visualizations

uv run python plot_paper_figures.py \
    --benchmark-dir "$BENCH" \
    --output-dir figs

```

The [`plot_paper_figures.py`](https://github.com/uber/ADR/blob/main/plot_paper_figures.py) script consumes analysis JSON files to recreate the exact figures from the ADR research paper.

---

## Configuration and Reproducibility

### Benchmark Configuration

[`Detection/config_benchmark.yaml`](https://github.com/uber/ADR/blob/main/Detection/config_benchmark.yaml) controls execution parameters:

```yaml
max_concurrent_tasks: 10      # Parallel task limit

timeout_seconds: 300          # Per-task timeout

restricted_tools: []          # Tools to disable for safety

agent_cli: "claude"           # Target agent interface

```

### Full Reproduction Workflow

The [`docs/REPRODUCIBILITY.md`](https://github.com/uber/ADR/blob/main/docs/REPRODUCIBILITY.md) document specifies the end-to-end pipeline:

1. **Inflate or generate** benchmark conversations
2. **Run detectors** on the conversation dataset
3. **Plot figures** from detection analysis

This ensures the 300+ tasks—including the additional benign task added post-publication—can be reproduced exactly.

---

## Key Files and Their Roles

| File | Purpose |
|------|---------|
| [`Detection/tasks.json`](https://github.com/uber/ADR/blob/main/Detection/tasks.json) | Master catalogue of 303 tasks with prompts, MCP requirements, and ground-truth labels |
| [`Detection/config_benchmark.yaml`](https://github.com/uber/ADR/blob/main/Detection/config_benchmark.yaml) | Execution parameters: concurrency, timeouts, tool restrictions |
| [`Detection/main_benchmark.py`](https://github.com/uber/ADR/blob/main/Detection/main_benchmark.py) | Core driver for task execution and conversation recording |
| [`Detection/main_detector.py`](https://github.com/uber/ADR/blob/main/Detection/main_detector.py) | Detector evaluation engine and metric computation |
| [`Detection/benchmark/benchmark_pack.py`](https://github.com/uber/ADR/blob/main/Detection/benchmark/benchmark_pack.py) | Pack/unpack utility for fast benchmark inflation |
| [`Detection/guardrail/base_detector.py`](https://github.com/uber/ADR/blob/main/Detection/guardrail/base_detector.py) | Abstract base class for all detector implementations |
| [`docs/REPRODUCIBILITY.md`](https://github.com/uber/ADR/blob/main/docs/REPRODUCIBILITY.md) | Step-by-step guide for full pipeline reproduction |

---

## Summary

- **ADR-Bench** evaluates AI agent security across **303 tasks** using a unified JSON schema with embedded ground-truth labels in [`tasks.json`](https://github.com/uber/ADR/blob/main/tasks.json)
- **MCP server fixtures** provide realistic tool-sandboxing through [`mcp_servers_registry.json`](https://github.com/uber/ADR/blob/main/mcp_servers_registry.json), capturing authentic agent behavior for detection analysis
- **Deterministic execution** via [`main_benchmark.py`](https://github.com/uber/ADR/blob/main/main_benchmark.py) with configurable concurrency prevents resource exhaustion during full-suite evaluation
- **Packed benchmarks** enable rapid iteration: inflate pre-generated conversations, run detectors, and plot results without costly API calls
- **Extensible detector interface** in [`base_detector.py`](https://github.com/uber/ADR/blob/main/base_detector.py) supports direct comparison between ADR and baseline security systems

---

## Frequently Asked Questions

### How many malicious vs. benign tasks does ADR-Bench include?

ADR-Bench contains **42 malicious attack scenarios** and **261 benign business-workflow tasks**, totaling 303 tasks. The malicious set covers all 17 agent-attack techniques described in the ADR research paper, while benign tasks represent realistic enterprise operations like calendar scheduling, document retrieval, and data analysis.

### Can I run ADR-Bench without making expensive API calls to Claude or OpenAI?

Yes. Use the **packed benchmark shortcut**: run `benchmark/benchmark_pack.py inflate` on `adr_bench_20251017_151604.jsonl` to unpack pre-generated conversation logs. This creates the same directory structure as live execution, enabling detector evaluation and figure generation without additional LLM API costs.

### What makes ADR-Bench different from AgentDojo and other agent security benchmarks?

ADR-Bench embeds ground-truth labels directly in [`tasks.json`](https://github.com/uber/ADR/blob/main/tasks.json) rather than maintaining separate ground-truth files. It uses **MCP servers** for realistic tool sandboxing rather than mocked environments. The benchmark also provides deterministic concurrency controls and a packed-archive workflow specifically designed for iterative detector development.

### How do I add a new detector to compare against ADR's dual-agent approach?

Implement the `BaseDetector` abstract class from [`Detection/guardrail/base_detector.py`](https://github.com/uber/ADR/blob/main/Detection/guardrail/base_detector.py). Your detector will automatically integrate with [`main_detector.py`](https://github.com/uber/ADR/blob/main/main_detector.py) and [`plot_paper_figures.py`](https://github.com/uber/ADR/blob/main/plot_paper_figures.py). The existing LlamaFirewall baseline demonstrates the required interface methods for loading, analyzing conversations, and returning detection scores.