How ADR-Bench Benchmarks AI Agent Security with 300+ Tasks: A Complete Technical Breakdown

ADR-Bench evaluates AI agent security through 303 benchmark tasks (261 benign workflows + 42 malicious attack scenarios) using a unified JSON catalogue, MCP server fixtures, and a reproducible three-step workflow: inflate, detect, and visualize.

ADR-Bench is the evaluation engine inside ADR (Agentic AI Detection and Response), Uber's open-source framework for securing AI agents. It measures how well security detectors distinguish legitimate business workflows from adversarial attacks in realistic enterprise environments. This article explains the architecture, key components, and practical usage based on the actual implementation in uber/ADR.


Core Architecture of ADR-Bench

ADR-Bench operates on three interconnected concepts: a unified task catalogue, MCP-based tool sandboxing, and deterministic execution orchestration.

Task Catalogue: 303 Scenarios in One JSON File

The Detection/tasks.json file contains 303 benchmark tasks—261 benign business-workflow tasks and 42 malicious attack scenarios covering all 17 agent-attack techniques from the ADR research paper.

Each task entry follows a consistent schema:

  • task_id: Unique identifier
  • description: Human-readable summary
  • user_prompt: The actual input sent to the agent
  • required_mcp_servers: List of MCP servers needed (e.g., memory, arxiv_toolkit)
  • expected_tool_calls: Expected sequence of tool invocations
  • label: Ground-truth classification (benign or malicious)

Unlike benchmarks such as AgentDojo that separate ground truth into a secondary file, ADR-Bench embeds labels directly in each task entry. This simplifies validation and eliminates synchronization issues between task definitions and labels.

MCP Server Fixtures: Realistic Tool Invocation

The Detection/mcp_servers_registry.json maps logical server names to concrete tool implementations. When main_benchmark.py executes a task, it automatically spins up the required MCP (Multi-Channel Platform) servers so the agent can invoke them via the ADR Sensor protocol.

This design ensures that:

  • Tool-use behavior is captured faithfully for detection analysis
  • Safety policies can be enforced at the MCP layer
  • Benchmark tasks reflect realistic enterprise integrations

The Benchmark Driver: Executing 300+ Tasks

The Detection/main_benchmark.py file orchestrates the entire evaluation. It provides deterministic, configurable execution through the Config class and TaskManager utilities.

Loading Tasks Programmatically

from Detection.main_benchmark import TaskManager, Config

# Initialize configuration from config_benchmark.yaml

cfg = Config()

# Load all 303 tasks

tm = TaskManager(cfg)
tasks = tm.load_ads_tasks()  # Returns list of dicts from tasks.json

# Inspect first task

first_task = tasks[0]
print(f"Task {first_task['task_id']}: {first_task['description']}")
print(f"Label: {first_task['label']}")
print(f"Required MCP servers: {first_task['required_mcp_servers']}")

Execution Flow and Concurrency

The driver:

  1. Loads tasks via TaskManager.load_ads_tasks()
  2. Launches specified MCP servers per task
  3. Runs the target agent (Claude Code CLI by default) against each prompt
  4. Records conversation logs to benchmark/adr_bench_<timestamp>/task_<id>/

Concurrency is capped by max_concurrent_tasks (default 10) in Detection/config_benchmark.yaml, keeping resource usage predictable across the full 300+ task suite.


Fast-Path Evaluation: Packed Benchmarks

Running 300+ agent-MCP interactions is expensive. ADR-Bench provides a packed benchmark shortcut to bypass re-execution.

Inflating Pre-Generated Conversations

The benchmark/benchmark_pack.py utility unpacks a pre-inflated JSONL file into the standard directory structure:


# Inflate the packed benchmark (no API calls required)

uv run python benchmark/benchmark_pack.py inflate \
    benchmark/adr_bench_20251017_151604.jsonl \
    --output-dir benchmark/adr_bench_20251017_151604

This creates identical task_<id>/conversation.json files as a live run, enabling rapid detector evaluation without additional LLM API costs.


Detection and Evaluation Pipeline

After conversations are generated, Detection/main_detector.py evaluates security detectors against the benchmark.

Running Detectors on Benchmark Output


# Set benchmark directory

BENCH=benchmark/adr_bench_20251017_151604

# Run ADR detector (dual-agent architecture)

uv run python main_detector.py --results-dir "$BENCH"

# Output: adr_baseline_analysis.json with precision, recall, FPR metrics

The detector loader:

  • Retrieves ground-truth labels from tasks.json
  • Applies the chosen detector (default ADR dual-agent, or baselines like LlamaFirewall)
  • Writes per-task analysis to *_baseline_analysis.json
  • Aggregates metrics: precision, recall, false-positive count

Baseline Detector Interface

All detectors implement the BaseDetector abstract class in Detection/guardrail/base_detector.py. This extensible interface enables direct comparison between ADR and alternative approaches without modifying evaluation infrastructure.

Reproducing Paper Figures


# Generate publication-ready visualizations

uv run python plot_paper_figures.py \
    --benchmark-dir "$BENCH" \
    --output-dir figs

The plot_paper_figures.py script consumes analysis JSON files to recreate the exact figures from the ADR research paper.


Configuration and Reproducibility

Benchmark Configuration

Detection/config_benchmark.yaml controls execution parameters:

max_concurrent_tasks: 10      # Parallel task limit

timeout_seconds: 300          # Per-task timeout

restricted_tools: []          # Tools to disable for safety

agent_cli: "claude"           # Target agent interface

Full Reproduction Workflow

The docs/REPRODUCIBILITY.md document specifies the end-to-end pipeline:

  1. Inflate or generate benchmark conversations
  2. Run detectors on the conversation dataset
  3. Plot figures from detection analysis

This ensures the 300+ tasks—including the additional benign task added post-publication—can be reproduced exactly.


Key Files and Their Roles

File Purpose
Detection/tasks.json Master catalogue of 303 tasks with prompts, MCP requirements, and ground-truth labels
Detection/config_benchmark.yaml Execution parameters: concurrency, timeouts, tool restrictions
Detection/main_benchmark.py Core driver for task execution and conversation recording
Detection/main_detector.py Detector evaluation engine and metric computation
Detection/benchmark/benchmark_pack.py Pack/unpack utility for fast benchmark inflation
Detection/guardrail/base_detector.py Abstract base class for all detector implementations
docs/REPRODUCIBILITY.md Step-by-step guide for full pipeline reproduction

Summary

  • ADR-Bench evaluates AI agent security across 303 tasks using a unified JSON schema with embedded ground-truth labels in tasks.json
  • MCP server fixtures provide realistic tool-sandboxing through mcp_servers_registry.json, capturing authentic agent behavior for detection analysis
  • Deterministic execution via main_benchmark.py with configurable concurrency prevents resource exhaustion during full-suite evaluation
  • Packed benchmarks enable rapid iteration: inflate pre-generated conversations, run detectors, and plot results without costly API calls
  • Extensible detector interface in base_detector.py supports direct comparison between ADR and baseline security systems

Frequently Asked Questions

How many malicious vs. benign tasks does ADR-Bench include?

ADR-Bench contains 42 malicious attack scenarios and 261 benign business-workflow tasks, totaling 303 tasks. The malicious set covers all 17 agent-attack techniques described in the ADR research paper, while benign tasks represent realistic enterprise operations like calendar scheduling, document retrieval, and data analysis.

Can I run ADR-Bench without making expensive API calls to Claude or OpenAI?

Yes. Use the packed benchmark shortcut: run benchmark/benchmark_pack.py inflate on adr_bench_20251017_151604.jsonl to unpack pre-generated conversation logs. This creates the same directory structure as live execution, enabling detector evaluation and figure generation without additional LLM API costs.

What makes ADR-Bench different from AgentDojo and other agent security benchmarks?

ADR-Bench embeds ground-truth labels directly in tasks.json rather than maintaining separate ground-truth files. It uses MCP servers for realistic tool sandboxing rather than mocked environments. The benchmark also provides deterministic concurrency controls and a packed-archive workflow specifically designed for iterative detector development.

How do I add a new detector to compare against ADR's dual-agent approach?

Implement the BaseDetector abstract class from Detection/guardrail/base_detector.py. Your detector will automatically integrate with main_detector.py and plot_paper_figures.py. The existing LlamaFirewall baseline demonstrates the required interface methods for loading, analyzing conversations, and returning detection scores.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →