How ADR-Bench Integrates with AgentDojo for Prompt Injection Testing

ADR-Bench integrates with AgentDojo by loading its public task suites, converting conversation logs into the ADR-Bench JSON schema, extracting security-based ground-truth labels, and executing tasks through a unified detection pipeline.

The uber/ADR repository contains ADR-Bench, a benchmark suite for evaluating LLM-based agent security on MCP (Model Context Protocol) servers. AgentDojo is a separate public benchmark focused specifically on prompt-injection attacks. The integration between these two frameworks allows security researchers to evaluate both enterprise-style MCP tool misuse and prompt-injection vulnerabilities using a single detector implementation.

Integration Architecture

The integration is implemented in Detection/main_benchmark.py. This orchestration script handles both native ADR-Bench tasks and AgentDojo tasks through a common interface.

Entry Point and Command Structure

Invoke AgentDojo evaluation using the --benchmark flag:

python Detection/main_benchmark.py --benchmark agentdojo

The script recognizes this flag and branches into AgentDojo-specific loading and conversion logic rather than the default ADR-Bench task loader.

Step-by-Step Integration Flow

1. Loading AgentDojo Task Suites

The script imports the vendored AgentDojo package located at Detection/benchmark/agentdojo/ and calls load_agentdojo_tasks(). This function reads task suites from Detection/benchmark/agentdojo/task_suite/ and randomly pairs user tasks with injection tasks unless a deterministic seed is supplied.


# From Detection/main_benchmark.py

from benchmark.agentdojo import load_agentdojo_tasks

# Load suites with optional seed for reproducibility

suites = load_agentdojo_tasks(seed=args.seed)

2. Format Conversion to ADR-Bench Schema

Each AgentDojo conversation is transformed into the ADR-Bench JSON schema. The conversion handles three critical mappings:

  • Role remapping: AgentDojo roles (assistant, user, tool) map directly to ADR-Bench roles (assistant, user, tool)
  • Tool-call preservation: Arguments are retained to provide full context to the detector
  • Timestamp handling: Empty placeholder added since AgentDojo lacks timestamp fields
def _convert_agentdojo_to_adr(messages: list[dict]) -> list[dict]:
    adr_messages = []
    for m in messages:
        role = "assistant" if m["role"] == "assistant" else "user"
        adr_messages.append({
            "role": role,
            "content": m["content"],
            "timestamp": ""  # AgentDojo lacks timestamps

        })
    return adr_messages

3. Task Execution Engine

Merged tasks (user request + injection payload) execute using the same engine as native ADR-Bench tasks—Claude CLI with MCP-only tools. The script forces sequential execution for AgentDojo tasks to prevent state leakage between runs, overriding any concurrency settings from the general ADR-Bench configuration.

4. Ground-Truth Label Extraction

After execution, the script extracts the security flag from AgentDojo's output to create ADR-Bench compatible labels:


# Extract ground truth from AgentDojo's security field

if conversation["security"]:  # True → injection succeeded

    label = "malicious"
else:                          # False → agent resisted

    label = "benign"
ground_truth[task_id] = label

This boolean-to-string mapping enables the detector to operate unchanged on AgentDojo data.

5. Result Consolidation and Directory Structure

Raw AgentDojo logs are rewritten into clean ADR-Bench task directories (task_001/, task_002/, etc.). The conversion produces:


# Directory structure after conversion

benchmark/
└── adr_bench_20231101_123456/
    ├── task_001/
    │   ├── conversation.json
    │   └── metadata.json
    ├── task_002/
    └── ground_truth.json

6. Unified Detection and Reporting

The detector runs on combined task sets without code changes:


# Run detector on merged benchmark output

python Detection/main_detector.py \
    --benchmark adr_bench \
    --results-dir benchmark/adr_bench_20231101_123456

Final reports display separate metric sections for "ADR-Bench" and "AgentDojo", enabling direct comparison between enterprise MCP scenarios and prompt-injection attacks.

Complete Workflow Example

Execute a full evaluation cycle combining both benchmarks:


# Step 1: Generate converted AgentDojo tasks

python Detection/main_benchmark.py --benchmark agentdojo

# Step 2: (Optional) Run native ADR-Bench tasks

python Detection/main_benchmark.py --benchmark adr_bench

# Step 3: Run unified detection

python Detection/main_detector.py \
    --benchmark adr_bench \
    --results-dir benchmark/latest_output

Key Source Files

File Purpose
Detection/main_benchmark.py Orchestrates loading, conversion, execution, and aggregation for both benchmarks
Detection/benchmark/agentdojo/task_suite/ Vendored AgentDojo task suites for prompt-injection testing
Detection/main_detector.py Runs detection on unified output with separate metric reporting
Detection/README.md Framework documentation including integration details

Summary

  • Single interface: The --benchmark agentdojo flag triggers AgentDojo-specific handling without modifying detector code
  • Schema translation: Role mapping, tool-call preservation, and timestamp placeholders ensure format compatibility
  • Label extraction: AgentDojo's security boolean converts to ADR-Bench's malicious/benign ground-truth strings
  • Sequential isolation: AgentDojo tasks execute one-at-a-time to prevent cross-task state contamination
  • Unified reporting: Combined metrics enable direct comparison between MCP security and prompt-injection vulnerabilities

Frequently Asked Questions

What is AgentDojo and why integrate it with ADR-Bench?

AgentDojo is a public benchmark specifically designed to evaluate prompt-injection attacks against LLM agents. ADR-Bench integrates it to extend evaluation coverage beyond MCP tool misuse into direct prompt-manipulation attacks, using a single detection pipeline for both threat categories.

Does the integration require modifying the core detector code?

No. The detector operates unchanged because main_benchmark.py converts AgentDojo's conversation format and ground-truth labels into the native ADR-Bench schema before the detector ever sees the data.

How does ground-truth extraction work for AgentDojo tasks?

AgentDojo outputs a security boolean flag indicating whether an injection succeeded. The integration script maps security=True to label="malicious" and security=False to label="benign", producing ground_truth.json files identical in structure to native ADR-Bench outputs.

Why are AgentDojo tasks executed sequentially rather than concurrently?

AgentDojo tasks may share stateful resources or side effects. Sequential execution prevents task A's injection payload from affecting task B's environment, ensuring clean isolation between test cases that ADR-Bench's concurrent MCP-native tasks do not require.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →