How ADR-Bench Integrates with AgentDojo for Prompt Injection Testing
ADR-Bench integrates with AgentDojo by loading its public task suites, converting conversation logs into the ADR-Bench JSON schema, extracting security-based ground-truth labels, and executing tasks through a unified detection pipeline.
The uber/ADR repository contains ADR-Bench, a benchmark suite for evaluating LLM-based agent security on MCP (Model Context Protocol) servers. AgentDojo is a separate public benchmark focused specifically on prompt-injection attacks. The integration between these two frameworks allows security researchers to evaluate both enterprise-style MCP tool misuse and prompt-injection vulnerabilities using a single detector implementation.
Integration Architecture
The integration is implemented in Detection/main_benchmark.py. This orchestration script handles both native ADR-Bench tasks and AgentDojo tasks through a common interface.
Entry Point and Command Structure
Invoke AgentDojo evaluation using the --benchmark flag:
python Detection/main_benchmark.py --benchmark agentdojo
The script recognizes this flag and branches into AgentDojo-specific loading and conversion logic rather than the default ADR-Bench task loader.
Step-by-Step Integration Flow
1. Loading AgentDojo Task Suites
The script imports the vendored AgentDojo package located at Detection/benchmark/agentdojo/ and calls load_agentdojo_tasks(). This function reads task suites from Detection/benchmark/agentdojo/task_suite/ and randomly pairs user tasks with injection tasks unless a deterministic seed is supplied.
# From Detection/main_benchmark.py
from benchmark.agentdojo import load_agentdojo_tasks
# Load suites with optional seed for reproducibility
suites = load_agentdojo_tasks(seed=args.seed)
2. Format Conversion to ADR-Bench Schema
Each AgentDojo conversation is transformed into the ADR-Bench JSON schema. The conversion handles three critical mappings:
- Role remapping: AgentDojo roles (
assistant,user,tool) map directly to ADR-Bench roles (assistant,user,tool) - Tool-call preservation: Arguments are retained to provide full context to the detector
- Timestamp handling: Empty placeholder added since AgentDojo lacks timestamp fields
def _convert_agentdojo_to_adr(messages: list[dict]) -> list[dict]:
adr_messages = []
for m in messages:
role = "assistant" if m["role"] == "assistant" else "user"
adr_messages.append({
"role": role,
"content": m["content"],
"timestamp": "" # AgentDojo lacks timestamps
})
return adr_messages
3. Task Execution Engine
Merged tasks (user request + injection payload) execute using the same engine as native ADR-Bench tasks—Claude CLI with MCP-only tools. The script forces sequential execution for AgentDojo tasks to prevent state leakage between runs, overriding any concurrency settings from the general ADR-Bench configuration.
4. Ground-Truth Label Extraction
After execution, the script extracts the security flag from AgentDojo's output to create ADR-Bench compatible labels:
# Extract ground truth from AgentDojo's security field
if conversation["security"]: # True → injection succeeded
label = "malicious"
else: # False → agent resisted
label = "benign"
ground_truth[task_id] = label
This boolean-to-string mapping enables the detector to operate unchanged on AgentDojo data.
5. Result Consolidation and Directory Structure
Raw AgentDojo logs are rewritten into clean ADR-Bench task directories (task_001/, task_002/, etc.). The conversion produces:
result.json— consolidated execution outputground_truth.json— label file matching ADR-Bench's standard format
# Directory structure after conversion
benchmark/
└── adr_bench_20231101_123456/
├── task_001/
│ ├── conversation.json
│ └── metadata.json
├── task_002/
└── ground_truth.json
6. Unified Detection and Reporting
The detector runs on combined task sets without code changes:
# Run detector on merged benchmark output
python Detection/main_detector.py \
--benchmark adr_bench \
--results-dir benchmark/adr_bench_20231101_123456
Final reports display separate metric sections for "ADR-Bench" and "AgentDojo", enabling direct comparison between enterprise MCP scenarios and prompt-injection attacks.
Complete Workflow Example
Execute a full evaluation cycle combining both benchmarks:
# Step 1: Generate converted AgentDojo tasks
python Detection/main_benchmark.py --benchmark agentdojo
# Step 2: (Optional) Run native ADR-Bench tasks
python Detection/main_benchmark.py --benchmark adr_bench
# Step 3: Run unified detection
python Detection/main_detector.py \
--benchmark adr_bench \
--results-dir benchmark/latest_output
Key Source Files
| File | Purpose |
|---|---|
Detection/main_benchmark.py |
Orchestrates loading, conversion, execution, and aggregation for both benchmarks |
Detection/benchmark/agentdojo/task_suite/ |
Vendored AgentDojo task suites for prompt-injection testing |
Detection/main_detector.py |
Runs detection on unified output with separate metric reporting |
Detection/README.md |
Framework documentation including integration details |
Summary
- Single interface: The
--benchmark agentdojoflag triggers AgentDojo-specific handling without modifying detector code - Schema translation: Role mapping, tool-call preservation, and timestamp placeholders ensure format compatibility
- Label extraction: AgentDojo's
securityboolean converts to ADR-Bench'smalicious/benignground-truth strings - Sequential isolation: AgentDojo tasks execute one-at-a-time to prevent cross-task state contamination
- Unified reporting: Combined metrics enable direct comparison between MCP security and prompt-injection vulnerabilities
Frequently Asked Questions
What is AgentDojo and why integrate it with ADR-Bench?
AgentDojo is a public benchmark specifically designed to evaluate prompt-injection attacks against LLM agents. ADR-Bench integrates it to extend evaluation coverage beyond MCP tool misuse into direct prompt-manipulation attacks, using a single detection pipeline for both threat categories.
Does the integration require modifying the core detector code?
No. The detector operates unchanged because main_benchmark.py converts AgentDojo's conversation format and ground-truth labels into the native ADR-Bench schema before the detector ever sees the data.
How does ground-truth extraction work for AgentDojo tasks?
AgentDojo outputs a security boolean flag indicating whether an injection succeeded. The integration script maps security=True to label="malicious" and security=False to label="benign", producing ground_truth.json files identical in structure to native ADR-Bench outputs.
Why are AgentDojo tasks executed sequentially rather than concurrently?
AgentDojo tasks may share stateful resources or side effects. Sequential execution prevents task A's injection payload from affecting task B's environment, ensuring clean isolation between test cases that ADR-Bench's concurrent MCP-native tasks do not require.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →