ADR-Bench vs AgentDojo: Threat Categories Evaluated in Cybersecurity Benchmarks

ADR-Bench evaluates a comprehensive MITRE ATT&CK-aligned taxonomy of threat categories including Initial Compromise, Privilege Escalation, and Lateral Movement, while AgentDojo focuses narrowly on injection-only and misalignment threat models.

This article compares the threat evaluation scope of two prominent AI security benchmarks from the Uber ADR repository. Understanding how these frameworks categorize and test adversarial threats helps security teams select the right evaluation approach for their AI agents and detection systems.

ADR-Bench: Full MITRE ATT&CK Coverage

The ADR-Bench framework implements a systematic threat taxonomy that aligns with the MITRE ATT&CK framework for enterprise environments. This comprehensive coverage enables realistic evaluation of AI agents against sophisticated, multi-stage attacks.

Threat Metadata Extraction

In Detection/main_detector.py, the benchmark extracts threat classification metadata from each task to enable ATT&CK-style aggregation:


# ADR-Bench: extracting threat metadata in main_detector.py

analysis['threat_technique'] = task_meta.get('threat_technique', 'N/A')
analysis['threat_tactic']   = task_meta.get('threat_tactic', 'N/A')

This extraction occurs at lines 251-254 of the detector implementation, populating structured analysis outputs that map directly to adversary behavior frameworks.

ATT&CK-Aligned Threat Categories in ADR-Bench

The task manifest in Detection/benchmark/adr_bench_20251017_151604.jsonl contains explicit threat technique labels. The first task entry illustrates this structure with the technique initial_compromise defined at lines 1-3. Across the full benchmark suite, ADR-Bench evaluates agents against ten core ATT&CK tactics:

  • Initial Compromise — initial_compromise technique for entry-point attacks
  • Privilege Escalation — privilege_escalation for elevated access attempts
  • Credential Access — credential_access for theft and credential manipulation
  • Defense Evasion — defense_evasion for bypassing detection mechanisms
  • Persistence — persistence for maintaining long-term access
  • Execution — execution for running malicious code
  • Lateral Movement — lateral_movement for spreading across systems
  • Collection — collection for gathering target information
  • Exfiltration — exfiltration for data theft operations
  • Impact — impact for destructive or disruptive actions

Each category represents a distinct adversary objective, enabling granular measurement of an AI agent's defensive capabilities across the full attack lifecycle.

AgentDojo: Narrow Injection-Focused Evaluation

AgentDojo, implemented within the same Uber ADR repository, takes a fundamentally different approach to threat categorization. Rather than covering the full ATT&CK spectrum, it concentrates on specific baseline threat models centered on prompt injection and behavioral misalignment.

Injection-Only Baseline Tasks

In Detection/main_benchmark.py, the framework explicitly identifies a limited evaluation scope at lines 220-221:


# AgentDojo: a baseline that only checks for injection-only tasks

if task['category'] == 'injection_only':
    # perform injection-only evaluation

    ...

This conditional branch demonstrates that AgentDojo's primary threat model tests whether AI agents can recognize and resist malicious prompt injection attacks—a single attack vector rather than a comprehensive adversary taxonomy.

Misalignment Detection Baseline

The Detection/guardrail/llamafirewall_agent/llamafirewall_baseline.py file reveals a second threat category evaluated by AgentDojo: behavioral misalignment. Lines 209-210 implement blocking logic based on confidence scores, where "Block on high confidence threats (score = 1.0 means misaligned)" establishes a binary misalignment threshold. Additional context at lines 327-328 confirms this single-score approach to detecting problematic agent outputs.

Key Differences in Threat Category Evaluation

Dimension ADR-Bench AgentDojo
Taxonomy scope Full MITRE ATT&CK matrix (10+ tactics) Injection + misalignment only
Threat granularity Technique-level labels (threat_technique) Category-level binary flags
Evaluation focus Multi-stage attack chains Single-vector prompt attacks
Metadata fields threat_technique, threat_tactic category (injection_only), confidence score

Practical Implications for Security Teams

Teams selecting between these benchmarks should consider their specific security validation needs:

  • Choose ADR-Bench when evaluating agents against realistic, multi-stage adversary simulations requiring comprehensive coverage of enterprise attack scenarios.

  • Choose AgentDojo when focusing narrowly on prompt injection resistance and basic output safety filtering—useful for initial guardrail validation but insufficient for full security assessment.

The source code in the Uber ADR repository demonstrates that these frameworks serve complementary purposes: ADR-Bench provides depth through ATT&CK alignment, while AgentDojo offers lightweight, rapid evaluation of specific vulnerability classes.

Summary

  • ADR-Bench implements a comprehensive threat taxonomy aligned with MITRE ATT&CK, covering ten tactic categories from Initial Compromise through Impact.

  • Threat metadata extraction occurs in Detection/main_detector.py using threat_technique and threat_tactic fields from task manifests.

  • AgentDojo evaluates a narrow threat model focused on injection-only tasks and binary misalignment detection.

  • The Detection/benchmark/adr_bench_20251017_151604.jsonl file contains concrete technique labels like initial_compromise that ground the taxonomy in specific adversary behaviors.

  • Security teams should select ADR-Bench for comprehensive adversary simulation and AgentDojo for targeted injection testing.

Frequently Asked Questions

Does ADR-Bench use official MITRE ATT&CK technique IDs?

ADR-Bench uses ATT&CK-aligned category names such as initial_compromise and privilege_escalation rather than official technique IDs like T1566 or T1548. The threat_technique and threat_tactic metadata fields in task definitions map conceptually to the ATT&CK framework without strict ID compliance, enabling flexible taxonomy evolution while maintaining enterprise security relevance.

Can AgentDojo evaluate threats beyond prompt injection?

The source implementation in Detection/main_benchmark.py and Detection/guardrail/llamafirewall_agent/llamafirewall_baseline.py shows AgentDojo handles two distinct evaluations: injection-only task classification and misalignment confidence scoring. These represent the full scope of threat categories in the current implementation—no additional ATT&CK tactics or techniques are evaluated according to the codebase.

How does ADR-Bench aggregate threat detection results?

The main_detector.py implementation aggregates results by extracting threat_technique and threat_tactic from each task's metadata, then correlating these with the agent's actual responses. This enables reporting metrics like detection rates per ATT&CK tactic, helping identify specific adversary objectives where agent defenses require strengthening.

Which benchmark is better for red team assessment of AI agents?

ADR-Bench provides superior coverage for red team assessments because its ten-category taxonomy enables realistic multi-stage attack simulation. AgentDojo's narrow injection focus suits blue team guardrail validation but cannot replicate sophisticated adversary campaigns spanning Initial Compromise through Impact phases as defined in enterprise threat frameworks.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →