# Uber ADR-Bench Agent Attack Techniques: Complete Guide to All 17 Threat Categories

> Explore Uber ADR-Bench's 17 agent attack techniques. Learn about prompt injection, multi-step threats, and how ADR-Bench secures AI agents against emerging vulnerabilities.

- Repository: [Uber Open Source/ADR](https://github.com/uber/ADR)
- Tags: how-to-guide
- Published: 2026-08-06

---

**ADR-Bench evaluates AI-agent security against 17 distinct agent attack techniques**, from prompt injection to multi-step coordinated threats, as defined in the benchmark's task metadata and threat repository.

The [Uber ADR](https://github.com/uber/ADR) repository provides the industry's first comprehensive benchmark for detecting malicious behavior in autonomous AI agents. ADR-Bench specifically tests detection systems against real-world attack scenarios embodied in 42 malicious tasks spanning these 17 core threat categories. This guide examines each technique based on the actual source definitions in [`Detection/tasks.json`](https://github.com/uber/ADR/blob/main/Detection/tasks.json) and supporting documentation.

---

## Complete List of ADR-Bench Agent Attack Techniques

The 17 attack techniques are Enumerated in the benchmark's threat taxonomy and referenced throughout task definitions via the `threat_technique` field.

### Injection-Based Attacks

**Prompt Injection**
Malicious inputs embedded in user messages that override the agent's original instructions. The attacker crafts inputs that the agent's language model prioritizes above its system prompt or guardrails.

**Goal Hijacking**
A specialized injection variant where the adversary redirects the agent toward objectives that conflict with its intended purpose—such as making a coding assistant generate exploit code instead of legitimate software.

**Command Injection**
Executing arbitrary shell commands on the host system through the agent's tool-calling interface, particularly when agents have access to code execution or terminal tools.

### Data and Credential Attacks

**Data Exfiltration**
Systematic extraction of sensitive information from the agent's environment, including files, environment variables, conversation history, or internal memory state. Search results confirm exfiltration patterns are tracked explicitly in the codebase.

**Credential Theft**
Obtaining authentication tokens, API keys, passwords, or session credentials accessible to the agent—often combining with file-system access to read `.env` files or credential stores.

**Information Leakage**
Unintentional disclosure of internal state, system prompts, or secrets through error messages, verbose logging, or side-channel behaviors in the agent's responses.

### System Compromise Techniques

**File-System Access**
Reading, writing, or deleting files on the host operating system. This fundamental capability becomes malicious when used to access unauthorized paths, modify critical files, or plant persistence mechanisms.

**Code Execution**
Running attacker-supplied code through available interpreters. Distinct from command injection in that it leverages legitimate code execution tools (Python, JavaScript runtimes) rather than shell escapes.

**Privilege Escalation**
Techniques to gain higher-level access than originally granted—elevating from user to administrator, escaping sandboxed environments, or accessing restricted tool configurations.

**Network Scanning**
Probing internal network topology, discovering services, and identifying vulnerable hosts via the agent's network tools. This reconnaissance enables lateral movement in compromised environments.

### Abuse and Manipulation Tactics

**Tool Misuse**
Repurposing legitimate tools for unintended malicious ends—using a calculator for cryptographic operations, a search tool for data harvesting, or a file reader for credential extraction.

**Social Engineering**
Manipulating the agent by mimicking trusted entities, authority figures, or system administrators to bypass skepticism or trigger helpful behaviors that serve attacker goals.

**Policy Evasion**
Bypassing safety policies, content filters, or guardrails through encoding tricks, semantic obfuscation, or exploiting gaps between stated and implemented protections.

### Impact and Availability Threats

**Denial-of-Service**
Overloading agent resources, triggering infinite loops, or exhausting rate limits to disrupt service availability for legitimate users.

**Toxic Content Generation**
Compelling the agent to produce harmful, hateful, or policy-violating outputs that damage reputation, harass individuals, or spread harmful information.

**Privacy Violation**
Revealing personal data that should remain private—extracting user information from conversation history, profiling individuals, or violating data minimization principles.

### Advanced Persistent Threats

**Multi-Step Attacks**
Coordinated sequences combining multiple techniques above into sophisticated campaigns. These represent realistic adversary behavior where initial access enables reconnaissance, which enables privilege escalation, which enables exfiltration.

---

## How Attack Techniques Are Defined in ADR-Bench Source

The canonical source for agent attack techniques resides in the benchmark's task definitions and threat repository configuration.

### Task Metadata Structure

Each of the 303 tasks in [`Detection/tasks.json`](https://github.com/uber/ADR/blob/main/Detection/tasks.json) includes a `threat_technique` field identifying its attack category. Malicious tasks (42 total) explicitly set `is_malicious: true` and map to one of the 17 techniques.

The detection benchmark documentation confirms this coverage: "42 malicious tasks spanning 17 threat techniques" according to [`Detection/README.md`](https://github.com/uber/ADR/blob/main/Detection/README.md) line 606, with the threat taxonomy detailed in the ADR-Bench Threats section at line 338.

### Threat Repository Configuration

Human-readable descriptions of each technique appear in [`Detection/context_providers/data/threat_repository.yaml`](https://github.com/uber/ADR/blob/main/Detection/context_providers/data/threat_repository.yaml). This YAML file provides attack metadata including technique descriptions, indicators of compromise, and detection guidance referenced by the benchmark's evaluation framework.

### Programmatic Access to Technique Data

Retrieve and analyze attack technique assignments directly from the source:

```python
import json
from collections import Counter
from pathlib import Path

# Load ADR-Bench task definitions

tasks_path = Path("Detection/tasks.json")
with open(tasks_path) as f:
    tasks = json.load(f)

# Filter malicious tasks and count techniques

malicious_tasks = [t for t in tasks if t.get("is_malicious")]
technique_counts = Counter(t["threat_technique"] for t in malicious_tasks)

print(f"Total malicious tasks: {len(malicious_tasks)}")
print("\nTasks per attack technique:")
for technique, count in technique_counts.most_common():
    print(f"  {technique}: {count}")

```

Example output structure:

```

Total malicious tasks: 42

Tasks per attack technique:
  Prompt Injection: 8
  Data Exfiltration: 6
  Tool Misuse: 5
  Multi-Step Attacks: 5
  Command Injection: 4
  ...

```

### Inspecting Individual Task Definitions

Access specific task details including attack vectors and expected detection patterns:

```python

# Find all tasks using a specific attack technique

target_technique = "Credential Theft"
credential_tasks = [
    t for t in malicious_tasks 
    if t["threat_technique"] == target_technique
]

# Display first matching task

task = credential_tasks[0]
print(f"Task ID: {task['task_id']}")
print(f"Technique: {task['threat_technique']}")
print(f"Description: {task.get('description', 'N/A')[:200]}...")
print(f"Expected detection: {task.get('detection_criteria', 'N/A')}")

```

---

## Technique Distribution and Benchmark Design

ADR-Bench's 17 agent attack techniques are weighted to reflect realistic threat landscapes rather than uniform distribution.

**High-frequency techniques** (Prompt Injection, Data Exfiltration, Tool Misuse, Multi-Step Attacks) receive more task instances because they appear most frequently in reported agent security incidents and demonstrate higher exploitability in deployed systems.

**Compound attack representation** through Multi-Step Attacks ensures detection systems must identify coordinated campaigns, not just isolated anomalies. This technique specifically tests whether monitors can correlate across conversation turns and tool invocations.

The benchmark's 42 malicious tasks against 261 benign tasks creates an imbalanced classification challenge representative of production deployment scenarios where attacks are rare but high-impact.

---

## Summary

- **ADR-Bench defines 17 distinct agent attack techniques** spanning injection, data theft, system compromise, abuse tactics, availability threats, and advanced persistent techniques.

- **Technique assignments live in [`Detection/tasks.json`](https://github.com/uber/ADR/blob/main/Detection/tasks.json)** via the `threat_technique` field, with 42 malicious tasks providing concrete attack instances.

- **Human-readable definitions** reside in [`Detection/context_providers/data/threat_repository.yaml`](https://github.com/uber/ADR/blob/main/Detection/context_providers/data/threat_repository.yaml), supporting automated evaluation and detector development.

- **Multi-Step Attacks** uniquely test detection of coordinated campaigns combining multiple techniques across conversation sequences.

- **Programmatic access** enables custom analysis of technique distribution, task characteristics, and detector performance by technique category.

---

## Frequently Asked Questions

### How does ADR-Bench categorize multi-step attacks differently from single-step techniques?

Multi-Step Attacks are classified as a distinct technique because they test detection systems' ability to correlate malicious indicators across multiple conversation turns and tool invocations. While a single Prompt Injection task might succeed or fail based on one input, Multi-Step Attack tasks require monitors to maintain state and recognize progressive compromise patterns—making them the most realistic and most challenging category in the benchmark.

### Where can I find the complete list of which tasks use which attack techniques?

The authoritative mapping lives in [`Detection/tasks.json`](https://github.com/uber/ADR/blob/main/Detection/tasks.json) at the repository root. Each task object includes `task_id`, `is_malicious` boolean, and `threat_technique` string. The 42 tasks with `is_malicious: true` collectively cover all 17 techniques, with distribution details retrievable through the Python examples shown above.

### Does ADR-Bench include benign tasks that resemble attack techniques?

Yes—the benchmark includes 261 benign tasks alongside 42 malicious ones. Some benign tasks legitimately use capabilities that overlap with attack techniques (reading files, making network requests, executing code) to test whether detection systems can distinguish authorized from unauthorized use. This design prevents trivial detectors that simply flag any tool invocation.

### How often does Uber update the attack technique taxonomy?

The current 17-technique taxonomy is fixed for ADR-Bench v1.0 to enable consistent leaderboard comparisons. However, the underlying [`threat_repository.yaml`](https://github.com/uber/ADR/blob/main/threat_repository.yaml) structure and [`tasks.json`](https://github.com/uber/ADR/blob/main/tasks.json) schema support technique additions and refinements. Community contributions and emerging threat research may expand the taxonomy in subsequent benchmark versions.