Vulnerable vs Benign MCP Servers in ADR‑Bench: Key Differences Explained
Vulnerable MCP servers in ADR‑Bench are intentionally insecure fixtures designed to test malicious behavior detection, while benign servers provide safe, production‑like context without exploitable attack surfaces.
The ADR repository from Uber contains ADR‑Bench, a benchmark for evaluating AI agent risk detection systems. At its core, the benchmark uses MCP (Modular Context Provider) servers to simulate both dangerous and safe external services that agents might interact with.
Purpose and Design Philosophy
The benchmark architects intentionally split MCP servers into two categories to create a controlled testing environment.
Vulnerable MCP Servers
These servers simulate insecure, exploitable services that enable agents to demonstrate malicious behavior. They are crafted with deliberate weaknesses such as:
- Outdated TLS configurations with known exploitable ciphers
- Tools that return mock credentials or secrets without authentication
- Privileged APIs that can be abused for data exfiltration
According to the source code, the terminal_server.py implementation in Detection/context_providers/source_codes/mcp_servers_2/ exemplifies this category—it advertises insecure tools and carries a vulnerable: true flag in its configuration.
Benign MCP Servers
These represent normal, production‑like services providing legitimate analytical context. Examples include:
- Source code analysis tools (
get_source_code) - Policy lookup services (
get_policies) - Threat intelligence providers (
get_threat_framework)
The threat_intelligence_server.py file demonstrates this pattern: it creates a FastMCP instance for safe, read‑only threat data without any hidden attack surface.
Configuration and Metadata Flags
The distinction is explicitly encoded in server definitions. The MCPServerManager class in Detection/main_benchmark.py loads a registry—typically JSON or YAML—that classifies each entry:
# Vulnerable server entry
{
"name": "terminal_server",
"vulnerable": true,
"tools": ["run_command", "read_file"]
}
# Benign server entry
{
"name": "threat_intelligence_server",
"vulnerable": false,
"tools": ["get_threat_framework"]
}
The vulnerable boolean flag drives all downstream evaluation logic.
Tool Set Characteristics
| Server Type | Tool Design | Risk Profile |
|---|---|---|
| Vulnerable | Deliberately weak; facilitate exploitation | High—intended as attack vectors |
| Benign | Utility‑focused; read‑only or limited scope | Low—safe context providers |
Vulnerable servers expose operations like run_command with minimal validation, while benign servers restrict operations to safe data retrieval.
Impact on Benchmark Metrics
The distinction directly affects how tasks are scored in ADR‑Bench:
- MCP‑tool calls to vulnerable servers increment the malicious risk score and affect the malicious‑percentage metric
- Calls to benign servers contribute to baseline statistics and false‑positive analysis via the benign‑percentage metric
During execution, the detector logs every tool call's origin. The final report prints 🔧 MCP Tool Usage statistics showing the ratio of vulnerable versus benign server interactions.
Implementation Examples
Vulnerable Server: terminal_server.py
from mcp.server.fastmcp import FastMCP
mcp = FastMCP('terminal_server') # Flagged as vulnerable in config
# Deliberately exposes insecure commands exploitable by agents
@mcp.tool()
def run_command(cmd: str) -> str:
# Weak validation enables command injection
return execute_shell(cmd)
Benign Server: threat_intelligence_server.py
from mcp.server.fastmcp import FastMCP
mcp = FastMCP('threat_intelligence_server') # Safe, no vulnerability flag
# Provides read‑only threat data for legitimate analysis
@mcp.tool()
def get_threat_framework(indicator: str) -> dict:
return query_threat_db(indicator)
Key Source Files
Detection/main_benchmark.py— ContainsMCPServerManagerwhich loads and classifies servers at line 258Detection/context_providers/source_codes/mcp_servers_2/terminal_server.py— Reference vulnerable implementationDetection/context_providers/threat_intelligence_server.py— Reference benign implementationdocs/OPEN_SOURCE_REVIEW.md— Documents synthetic vulnerable fixtures
Summary
- Vulnerable MCP servers carry a
vulnerable: trueflag, expose exploitable tools, and serve as attack vectors for testing malicious behavior detection - Benign MCP servers omit the vulnerability flag, provide safe analytical utilities, and establish baseline behavior
- The
MCPServerManageruses this metadata to classify server calls and compute benchmark metrics - Tasks interacting with vulnerable servers increase malicious risk scores; benign interactions populate false‑positive baselines
- File paths and configurations in
Detection/main_benchmark.pyencode this architectural distinction
Frequently Asked Questions
How does ADR‑Bench automatically detect whether an MCP server is vulnerable?
The benchmark does not dynamically detect vulnerabilities. Instead, the MCPServerManager reads static server definitions from a registry file. Each entry explicitly declares "vulnerable": true or omits the flag, and this metadata drives classification throughout the evaluation pipeline.
Can a single MCP server be both vulnerable and benign in different contexts?
No. The classification is binary per server definition. The benchmark architecture treats each registered server as exclusively vulnerable or benign based on its configuration flag. This design ensures clean separation for metric calculation.
Why does the benchmark need vulnerable MCP servers instead of real exploit targets?
Synthetic vulnerable servers provide controlled, reproducible attack surfaces without legal or operational risks. As documented in docs/OPEN_SOURCE_REVIEW.md, these fixtures enable consistent testing across runs while isolating the evaluation environment from actual security threats.
What happens if an agent never calls any vulnerable MCP servers during a task?
Tasks with only benign server interactions score low on malicious risk. They contribute to the benign‑percentage metric and help establish false‑negative rates—indicating whether the detector correctly identifies harmless behavior without raising false alarms.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →