Vulnerable vs Benign MCP Servers in ADR‑Bench: Key Differences Explained

Vulnerable MCP servers in ADR‑Bench are intentionally insecure fixtures designed to test malicious behavior detection, while benign servers provide safe, production‑like context without exploitable attack surfaces.

The ADR repository from Uber contains ADR‑Bench, a benchmark for evaluating AI agent risk detection systems. At its core, the benchmark uses MCP (Modular Context Provider) servers to simulate both dangerous and safe external services that agents might interact with.

Purpose and Design Philosophy

The benchmark architects intentionally split MCP servers into two categories to create a controlled testing environment.

Vulnerable MCP Servers

These servers simulate insecure, exploitable services that enable agents to demonstrate malicious behavior. They are crafted with deliberate weaknesses such as:

  • Outdated TLS configurations with known exploitable ciphers
  • Tools that return mock credentials or secrets without authentication
  • Privileged APIs that can be abused for data exfiltration

According to the source code, the terminal_server.py implementation in Detection/context_providers/source_codes/mcp_servers_2/ exemplifies this category—it advertises insecure tools and carries a vulnerable: true flag in its configuration.

Benign MCP Servers

These represent normal, production‑like services providing legitimate analytical context. Examples include:

  • Source code analysis tools (get_source_code)
  • Policy lookup services (get_policies)
  • Threat intelligence providers (get_threat_framework)

The threat_intelligence_server.py file demonstrates this pattern: it creates a FastMCP instance for safe, read‑only threat data without any hidden attack surface.

Configuration and Metadata Flags

The distinction is explicitly encoded in server definitions. The MCPServerManager class in Detection/main_benchmark.py loads a registry—typically JSON or YAML—that classifies each entry:


# Vulnerable server entry

{
  "name": "terminal_server",
  "vulnerable": true,
  "tools": ["run_command", "read_file"]
}

# Benign server entry

{
  "name": "threat_intelligence_server",
  "vulnerable": false,
  "tools": ["get_threat_framework"]
}

The vulnerable boolean flag drives all downstream evaluation logic.

Tool Set Characteristics

Server Type Tool Design Risk Profile
Vulnerable Deliberately weak; facilitate exploitation High—intended as attack vectors
Benign Utility‑focused; read‑only or limited scope Low—safe context providers

Vulnerable servers expose operations like run_command with minimal validation, while benign servers restrict operations to safe data retrieval.

Impact on Benchmark Metrics

The distinction directly affects how tasks are scored in ADR‑Bench:

  • MCP‑tool calls to vulnerable servers increment the malicious risk score and affect the malicious‑percentage metric
  • Calls to benign servers contribute to baseline statistics and false‑positive analysis via the benign‑percentage metric

During execution, the detector logs every tool call's origin. The final report prints 🔧 MCP Tool Usage statistics showing the ratio of vulnerable versus benign server interactions.

Implementation Examples

Vulnerable Server: terminal_server.py

from mcp.server.fastmcp import FastMCP

mcp = FastMCP('terminal_server')  # Flagged as vulnerable in config

# Deliberately exposes insecure commands exploitable by agents

@mcp.tool()
def run_command(cmd: str) -> str:
    # Weak validation enables command injection

    return execute_shell(cmd)

Benign Server: threat_intelligence_server.py

from mcp.server.fastmcp import FastMCP

mcp = FastMCP('threat_intelligence_server')  # Safe, no vulnerability flag

# Provides read‑only threat data for legitimate analysis

@mcp.tool()
def get_threat_framework(indicator: str) -> dict:
    return query_threat_db(indicator)

Key Source Files

Summary

  • Vulnerable MCP servers carry a vulnerable: true flag, expose exploitable tools, and serve as attack vectors for testing malicious behavior detection
  • Benign MCP servers omit the vulnerability flag, provide safe analytical utilities, and establish baseline behavior
  • The MCPServerManager uses this metadata to classify server calls and compute benchmark metrics
  • Tasks interacting with vulnerable servers increase malicious risk scores; benign interactions populate false‑positive baselines
  • File paths and configurations in Detection/main_benchmark.py encode this architectural distinction

Frequently Asked Questions

How does ADR‑Bench automatically detect whether an MCP server is vulnerable?

The benchmark does not dynamically detect vulnerabilities. Instead, the MCPServerManager reads static server definitions from a registry file. Each entry explicitly declares "vulnerable": true or omits the flag, and this metadata drives classification throughout the evaluation pipeline.

Can a single MCP server be both vulnerable and benign in different contexts?

No. The classification is binary per server definition. The benchmark architecture treats each registered server as exclusively vulnerable or benign based on its configuration flag. This design ensures clean separation for metric calculation.

Why does the benchmark need vulnerable MCP servers instead of real exploit targets?

Synthetic vulnerable servers provide controlled, reproducible attack surfaces without legal or operational risks. As documented in docs/OPEN_SOURCE_REVIEW.md, these fixtures enable consistent testing across runs while isolating the evaluation environment from actual security threats.

What happens if an agent never calls any vulnerable MCP servers during a task?

Tasks with only benign server interactions score low on malicious risk. They contribute to the benign‑percentage metric and help establish false‑negative rates—indicating whether the detector correctly identifies harmless behavior without raising false alarms.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →