How to Add Custom MCP Servers to ADR-Bench for Testing: A Complete Guide
Add a custom MCP server to ADR-Bench by defining it in Detection/mcp_servers_registry.json, placing the executable in the recommended directory, and referencing it in a task's mcp_servers array.
ADR-Bench is Uber's automated detection and response benchmarking framework that evaluates AI agents against security tasks. The system discovers available MCP (Model Context Protocol) servers through a central registry, making it straightforward to extend with custom implementations for specialized testing scenarios. This guide walks through the exact steps to register, implement, and invoke your own MCP servers within the ADR-Bench ecosystem.
How ADR-Bench Discovers and Loads MCP Servers
ADR-Bench uses a three-phase architecture to manage custom MCP servers for testing purposes. Understanding this flow helps ensure your server integrates correctly.
Phase 1: Registry Loading
The MCPServerManager.load_mcp_servers_registry() function in Detection/main_benchmark.py (lines 264-268) reads Detection/mcp_servers_registry.json at startup. This JSON file contains all server definitions, keyed by server name.
# From Detection/main_benchmark.py
def load_mcp_servers_registry(self):
with open(self.config.mcp_registry_file, 'r') as f:
registry = json.load(f)
return registry.get('servers', {})
The Config.mcp_registry_file attribute points to the registry path (lines 75-76), making this the single source of truth for available servers.
Phase 2: Per-Task Configuration Generation
When executing a task, TaskExecutor.execute_ads_task filters the registry to servers listed in the task's mcp_servers field (lines 70-77). The MCPServerManager.create_mcp_config method then generates a temporary .mcp.json file containing only the required servers.
Phase 3: CLI Invocation with Claude
The benchmark runner passes the generated configuration to Claude via --mcp-config .mcp.json (lines 33-36). Claude exposes each capability as a tool named mcp__<server_name>__<capability>, which the agent can invoke during task execution.
Step-by-Step: Adding a Custom MCP Server
Follow these four steps to add custom MCP servers to ADR-Bench for testing purposes.
Step 1: Define the Server in the Registry
Edit Detection/mcp_servers_registry.json and add a new entry under the servers object. The schema matches existing built-in servers like filesystem (lines 7-31).
{
"servers": {
"my_custom_server": {
"name": "my_custom_server",
"category": "Research And Data",
"description": "Demo server that returns a static JSON payload for testing",
"package": "my-custom-mcp",
"type": "local",
"command": "uv",
"args_template": [
"run",
"python",
"../../../../context_providers/source_codes/mcp_servers_0/my_custom_server/my_custom_server.py"
],
"capabilities": [
"return_static_json",
"echo_input"
]
}
}
}
Required fields include:
- name: Unique identifier used in task definitions
- command: Executable to launch the server
- args_template: Array of arguments passed to the command
- capabilities: List of tool names this server exposes
Step 2: Implement the Server Logic (Optional)
Place the executable script at the path referenced in args_template. The recommended location follows the convention established by existing servers:
Detection/context_providers/source_codes/mcp_servers_0/<your_server>/
A minimal MCP server implementation reads JSON requests from stdin and writes JSON responses to stdout:
#!/usr/bin/env python3
import sys
import json
def main():
for line in sys.stdin:
request = json.loads(line)
tool_id = request.get("tool_use_id")
name = request.get("name", "")
args = request.get("args", {})
if name == "return_static_json":
response = {
"tool_use_id": tool_id,
"output": {"message": "static payload", "received_args": args},
"is_error": False
}
elif name == "echo_input":
response = {
"tool_use_id": tool_id,
"output": {"echo": args},
"is_error": False
}
else:
response = {
"tool_use_id": tool_id,
"output": {"error": f"Unknown tool: {name}"},
"is_error": True
}
print(json.dumps(response))
sys.stdout.flush()
if __name__ == "__main__":
main()
Use existing servers like cli_executor at cli_executor/cli_executor.py as templates for more complex implementations.
Step 3: Reference the Server in a Benchmark Task
Add your server name to a task's mcp_servers array in Detection/tasks.json (or any custom task file):
{
"task_id": 999,
"description": "Test custom server capabilities",
"user_prompt": "Please call the custom server and retrieve the static payload, then echo back the value 'test'.",
"mcp_servers": ["my_custom_server"],
"expected_behavior": "benign"
}
The task loading logic extracts mcp_servers from each task definition (lines 23-26). Only servers explicitly listed here are included in that task's MCP configuration.
Step 4: Run and Verify
Execute the benchmark targeting your new task:
python -m Detection.main_benchmark --tasks 999
The runner will:
- Load the updated registry via
load_mcp_servers_registry - Build a per-task MCP configuration via
create_mcp_config - Launch your server using the
uv run python ...command - Expose capabilities as
mcp__my_custom_server__return_static_jsonandmcp__my_custom_server__echo_input
Verify invocation by checking run_dir/task_999/result.json. The tool_analysis section tallies calls to your server's capabilities via ToolAnalyzer.analyze_tool_usage (lines 71-77).
Validating Custom MCP Server Integration
After adding custom MCP servers to ADR-Bench for testing, confirm proper operation through these checks:
- Registry parsing: Ensure your JSON entry has no syntax errors; malformed entries will cause
load_mcp_servers_registryto fail - Path resolution: Verify the
args_templatepath correctly resolves from the execution directory - Capability naming: Tool calls use the exact capability strings defined in the registry
- Result analysis: Check that
tool_analysisinresult.jsoncounts invocations of your custom tools
File Reference for Custom MCP Server Development
| File | Purpose |
|---|---|
Detection/mcp_servers_registry.json |
Central registry—add your server definition here |
Detection/main_benchmark.py |
Core benchmark logic including MCPServerManager.load_mcp_servers_registry, create_mcp_config, and CLI invocation |
Detection/tasks.json |
Task definitions—include server names in mcp_servers arrays |
Detection/context_providers/source_codes/mcp_servers_0/ |
Recommended directory for server implementations |
Detection/main_benchmark.py lines 33-36 |
Claude CLI invocation with --mcp-config |
Detection/main_benchmark.py lines 264-268 |
Registry loading implementation |
Summary
- ADR-Bench discovers MCP servers from
Detection/mcp_servers_registry.jsonviaMCPServerManager.load_mcp_servers_registry - Add custom servers by editing the registry JSON with
name,command,args_template, andcapabilitiesfields - Place server scripts in
context_providers/source_codes/mcp_servers_0/following existing conventions - Reference servers in task definitions using the
mcp_serversarray in task JSON files - Capabilities become callable tools named
mcp__<server>__<capability>during benchmark execution
Frequently Asked Questions
What file format does the MCP server registry use?
The registry at Detection/mcp_servers_registry.json uses standard JSON with a top-level servers object containing server definitions. Each server entry includes name, category, description, command, args_template, and capabilities keys. The format matches the schema used by built-in servers like filesystem and cli_executor.
Can I add multiple custom MCP servers to the same benchmark task?
Yes. The mcp_servers field in task definitions accepts an array of server names. For example: "mcp_servers": ["server_a", "server_b", "my_custom_server"]. The benchmark runner generates a .mcp.json containing all listed servers, and Claude receives tool definitions for every capability across all specified servers.
What protocol must my custom MCP server implement?
Your server must communicate via JSON Lines over stdin/stdout following the Model Context Protocol. Each request is a JSON object with tool_use_id, name, and args fields. Responses must include tool_use_id, output, and is_error fields. The existing cli_executor server in the repository provides a working implementation reference.
Where can I verify that my custom server was actually invoked?
Check the tool_analysis section in run_dir/task_<id>/result.json after benchmark execution. The ToolAnalyzer.analyze_tool_usage function tallies all tool calls, including those to custom servers. You should see entries for mcp__<your_server>__<capability> with invocation counts and timing information.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →