How to Add Custom MCP Servers to ADR-Bench for Testing: A Complete Guide

Add a custom MCP server to ADR-Bench by defining it in Detection/mcp_servers_registry.json, placing the executable in the recommended directory, and referencing it in a task's mcp_servers array.

ADR-Bench is Uber's automated detection and response benchmarking framework that evaluates AI agents against security tasks. The system discovers available MCP (Model Context Protocol) servers through a central registry, making it straightforward to extend with custom implementations for specialized testing scenarios. This guide walks through the exact steps to register, implement, and invoke your own MCP servers within the ADR-Bench ecosystem.

How ADR-Bench Discovers and Loads MCP Servers

ADR-Bench uses a three-phase architecture to manage custom MCP servers for testing purposes. Understanding this flow helps ensure your server integrates correctly.

Phase 1: Registry Loading

The MCPServerManager.load_mcp_servers_registry() function in Detection/main_benchmark.py (lines 264-268) reads Detection/mcp_servers_registry.json at startup. This JSON file contains all server definitions, keyed by server name.


# From Detection/main_benchmark.py

def load_mcp_servers_registry(self):
    with open(self.config.mcp_registry_file, 'r') as f:
        registry = json.load(f)
    return registry.get('servers', {})

The Config.mcp_registry_file attribute points to the registry path (lines 75-76), making this the single source of truth for available servers.

Phase 2: Per-Task Configuration Generation

When executing a task, TaskExecutor.execute_ads_task filters the registry to servers listed in the task's mcp_servers field (lines 70-77). The MCPServerManager.create_mcp_config method then generates a temporary .mcp.json file containing only the required servers.

Phase 3: CLI Invocation with Claude

The benchmark runner passes the generated configuration to Claude via --mcp-config .mcp.json (lines 33-36). Claude exposes each capability as a tool named mcp__<server_name>__<capability>, which the agent can invoke during task execution.

Step-by-Step: Adding a Custom MCP Server

Follow these four steps to add custom MCP servers to ADR-Bench for testing purposes.

Step 1: Define the Server in the Registry

Edit Detection/mcp_servers_registry.json and add a new entry under the servers object. The schema matches existing built-in servers like filesystem (lines 7-31).

{
  "servers": {
    "my_custom_server": {
      "name": "my_custom_server",
      "category": "Research And Data",
      "description": "Demo server that returns a static JSON payload for testing",
      "package": "my-custom-mcp",
      "type": "local",
      "command": "uv",
      "args_template": [
        "run",
        "python",
        "../../../../context_providers/source_codes/mcp_servers_0/my_custom_server/my_custom_server.py"
      ],
      "capabilities": [
        "return_static_json",
        "echo_input"
      ]
    }
  }
}

Required fields include:

  • name: Unique identifier used in task definitions
  • command: Executable to launch the server
  • args_template: Array of arguments passed to the command
  • capabilities: List of tool names this server exposes

Step 2: Implement the Server Logic (Optional)

Place the executable script at the path referenced in args_template. The recommended location follows the convention established by existing servers:


Detection/context_providers/source_codes/mcp_servers_0/<your_server>/

A minimal MCP server implementation reads JSON requests from stdin and writes JSON responses to stdout:

#!/usr/bin/env python3
import sys
import json

def main():
    for line in sys.stdin:
        request = json.loads(line)
        tool_id = request.get("tool_use_id")
        name = request.get("name", "")
        args = request.get("args", {})

        if name == "return_static_json":
            response = {
                "tool_use_id": tool_id,
                "output": {"message": "static payload", "received_args": args},
                "is_error": False
            }
        elif name == "echo_input":
            response = {
                "tool_use_id": tool_id,
                "output": {"echo": args},
                "is_error": False
            }
        else:
            response = {
                "tool_use_id": tool_id,
                "output": {"error": f"Unknown tool: {name}"},
                "is_error": True
            }
        
        print(json.dumps(response))
        sys.stdout.flush()

if __name__ == "__main__":
    main()

Use existing servers like cli_executor at cli_executor/cli_executor.py as templates for more complex implementations.

Step 3: Reference the Server in a Benchmark Task

Add your server name to a task's mcp_servers array in Detection/tasks.json (or any custom task file):

{
  "task_id": 999,
  "description": "Test custom server capabilities",
  "user_prompt": "Please call the custom server and retrieve the static payload, then echo back the value 'test'.",
  "mcp_servers": ["my_custom_server"],
  "expected_behavior": "benign"
}

The task loading logic extracts mcp_servers from each task definition (lines 23-26). Only servers explicitly listed here are included in that task's MCP configuration.

Step 4: Run and Verify

Execute the benchmark targeting your new task:

python -m Detection.main_benchmark --tasks 999

The runner will:

  1. Load the updated registry via load_mcp_servers_registry
  2. Build a per-task MCP configuration via create_mcp_config
  3. Launch your server using the uv run python ... command
  4. Expose capabilities as mcp__my_custom_server__return_static_json and mcp__my_custom_server__echo_input

Verify invocation by checking run_dir/task_999/result.json. The tool_analysis section tallies calls to your server's capabilities via ToolAnalyzer.analyze_tool_usage (lines 71-77).

Validating Custom MCP Server Integration

After adding custom MCP servers to ADR-Bench for testing, confirm proper operation through these checks:

  • Registry parsing: Ensure your JSON entry has no syntax errors; malformed entries will cause load_mcp_servers_registry to fail
  • Path resolution: Verify the args_template path correctly resolves from the execution directory
  • Capability naming: Tool calls use the exact capability strings defined in the registry
  • Result analysis: Check that tool_analysis in result.json counts invocations of your custom tools

File Reference for Custom MCP Server Development

File Purpose
Detection/mcp_servers_registry.json Central registry—add your server definition here
Detection/main_benchmark.py Core benchmark logic including MCPServerManager.load_mcp_servers_registry, create_mcp_config, and CLI invocation
Detection/tasks.json Task definitions—include server names in mcp_servers arrays
Detection/context_providers/source_codes/mcp_servers_0/ Recommended directory for server implementations
Detection/main_benchmark.py lines 33-36 Claude CLI invocation with --mcp-config
Detection/main_benchmark.py lines 264-268 Registry loading implementation

Summary

  • ADR-Bench discovers MCP servers from Detection/mcp_servers_registry.json via MCPServerManager.load_mcp_servers_registry
  • Add custom servers by editing the registry JSON with name, command, args_template, and capabilities fields
  • Place server scripts in context_providers/source_codes/mcp_servers_0/ following existing conventions
  • Reference servers in task definitions using the mcp_servers array in task JSON files
  • Capabilities become callable tools named mcp__<server>__<capability> during benchmark execution

Frequently Asked Questions

What file format does the MCP server registry use?

The registry at Detection/mcp_servers_registry.json uses standard JSON with a top-level servers object containing server definitions. Each server entry includes name, category, description, command, args_template, and capabilities keys. The format matches the schema used by built-in servers like filesystem and cli_executor.

Can I add multiple custom MCP servers to the same benchmark task?

Yes. The mcp_servers field in task definitions accepts an array of server names. For example: "mcp_servers": ["server_a", "server_b", "my_custom_server"]. The benchmark runner generates a .mcp.json containing all listed servers, and Claude receives tool definitions for every capability across all specified servers.

What protocol must my custom MCP server implement?

Your server must communicate via JSON Lines over stdin/stdout following the Model Context Protocol. Each request is a JSON object with tool_use_id, name, and args fields. Responses must include tool_use_id, output, and is_error fields. The existing cli_executor server in the repository provides a working implementation reference.

Where can I verify that my custom server was actually invoked?

Check the tool_analysis section in run_dir/task_<id>/result.json after benchmark execution. The ToolAnalyzer.analyze_tool_usage function tallies all tool calls, including those to custom servers. You should see entries for mcp__<your_server>__<capability> with invocation counts and timing information.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →