# How to Add Custom MCP Servers to ADR-Bench for Testing: A Complete Guide

> Learn how to add custom MCP servers to ADR-Bench for testing. Follow this guide to define servers in mcp_servers_registry.json and reference them in your tasks.

- Repository: [Uber Open Source/ADR](https://github.com/uber/ADR)
- Tags: how-to-guide
- Published: 2026-08-06

---

**Add a custom MCP server to ADR-Bench by defining it in [`Detection/mcp_servers_registry.json`](https://github.com/uber/ADR/blob/main/Detection/mcp_servers_registry.json), placing the executable in the recommended directory, and referencing it in a task's `mcp_servers` array.**

ADR-Bench is Uber's automated detection and response benchmarking framework that evaluates AI agents against security tasks. The system discovers available MCP (Model Context Protocol) servers through a central registry, making it straightforward to extend with custom implementations for specialized testing scenarios. This guide walks through the exact steps to register, implement, and invoke your own MCP servers within the ADR-Bench ecosystem.

## How ADR-Bench Discovers and Loads MCP Servers

ADR-Bench uses a three-phase architecture to manage custom MCP servers for testing purposes. Understanding this flow helps ensure your server integrates correctly.

### Phase 1: Registry Loading

The `MCPServerManager.load_mcp_servers_registry()` function in [`Detection/main_benchmark.py`](https://github.com/uber/ADR/blob/main/Detection/main_benchmark.py) (lines 264-268) reads [`Detection/mcp_servers_registry.json`](https://github.com/uber/ADR/blob/main/Detection/mcp_servers_registry.json) at startup. This JSON file contains all server definitions, keyed by server name.

```python

# From Detection/main_benchmark.py

def load_mcp_servers_registry(self):
    with open(self.config.mcp_registry_file, 'r') as f:
        registry = json.load(f)
    return registry.get('servers', {})

```

The `Config.mcp_registry_file` attribute points to the registry path (lines 75-76), making this the single source of truth for available servers.

### Phase 2: Per-Task Configuration Generation

When executing a task, `TaskExecutor.execute_ads_task` filters the registry to servers listed in the task's `mcp_servers` field (lines 70-77). The `MCPServerManager.create_mcp_config` method then generates a temporary [`.mcp.json`](https://github.com/uber/ADR/blob/main/.mcp.json) file containing only the required servers.

### Phase 3: CLI Invocation with Claude

The benchmark runner passes the generated configuration to Claude via `--mcp-config .mcp.json` (lines 33-36). Claude exposes each capability as a tool named `mcp__<server_name>__<capability>`, which the agent can invoke during task execution.

## Step-by-Step: Adding a Custom MCP Server

Follow these four steps to add custom MCP servers to ADR-Bench for testing purposes.

### Step 1: Define the Server in the Registry

Edit [`Detection/mcp_servers_registry.json`](https://github.com/uber/ADR/blob/main/Detection/mcp_servers_registry.json) and add a new entry under the `servers` object. The schema matches existing built-in servers like `filesystem` (lines 7-31).

```json
{
  "servers": {
    "my_custom_server": {
      "name": "my_custom_server",
      "category": "Research And Data",
      "description": "Demo server that returns a static JSON payload for testing",
      "package": "my-custom-mcp",
      "type": "local",
      "command": "uv",
      "args_template": [
        "run",
        "python",
        "../../../../context_providers/source_codes/mcp_servers_0/my_custom_server/my_custom_server.py"
      ],
      "capabilities": [
        "return_static_json",
        "echo_input"
      ]
    }
  }
}

```

Required fields include:
- **name**: Unique identifier used in task definitions
- **command**: Executable to launch the server
- **args_template**: Array of arguments passed to the command
- **capabilities**: List of tool names this server exposes

### Step 2: Implement the Server Logic (Optional)

Place the executable script at the path referenced in `args_template`. The recommended location follows the convention established by existing servers:

```

Detection/context_providers/source_codes/mcp_servers_0/<your_server>/

```

A minimal MCP server implementation reads JSON requests from stdin and writes JSON responses to stdout:

```python
#!/usr/bin/env python3
import sys
import json

def main():
    for line in sys.stdin:
        request = json.loads(line)
        tool_id = request.get("tool_use_id")
        name = request.get("name", "")
        args = request.get("args", {})

        if name == "return_static_json":
            response = {
                "tool_use_id": tool_id,
                "output": {"message": "static payload", "received_args": args},
                "is_error": False
            }
        elif name == "echo_input":
            response = {
                "tool_use_id": tool_id,
                "output": {"echo": args},
                "is_error": False
            }
        else:
            response = {
                "tool_use_id": tool_id,
                "output": {"error": f"Unknown tool: {name}"},
                "is_error": True
            }
        
        print(json.dumps(response))
        sys.stdout.flush()

if __name__ == "__main__":
    main()

```

Use existing servers like `cli_executor` at [`cli_executor/cli_executor.py`](https://github.com/uber/ADR/blob/main/cli_executor/cli_executor.py) as templates for more complex implementations.

### Step 3: Reference the Server in a Benchmark Task

Add your server name to a task's `mcp_servers` array in [`Detection/tasks.json`](https://github.com/uber/ADR/blob/main/Detection/tasks.json) (or any custom task file):

```json
{
  "task_id": 999,
  "description": "Test custom server capabilities",
  "user_prompt": "Please call the custom server and retrieve the static payload, then echo back the value 'test'.",
  "mcp_servers": ["my_custom_server"],
  "expected_behavior": "benign"
}

```

The task loading logic extracts `mcp_servers` from each task definition (lines 23-26). Only servers explicitly listed here are included in that task's MCP configuration.

### Step 4: Run and Verify

Execute the benchmark targeting your new task:

```bash
python -m Detection.main_benchmark --tasks 999

```

The runner will:
1. Load the updated registry via `load_mcp_servers_registry`
2. Build a per-task MCP configuration via `create_mcp_config`
3. Launch your server using the `uv run python ...` command
4. Expose capabilities as `mcp__my_custom_server__return_static_json` and `mcp__my_custom_server__echo_input`

Verify invocation by checking [`run_dir/task_999/result.json`](https://github.com/uber/ADR/blob/main/run_dir/task_999/result.json). The `tool_analysis` section tallies calls to your server's capabilities via `ToolAnalyzer.analyze_tool_usage` (lines 71-77).

## Validating Custom MCP Server Integration

After adding custom MCP servers to ADR-Bench for testing, confirm proper operation through these checks:

- **Registry parsing**: Ensure your JSON entry has no syntax errors; malformed entries will cause `load_mcp_servers_registry` to fail
- **Path resolution**: Verify the `args_template` path correctly resolves from the execution directory
- **Capability naming**: Tool calls use the exact capability strings defined in the registry
- **Result analysis**: Check that `tool_analysis` in [`result.json`](https://github.com/uber/ADR/blob/main/result.json) counts invocations of your custom tools

## File Reference for Custom MCP Server Development

| File | Purpose |
|------|---------|
| [`Detection/mcp_servers_registry.json`](https://github.com/uber/ADR/blob/main/Detection/mcp_servers_registry.json) | Central registry—add your server definition here |
| [`Detection/main_benchmark.py`](https://github.com/uber/ADR/blob/main/Detection/main_benchmark.py) | Core benchmark logic including `MCPServerManager.load_mcp_servers_registry`, `create_mcp_config`, and CLI invocation |
| [`Detection/tasks.json`](https://github.com/uber/ADR/blob/main/Detection/tasks.json) | Task definitions—include server names in `mcp_servers` arrays |
| `Detection/context_providers/source_codes/mcp_servers_0/` | Recommended directory for server implementations |
| [`Detection/main_benchmark.py`](https://github.com/uber/ADR/blob/main/Detection/main_benchmark.py) lines 33-36 | Claude CLI invocation with `--mcp-config` |
| [`Detection/main_benchmark.py`](https://github.com/uber/ADR/blob/main/Detection/main_benchmark.py) lines 264-268 | Registry loading implementation |

## Summary

- **ADR-Bench discovers MCP servers from [`Detection/mcp_servers_registry.json`](https://github.com/uber/ADR/blob/main/Detection/mcp_servers_registry.json)** via `MCPServerManager.load_mcp_servers_registry`
- **Add custom servers by editing the registry JSON** with `name`, `command`, `args_template`, and `capabilities` fields
- **Place server scripts in `context_providers/source_codes/mcp_servers_0/`** following existing conventions
- **Reference servers in task definitions** using the `mcp_servers` array in task JSON files
- **Capabilities become callable tools** named `mcp__<server>__<capability>` during benchmark execution

## Frequently Asked Questions

### What file format does the MCP server registry use?

The registry at [`Detection/mcp_servers_registry.json`](https://github.com/uber/ADR/blob/main/Detection/mcp_servers_registry.json) uses standard JSON with a top-level `servers` object containing server definitions. Each server entry includes `name`, `category`, `description`, `command`, `args_template`, and `capabilities` keys. The format matches the schema used by built-in servers like `filesystem` and `cli_executor`.

### Can I add multiple custom MCP servers to the same benchmark task?

Yes. The `mcp_servers` field in task definitions accepts an array of server names. For example: `"mcp_servers": ["server_a", "server_b", "my_custom_server"]`. The benchmark runner generates a [`.mcp.json`](https://github.com/uber/ADR/blob/main/.mcp.json) containing all listed servers, and Claude receives tool definitions for every capability across all specified servers.

### What protocol must my custom MCP server implement?

Your server must communicate via JSON Lines over stdin/stdout following the Model Context Protocol. Each request is a JSON object with `tool_use_id`, `name`, and `args` fields. Responses must include `tool_use_id`, `output`, and `is_error` fields. The existing `cli_executor` server in the repository provides a working implementation reference.

### Where can I verify that my custom server was actually invoked?

Check the `tool_analysis` section in `run_dir/task_<id>/result.json` after benchmark execution. The `ToolAnalyzer.analyze_tool_usage` function tallies all tool calls, including those to custom servers. You should see entries for `mcp__<your_server>__<capability>` with invocation counts and timing information.