How to Debug MCP Server Failures During ADR Benchmark Execution: A Complete Guide
To debug MCP server failures during ADR benchmark execution, enable the --debug flag to expose registry loading details, verify server connectivity via health endpoints, validate tool availability against the task configuration, and isolate failing tasks by ID for targeted diagnosis.
The MCP (Modular Compute Provider) architecture in Uber's ADR repository powers tool integrations ranging from source code analysis to threat intelligence gathering. When benchmark runs fail due to MCP server issues, the root cause typically hides in one of four layers: registry loading, server instantiation, tool call dispatch, or benchmark orchestration. This guide walks through systematic debugging based on the actual implementation in Detection/main_benchmark.py.
Understanding the MCP Failure Layers
The ADR benchmark pipeline structures MCP interactions into distinct phases. Identifying where your failure occurs accelerates resolution.
Registry Loading Failures
Symptom: Empty server list, "No MCP servers configured" messages, or missing tool errors before any server starts.
Investigation: Check MCPServerManager._load_registry() (lines ~265-275 in Detection/main_benchmark.py). This method parses config/mcp_registry.yaml and builds the server configuration list. Debug output shows:
✅ Loaded 3 MCP server configurations
If this line shows zero servers or the wrong count, inspect the YAML structure. Common issues include:
- Missing
nameortool_allowlistfields in server entries - Syntax errors (tabs instead of spaces, unclosed quotes)
- Incorrect file path resolution
Server Instantiation Failures
Symptom: ConnectionError, TimeoutError, or FastMCP constructor exceptions during benchmark initialization.
Investigation: Each server module under Detection/context_providers/source_codes/mcp_servers_*/ creates a FastMCP instance. For example, in Detection/context_providers/source_codes/mcp_servers_2/slack_server.py (lines 10-16):
from fastmcp import FastMCP
mcp = FastMCP("slack_server")
# ... tool definitions follow
The server must bind to a port and start a local HTTP endpoint. Verify reachability:
# Extract port from registry and test health endpoint
SERVER_NAME=slack_server
PORT=$(grep -A5 "name: $SERVER_NAME" config/mcp_registry.yaml | grep port | awk '{print $2}')
curl -s http://localhost:${PORT}/health || echo "Server $SERVER_NAME unreachable"
Non-200 responses indicate the server process crashed or failed to start. Check logs/<server_name>.log for stack traces.
Tool Call Dispatch Failures
Symptom: "MCP tool call failed" in CLI output, missing tool names, or task execution halts mid-run.
Investigation: MCPServerManager._track_tool_usage() (lines 498-507) parses Claude CLI output to correlate tool calls with servers. The benchmark task JSON specifies expected tools in an "mcp_servers" field. Discrepancies between:
allowed_toolsin the server configuration (printed in debug logs)- Tools referenced by the task
...trigger "MCP tool not found" errors. Cross-reference these lists before execution.
Benchmark Orchestration Failures
Symptom: Complete benchmark abort with stack trace from main_benchmark.run_task().
Investigation: The run_task method (lines 640-660) orchestrates MCP server lifecycle management. Exceptions here often indicate resource exhaustion, port conflicts, or unhandled exceptions in server cleanup.
Step-by-Step Debugging Workflow
1. Enable Verbose Logging
Run the benchmark with the --debug flag to expose internal MCP state:
python -m Detection.main_benchmark \
--benchmark agentdojo \
--task-id 42 \
--debug
Debug output includes:
- Parsed registry entries with server names and ports
- Per-server
FastMCPinitialization status - Complete tool call logs with timing
2. Inspect the MCP Registry Programmatically
For programmatic verification, load the registry directly:
from Detection.main_benchmark import MCPServerManager, Config
cfg = Config.from_yaml("config/benchmark.yaml")
mcp_mgr = MCPServerManager(cfg)
# Verify loaded server names match filesystem
for server in mcp_mgr.registry:
print(f"Server: {server.name}, Tools: {server.allowed_tools}")
# Expected output: Server: slack_server, Tools: ['send_message', 'read_channel']
3. Validate Server Health Endpoints
Each FastMCP server exposes a health check. Automate verification:
#!/bin/bash
# validate_mcp_servers.sh
REGISTRY="config/mcp_registry.yaml"
for server in $(yq '.servers[].name' "$REGISTRY"); do
port=$(yq ".servers[] | select(.name == \"$server\") | .port" "$REGISTRY")
status=$(curl -s -o /dev/null -w "%{http_code}" "http://localhost:${port}/health")
if [ "$status" -eq 200 ]; then
echo "✅ $server: healthy (port $port)"
else
echo "❌ $server: unreachable (port $port, status $status)"
echo " Check logs/${server}.log for details"
fi
done
4. Isolate Failing Tasks
Benchmark outputs tag results with task IDs ([task_XXX]). Re-run single tasks to reduce noise:
python -m Detection.main_benchmark \
--task-id 157 \
--debug \
2>&1 | tee task_157_debug.log
Examine task_157_debug.log for MCPServerManager activity surrounding the failure timestamp.
5. Fix Tool Usage Parsing Regressions
If debug shows warnings like:
Failed to extract MCP tool usage: KeyError('tool_name')
The issue lies in _extract_mcp_tool_name() within adr_agent.py (lines 891-901). Claude CLI output format changes can break the regex. Verify the expected pattern against actual CLI output in your logs.
6. Deploy Mock Servers for Isolation
When external dependencies (OpenAI API, Slack, etc.) are unavailable, substitute mocks:
# Detection/context_providers/mock_server.py
from fastmcp import FastMCP
from unittest.mock import MagicMock
mcp = FastMCP("mock_server")
@mcp.tool()
def get_source_code(path: str) -> str:
"""Return static source for testing."""
return f"# Mock source for {path}\ndef example(): pass"
# In benchmark config, temporarily replace real server:
# servers:
# - name: source_code_analyzer_server
# module: context_providers.mock_server # was: source_codes.mcp_servers_1.source_code_analyzer_server
7. Verify Configuration Integrity
The BenchmarkConfig object (instantiated lines 258-275 in main_benchmark.py) validates server entries. Ensure every server specification includes:
servers:
- name: threat_intelligence_server
module: context_providers.threat_intelligence_server
tool_allowlist:
- query_threat_intel
- check_ip_reputation
port: 8081 # optional, auto-assigned if omitted
8. Run Unit Tests for Validation
After fixes, confirm stability with the repository's test suite:
pytest Detection/tests/test_main_benchmark.py::TestMCPServerManager -v
The TestMCPServerManager class (lines 71-93) mocks a complete MCP server lifecycle, validating:
- Registry registration
- Configuration object creation
- Tool usage tracking accuracy
Common Error Patterns and Solutions
| Error Message | Root Cause | Fix |
|---|---|---|
| "No MCP servers in registry" | Empty or malformed mcp_registry.yaml |
Verify YAML syntax and server list structure |
ConnectionError: [Errno 111] Connection refused |
Target server not running | Start server process or check port conflicts |
| "MCP tool 'X' not found" | Tool missing from allowed_tools or server mismatch |
Update tool_allowlist or correct server assignment |
| "Failed to extract MCP tool usage" | Claude CLI output format changed | Adjust regex in _extract_mcp_tool_name() |
TimeoutError connecting to server |
Server overloaded or deadlocked | Increase timeout or review server logs for blocking calls |
Key Source Files for Deep Debugging
-
Detection/main_benchmark.py— CoreMCPServerManagerimplementation with registry loading, server lifecycle, and tracking: github.com/uber/ADR/blob/main/Detection/main_benchmark.py -
Detection/guardrail/adr_agent/adr_baseline.py— Agent integration that invokes MCP tools and parses usage: github.com/uber/ADR/blob/main/Detection/guardrail/adr_agent/adr_baseline.py -
Detection/context_providers/source_codes/mcp_servers_2/slack_server.py— ExampleFastMCPserver implementation: github.com/uber/ADR/blob/main/Detection/context_providers/source_codes/mcp_servers_2/slack_server.py -
Detection/tests/test_main_benchmark.py— Unit tests including mocked MCP server validation: github.com/uber/ADR/blob/main/Detection/tests/test_main_benchmark.py#L71
Summary
Debugging MCP server failures during ADR benchmark execution follows a layered approach:
- Enable
--debugto expose registry state and tool call details - Verify server health endpoints to confirm
FastMCPprocesses are running - Cross-reference tool lists between server configurations and task requirements
- Isolate by task ID to reduce diagnostic noise
- Use mock servers when external dependencies are unavailable
- Validate with unit tests after implementing fixes
The MCPServerManager class in Detection/main_benchmark.py centralizes all MCP operations—registry loading through __init__, server management through _load_registry(), and usage tracking through _track_tool_usage(). Understanding these three methods provides complete visibility into the MCP pipeline.
Frequently Asked Questions
How do I know which MCP server is causing a benchmark failure?
Run with --debug and examine the server initialization log lines. Each FastMCP instance logs its name and bound port. If a server fails to start, the exception propagates immediately with the server name in the traceback. For runtime failures, check logs/<server_name>.log for the specific process output.
What does "MCP tool not found" mean during task execution?
This error indicates a mismatch between the tool name requested by the benchmark task and the allowed_tools list in the server configuration. The task JSON specifies required tools under "mcp_servers". Compare this against the output from your debug logs showing each server's allowed_tools. Add missing tools to the server's tool_allowlist or reassign the task to a different server.
Can I run benchmarks without all MCP servers available?
Yes. Temporarily replace unavailable servers with mock implementations that return static data. The repository includes a mock_server.py skeleton in Detection/context_providers/. Update config/mcp_registry.yaml to point the problematic server entry to the mock module, or remove the server from the task's "mcp_servers" list if the task permits reduced functionality.
Why does tool usage tracking fail with parsing errors?
The _extract_mcp_tool_name() function in adr_agent.py uses regex to parse Claude CLI output. Anthropic updates the CLI format periodically, which can break this parsing. Check the actual CLI output in your debug logs against the expected pattern in lines 891-901. Update the regex to match the new format, or consider using structured logging if available in your Claude version.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →