How the ADR Benchmark Handles Concurrency and Performance Optimization

The Uber ADR benchmark uses configurable asyncio semaphores, runtime-adjustable concurrency limits, and real-world speed-up metrics to parallelize security-research tasks while preventing resource exhaustion and measuring actual scaling efficiency.

The Uber ADR (Agent Detection & Response) benchmark is designed to evaluate security systems at scale. Its concurrency implementation balances throughput against system stability, with particular safeguards for stateful workloads like AgentDojo. This article examines how the BenchmarkRunner orchestrates parallel execution, enforces limits, and quantifies performance gains.

Configurable Concurrency Limits

The Config class in Detection/main_benchmark.py centralizes concurrency tuning through the config_benchmark.yaml file. Two parameters control behavior:

  • max_concurrent_tasks – caps simultaneous coroutines (default 3, commonly set to 10)
  • session_collection_delay – pause duration between Claude session creations to reduce API rate-limit hits

The property implementation reads nested configuration with safe defaults:

@property
def max_concurrent_tasks(self) -> int:
    return self._config.get("concurrency", {}).get("max_concurrent_tasks", 3)

Modify values in Detection/config_benchmark.yaml:

concurrency:
  max_concurrent_tasks: 5
  session_collection_delay: 2

Or adjust programmatically at runtime:

from Detection.main_benchmark import Config

cfg = Config()
cfg.update_concurrent_tasks(8)  # dynamic adjustment

Async Semaphore-Driven Execution

The BenchmarkRunner._execute_tasks_concurrently method (lines 909–928) implements the core throttling mechanism. It creates an asyncio.Semaphore sized to the configured limit, with one critical exception: AgentDojo tasks force sequential execution.

max_concurrent = 1 if is_agentdojo else self.config.max_concurrent_tasks
semaphore = asyncio.Semaphore(max_concurrent)

async def run_task_with_semaphore(task_prep, index):
    async with semaphore:
        return await self._execute_single_task(...)

This pattern ensures:

  • Resource protection – OS process and API limits are never exceeded
  • Automatic backpressure – coroutines await semaphore acquisition without manual queue management
  • State safety – AgentDojo's shared Claude project state (lines 925–928) prevents race conditions by design

AgentDojo Sequential Execution Exception

When benchmark_type="agentdojo", the runner overrides any concurrency configuration. The AgentDojo attack suite maintains global state across tasks—specifically a shared Claude "project" context—that would corrupt results if accessed concurrently. This hardcoded max_concurrent = 1 safeguard operates independently of YAML settings or CLI flags.

Real-World Speed-Up Measurement

ADR computes performance impact using two derived metrics after benchmark completion (lines 1155–1177). First, it establishes a sequential baseline:

sequential_time = sum(r.get("execution_time", 0) for r in results)

Then it compares against actual wall-clock time:

actual_speedup = sequential_time / wall_clock_time if wall_clock_time > 0 else 1.0

The concurrency efficiency score normalizes this by theoretical maximum parallelism:

concurrency_efficiency = actual_speedup / min(len(results), self.config.max_concurrent_tasks)

This reveals whether speed-ups achieve ideal scaling. An efficiency of 85% with 10 configured tasks indicates overheads (context switching, API latency, I/O waits) consume 15% of potential gains.

MCP Tool Ratio Tracking

The benchmark aggregates MCP-specific tool calls versus non-MCP invocations (lines 1158–1160). This metric identifies whether parallelism bottlenecks stem from:

  • LLM provider rate limits
  • Tool execution latency
  • Local resource constraints

High MCP ratios with low efficiency suggest tool-server saturation rather than API throttling.

Performance Optimization Configuration

Mechanism Purpose
Semaphore sizing Prevents process explosions and OOM crashes
Session collection delay Spaces Claude session creation to respect provider limits
Concurrency efficiency metric Quantifies real versus theoretical scaling
AgentDojo hardcoded sequentiality Eliminates state corruption in shared-context benchmarks

Inspecting Results

Performance data persists to summary.json in each benchmark run directory:

import json
import pathlib

summary = json.load(open(
    pathlib.Path("benchmark/adr_bench_20230806_123456/summary.json")
))

print(f"Speed-up: {summary['actual_speedup']:.2f}×")
print(f"Efficiency: {summary['concurrency_efficiency']:.1%}")

Typical outputs range from 3–8× speed-up for I/O-bound LLM benchmarks, with efficiency declining as max_concurrent_tasks exceeds available CPU cores or API quota.

Key Source Files

Path Responsibility
Detection/main_benchmark.py Config class, BenchmarkRunner, semaphore orchestration, metrics computation
Detection/config_benchmark.yaml Default concurrency parameters
Detection/benchmark/benchmark_pack.py Task pack loading for TaskManager
Detection/benchmark/agentdojo/benchmarks/agentdojo/benchmark.py AgentDojo sequential execution implementation
Detection/benchmark/agentdojo/benchmarks/agentdojo/task_suite/load_suites.py Per-suite task enumeration

Summary

  • Configurable limits via Config.max_concurrent_tasks and YAML allow environment-specific tuning without code changes
  • Asyncio semaphores in _execute_tasks_concurrently provide lightweight, Python-native concurrency control
  • AgentDojo exception forces max_concurrent = 1 to protect shared state, overriding all other settings
  • Actual speed-up and efficiency metrics measure real performance rather than assuming linear scaling
  • MCP tool ratios help diagnose whether bottlenecks originate from LLM APIs or tool execution layers

Frequently Asked Questions

How do I change the number of concurrent tasks in the ADR benchmark?

Edit Detection/config_benchmark.yaml under the concurrency section, or call Config.update_concurrent_tasks(n) at runtime. The default is 3 if unspecified, though production deployments typically override this to 10.

Why does AgentDojo run sequentially even when I set high concurrency?

AgentDojo tasks share a global Claude project state. Running them concurrently would cause state collisions and invalid results. The code at lines 925–928 of main_benchmark.py hardcodes max_concurrent = 1 when is_agentdojo is true, ignoring configuration values.

What does concurrency efficiency measure?

It compares actual speed-up against theoretical maximum. A value of 0.75 with 10 configured tasks means you achieved 7.5× speed-up instead of the ideal 10×, revealing 25% overhead from context switching, API latency, or I/O waits.

Where can I find the performance results after a benchmark run?

Check benchmark/adr_bench_{timestamp}/summary.json in your working directory. This file contains actual_speedup, concurrency_efficiency, sequential_time, and MCP tool call statistics from the most recent execution.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →