How to Execute Specific Benchmark Tasks or a Task Range in ADR-Bench
Pass the --tasks argument to Detection.main_benchmark with a single ID, comma-separated list, or dash-separated range (e.g., 5-10), and optionally control parallelism with --concurrent.
ADR-Bench, the evaluation framework in Uber's ADR repository, provides flexible task selection through command-line options defined in Detection/main_benchmark.py. This architecture lets you run precise subsets of the benchmark suite without modifying source code or filtering results post-hoc.
Understanding the Task Selection Architecture
The benchmark system implements task filtering through three coordinated components:
Detection/main_benchmark.py— Parses CLI arguments and constructs theBenchmarkRunnerDetection/benchmark/TaskManager.py— Loads task definitions fromtasks.jsonand filters by rangeDetection/benchmark/BenchmarkRunner.py— Orchestrates concurrent or sequential execution
In main_benchmark.py (lines 1204‑1209), the argument parser defines the --tasks flag:
parser.add_argument('--tasks', type=str, default='',
help='Comma-separated list of tasks or range (e.g., 1-5)')
The TaskManager.filter_tasks_by_range method (lines 28‑54) handles the actual parsing logic, expanding ranges and deduplicating IDs before execution.
Syntax for Specifying Tasks
ADR-Bench accepts three formats for the --tasks parameter:
| Format | Example | Description |
|---|---|---|
| Single task | 42 |
Execute one specific task ID |
| List | 7,15,23 |
Run multiple discrete tasks |
| Range | 101-110 |
Execute all tasks in a contiguous range |
| Mixed | 5-9,20 |
Combine ranges and individual IDs |
The parser automatically handles whitespace and normalizes the input into a set of integers for efficient lookup against the task catalog.
Running Specific Tasks: Code Examples
Single Task Execution
python -m Detection.main_benchmark --tasks 42
This loads only task 42 from tasks.json and executes it with the default concurrency of 10.
Multiple Specific Tasks
python -m Detection.main_benchmark --tasks 7,15,23
The runner expands this to tasks {7, 15, 23} and skips all others in the benchmark suite.
Contiguous Task Ranges
python -m Detection.main_benchmark --tasks 101-110
This executes ten consecutive tasks—useful for systematic coverage testing or slicing large benchmarks across machines.
Mixed Ranges and Individual IDs
python -m Detection.main_benchmark --tasks 5-9,20
Valid combinations include any permutation: 1,3,5-10,12,15-20.
Controlling Concurrency with --concurrent
By default, ADR-Bench runs up to 10 tasks in parallel. Adjust this with the --concurrent flag:
python -m Detection.main_benchmark --tasks 1-5 --concurrent 3
This reduces worker pool size to 3, which helps when:
- Tasks are resource-intensive (large models, heavy I/O)
- Debugging requires serialized output
- Rate limits or API quotas constrain parallelism
Note that the agentdojo benchmark mode ignores --concurrent and always runs sequentially due to its architectural requirements.
Switching Benchmark Modes
The --benchmark argument selects the evaluation backend:
# Default MCP-based benchmark (parallel-capable)
python -m Detection.main_benchmark --tasks 1-10 --benchmark adr_bench
# AgentDojo evaluation (always sequential)
python -m Detection.main_benchmark --tasks 1-3 --benchmark agentdojo
How Task Filtering Works Internally
The execution flow follows this sequence:
main_benchmark.pyparses--tasksand passes it toBenchmarkRunner.run_benchmarkastask_rangeTaskManager.filter_tasks_by_rangeparses the string, expands ranges, and builds asetof valid IDs- The full task list from
tasks.jsonis filtered to include only matchingtask_idvalues - A summary prints to stdout (lines 54‑55):
"Filtered N tasks from M total" BenchmarkRunnerschedules filtered tasks according to--concurrentand--benchmarksettings
This design ensures minimal memory overhead—even million-task catalogs filter efficiently through set operations.
Common Patterns and Workarounds
- Empty
--tasks: Runs the full benchmark suite (backward-compatible default behavior) - Invalid IDs: Silently ignored; only existing
task_idvalues intasks.jsonexecute - Overlapping ranges: Automatically deduplicated by set construction in
filter_tasks_by_range
Summary
- Use
--taskswith single IDs, comma lists, or dash ranges to select specific tasks in ADR-Bench - Control parallelism via
--concurrent(default 10, ignored inagentdojomode) - The filtering pipeline in
TaskManager.filter_tasks_by_rangehandles all parsing and validation - Mixed syntax like
5-9,20is fully supported for complex selection patterns
Frequently Asked Questions
What happens if I specify a task ID that doesn't exist?
ADR-Bench silently ignores invalid IDs. The filter_tasks_by_range method builds a set of requested IDs and intersects it with actual task IDs from tasks.json. Only matching tasks execute; no error or warning appears for missing entries.
Can I use negative numbers or reverse ranges (e.g., 10-1)?
No. The parser expects positive integers with the lower bound first. Reverse ranges like 10-1 evaluate to an empty set, resulting in zero tasks executed. Always specify ranges as start-end where start ≤ end.
Does task order matter for execution?
Task execution order depends on the underlying task list in tasks.json, not the --tasks argument order. The filtering preserves the original catalog's sequence. For deterministic ordering, ensure tasks.json is sorted before benchmark invocation.
How do I verify which tasks will run before executing?
Currently, ADR-Bench does not provide a dry-run mode. The first output line (lines 54‑55 in TaskManager) prints the filtered count: "Filtered N tasks from M total". Monitor this to confirm your range syntax selected the expected tasks.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →