How to Execute Specific Benchmark Tasks or a Task Range in ADR-Bench

Pass the --tasks argument to Detection.main_benchmark with a single ID, comma-separated list, or dash-separated range (e.g., 5-10), and optionally control parallelism with --concurrent.

ADR-Bench, the evaluation framework in Uber's ADR repository, provides flexible task selection through command-line options defined in Detection/main_benchmark.py. This architecture lets you run precise subsets of the benchmark suite without modifying source code or filtering results post-hoc.

Understanding the Task Selection Architecture

The benchmark system implements task filtering through three coordinated components:

In main_benchmark.py (lines 1204‑1209), the argument parser defines the --tasks flag:

parser.add_argument('--tasks', type=str, default='',
                    help='Comma-separated list of tasks or range (e.g., 1-5)')

The TaskManager.filter_tasks_by_range method (lines 28‑54) handles the actual parsing logic, expanding ranges and deduplicating IDs before execution.

Syntax for Specifying Tasks

ADR-Bench accepts three formats for the --tasks parameter:

Format Example Description
Single task 42 Execute one specific task ID
List 7,15,23 Run multiple discrete tasks
Range 101-110 Execute all tasks in a contiguous range
Mixed 5-9,20 Combine ranges and individual IDs

The parser automatically handles whitespace and normalizes the input into a set of integers for efficient lookup against the task catalog.

Running Specific Tasks: Code Examples

Single Task Execution

python -m Detection.main_benchmark --tasks 42

This loads only task 42 from tasks.json and executes it with the default concurrency of 10.

Multiple Specific Tasks

python -m Detection.main_benchmark --tasks 7,15,23

The runner expands this to tasks {7, 15, 23} and skips all others in the benchmark suite.

Contiguous Task Ranges

python -m Detection.main_benchmark --tasks 101-110

This executes ten consecutive tasks—useful for systematic coverage testing or slicing large benchmarks across machines.

Mixed Ranges and Individual IDs

python -m Detection.main_benchmark --tasks 5-9,20

Valid combinations include any permutation: 1,3,5-10,12,15-20.

Controlling Concurrency with --concurrent

By default, ADR-Bench runs up to 10 tasks in parallel. Adjust this with the --concurrent flag:

python -m Detection.main_benchmark --tasks 1-5 --concurrent 3

This reduces worker pool size to 3, which helps when:

  • Tasks are resource-intensive (large models, heavy I/O)
  • Debugging requires serialized output
  • Rate limits or API quotas constrain parallelism

Note that the agentdojo benchmark mode ignores --concurrent and always runs sequentially due to its architectural requirements.

Switching Benchmark Modes

The --benchmark argument selects the evaluation backend:


# Default MCP-based benchmark (parallel-capable)

python -m Detection.main_benchmark --tasks 1-10 --benchmark adr_bench

# AgentDojo evaluation (always sequential)

python -m Detection.main_benchmark --tasks 1-3 --benchmark agentdojo

How Task Filtering Works Internally

The execution flow follows this sequence:

  1. main_benchmark.py parses --tasks and passes it to BenchmarkRunner.run_benchmark as task_range
  2. TaskManager.filter_tasks_by_range parses the string, expands ranges, and builds a set of valid IDs
  3. The full task list from tasks.json is filtered to include only matching task_id values
  4. A summary prints to stdout (lines 54‑55): "Filtered N tasks from M total"
  5. BenchmarkRunner schedules filtered tasks according to --concurrent and --benchmark settings

This design ensures minimal memory overhead—even million-task catalogs filter efficiently through set operations.

Common Patterns and Workarounds

  • Empty --tasks: Runs the full benchmark suite (backward-compatible default behavior)
  • Invalid IDs: Silently ignored; only existing task_id values in tasks.json execute
  • Overlapping ranges: Automatically deduplicated by set construction in filter_tasks_by_range

Summary

  • Use --tasks with single IDs, comma lists, or dash ranges to select specific tasks in ADR-Bench
  • Control parallelism via --concurrent (default 10, ignored in agentdojo mode)
  • The filtering pipeline in TaskManager.filter_tasks_by_range handles all parsing and validation
  • Mixed syntax like 5-9,20 is fully supported for complex selection patterns

Frequently Asked Questions

What happens if I specify a task ID that doesn't exist?

ADR-Bench silently ignores invalid IDs. The filter_tasks_by_range method builds a set of requested IDs and intersects it with actual task IDs from tasks.json. Only matching tasks execute; no error or warning appears for missing entries.

Can I use negative numbers or reverse ranges (e.g., 10-1)?

No. The parser expects positive integers with the lower bound first. Reverse ranges like 10-1 evaluate to an empty set, resulting in zero tasks executed. Always specify ranges as start-end where start ≤ end.

Does task order matter for execution?

Task execution order depends on the underlying task list in tasks.json, not the --tasks argument order. The filtering preserves the original catalog's sequence. For deterministic ordering, ensure tasks.json is sorted before benchmark invocation.

How do I verify which tasks will run before executing?

Currently, ADR-Bench does not provide a dry-run mode. The first output line (lines 54‑55 in TaskManager) prints the filtered count: "Filtered N tasks from M total". Monitor this to confirm your range syntax selected the expected tasks.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →