# How to Execute Specific Benchmark Tasks or a Task Range in ADR-Bench

> Learn how to execute specific benchmark tasks or a range of tasks in ADR-Bench. Pass the --tasks argument with IDs, lists, or ranges and control parallelism for efficient testing.

- Repository: [Uber Open Source/ADR](https://github.com/uber/ADR)
- Tags: how-to-guide
- Published: 2026-08-06

---

**Pass the `--tasks` argument to `Detection.main_benchmark` with a single ID, comma-separated list, or dash-separated range (e.g., `5-10`), and optionally control parallelism with `--concurrent`.**

ADR-Bench, the evaluation framework in Uber's ADR repository, provides flexible **task selection** through command-line options defined in [`Detection/main_benchmark.py`](https://github.com/uber/ADR/blob/main/Detection/main_benchmark.py). This architecture lets you run precise subsets of the benchmark suite without modifying source code or filtering results post-hoc.

## Understanding the Task Selection Architecture

The benchmark system implements task filtering through three coordinated components:

- **[`Detection/main_benchmark.py`](https://github.com/uber/ADR/blob/main/Detection/main_benchmark.py)** — Parses CLI arguments and constructs the `BenchmarkRunner`
- **[`Detection/benchmark/TaskManager.py`](https://github.com/uber/ADR/blob/main/Detection/benchmark/TaskManager.py)** — Loads task definitions from [`tasks.json`](https://github.com/uber/ADR/blob/main/tasks.json) and filters by range
- **[`Detection/benchmark/BenchmarkRunner.py`](https://github.com/uber/ADR/blob/main/Detection/benchmark/BenchmarkRunner.py)** — Orchestrates concurrent or sequential execution

In [`main_benchmark.py`](https://github.com/uber/ADR/blob/main/main_benchmark.py) (lines 1204‑1209), the argument parser defines the `--tasks` flag:

```python
parser.add_argument('--tasks', type=str, default='',
                    help='Comma-separated list of tasks or range (e.g., 1-5)')

```

The `TaskManager.filter_tasks_by_range` method (lines 28‑54) handles the actual parsing logic, expanding ranges and deduplicating IDs before execution.

## Syntax for Specifying Tasks

ADR-Bench accepts three formats for the `--tasks` parameter:

| Format | Example | Description |
|--------|---------|-------------|
| Single task | `42` | Execute one specific task ID |
| List | `7,15,23` | Run multiple discrete tasks |
| Range | `101-110` | Execute all tasks in a contiguous range |
| Mixed | `5-9,20` | Combine ranges and individual IDs |

The parser automatically handles whitespace and normalizes the input into a set of integers for efficient lookup against the task catalog.

## Running Specific Tasks: Code Examples

### Single Task Execution

```bash
python -m Detection.main_benchmark --tasks 42

```

This loads only task 42 from [`tasks.json`](https://github.com/uber/ADR/blob/main/tasks.json) and executes it with the default concurrency of 10.

### Multiple Specific Tasks

```bash
python -m Detection.main_benchmark --tasks 7,15,23

```

The runner expands this to tasks `{7, 15, 23}` and skips all others in the benchmark suite.

### Contiguous Task Ranges

```bash
python -m Detection.main_benchmark --tasks 101-110

```

This executes ten consecutive tasks—useful for systematic coverage testing or slicing large benchmarks across machines.

### Mixed Ranges and Individual IDs

```bash
python -m Detection.main_benchmark --tasks 5-9,20

```

Valid combinations include any permutation: `1,3,5-10,12,15-20`.

## Controlling Concurrency with `--concurrent`

By default, ADR-Bench runs up to **10 tasks in parallel**. Adjust this with the `--concurrent` flag:

```bash
python -m Detection.main_benchmark --tasks 1-5 --concurrent 3

```

This reduces worker pool size to 3, which helps when:
- Tasks are resource-intensive (large models, heavy I/O)
- Debugging requires serialized output
- Rate limits or API quotas constrain parallelism

Note that the `agentdojo` benchmark mode ignores `--concurrent` and always runs sequentially due to its architectural requirements.

## Switching Benchmark Modes

The `--benchmark` argument selects the evaluation backend:

```bash

# Default MCP-based benchmark (parallel-capable)

python -m Detection.main_benchmark --tasks 1-10 --benchmark adr_bench

# AgentDojo evaluation (always sequential)

python -m Detection.main_benchmark --tasks 1-3 --benchmark agentdojo

```

## How Task Filtering Works Internally

The execution flow follows this sequence:

1. [`main_benchmark.py`](https://github.com/uber/ADR/blob/main/main_benchmark.py) parses `--tasks` and passes it to `BenchmarkRunner.run_benchmark` as `task_range`
2. `TaskManager.filter_tasks_by_range` parses the string, expands ranges, and builds a `set` of valid IDs
3. The full task list from [`tasks.json`](https://github.com/uber/ADR/blob/main/tasks.json) is filtered to include only matching `task_id` values
4. A summary prints to stdout (lines 54‑55): `"Filtered N tasks from M total"`
5. `BenchmarkRunner` schedules filtered tasks according to `--concurrent` and `--benchmark` settings

This design ensures minimal memory overhead—even million-task catalogs filter efficiently through set operations.

## Common Patterns and Workarounds

- **Empty `--tasks`**: Runs the full benchmark suite (backward-compatible default behavior)
- **Invalid IDs**: Silently ignored; only existing `task_id` values in [`tasks.json`](https://github.com/uber/ADR/blob/main/tasks.json) execute
- **Overlapping ranges**: Automatically deduplicated by set construction in `filter_tasks_by_range`

## Summary

- Use `--tasks` with single IDs, comma lists, or dash ranges to select specific tasks in ADR-Bench
- Control parallelism via `--concurrent` (default 10, ignored in `agentdojo` mode)
- The filtering pipeline in `TaskManager.filter_tasks_by_range` handles all parsing and validation
- Mixed syntax like `5-9,20` is fully supported for complex selection patterns

## Frequently Asked Questions

### What happens if I specify a task ID that doesn't exist?

ADR-Bench silently ignores invalid IDs. The `filter_tasks_by_range` method builds a set of requested IDs and intersects it with actual task IDs from [`tasks.json`](https://github.com/uber/ADR/blob/main/tasks.json). Only matching tasks execute; no error or warning appears for missing entries.

### Can I use negative numbers or reverse ranges (e.g., `10-1`)?

No. The parser expects positive integers with the lower bound first. Reverse ranges like `10-1` evaluate to an empty set, resulting in zero tasks executed. Always specify ranges as `start-end` where `start ≤ end`.

### Does task order matter for execution?

Task execution order depends on the underlying task list in [`tasks.json`](https://github.com/uber/ADR/blob/main/tasks.json), not the `--tasks` argument order. The filtering preserves the original catalog's sequence. For deterministic ordering, ensure [`tasks.json`](https://github.com/uber/ADR/blob/main/tasks.json) is sorted before benchmark invocation.

### How do I verify which tasks will run before executing?

Currently, ADR-Bench does not provide a dry-run mode. The first output line (lines 54‑55 in `TaskManager`) prints the filtered count: `"Filtered N tasks from M total"`. Monitor this to confirm your range syntax selected the expected tasks.