# How the ADHD Concurrency Parameter Limits Parallel LLM Calls

> Discover how the ADHD concurrency parameter limits parallel LLM calls by routing requests through a single p-limit instance, preventing over-saturation.

- Repository: [Udit Akhouri/adhd](https://github.com/UditAkhourii/adhd)
- Tags: how-to-guide
- Published: 2026-08-19

---

**The `concurrency` parameter in ADHD caps parallel LLM calls by routing every request through a single `p-limit` instance, ensuring that no more than the configured number of promises run simultaneously during divergence and deepening phases.**

The ADHD repository by UditAkhourii is an open-source reasoning engine that generates multiple solution branches by executing numerous LLM prompts in parallel. To prevent API rate-limit errors and network congestion, the engine exposes a `--concurrency` flag that controls exactly how many LLM requests can be in flight at once. Understanding how this concurrency parameter limits parallel LLM calls is essential for tuning performance and staying within provider quotas.

## CLI Parsing and Type Definitions

The concurrency journey begins at the command line and flows through the type system before reaching the execution engine.

### Parsing the `--concurrency` Flag in [`cli.ts`](https://github.com/UditAkhourii/adhd/blob/main/cli.ts)

In [`src/cli.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/cli.ts), the CLI reads the `--concurrency` flag and stores it inside the `RunOptions` object that drives each run. The flag is defined at line 53, making it available to users without touching the source code.

### Declaring `concurrency` in [`types.ts`](https://github.com/UditAkhourii/adhd/blob/main/types.ts)

The shape of the configuration is declared in [`src/types.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/types.ts) at line 55, where the `RunOptions` interface documents the `concurrency` property and its default value. This type safety ensures that every downstream consumer expects a numeric limit.

## Engine Setup and the Promise Limiter

Once the options reach the core engine, ADHD translates the numeric concurrency value into a hard ceiling on parallel LLM requests.

### Destructuring `RunOptions` with a Fallback

Inside [`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts), the `run()` function destructures the incoming options and applies a fallback of `4` when no value is supplied. This logic appears at lines 19-22, guaranteeing that the engine always has a bounded concurrency setting even if the caller omits the flag.

### Creating the Limiter with `p-limit`

At line 49 of [`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts), the engine instantiates a single shared limiter by calling `pLimit(concurrency)`. This utility from the **p-limit** library creates a promise pool that permits only `concurrency` number of active promises at any moment. All LLM-bound work is then funneled through this one instance.

## Applying the Limit to Divergence and Deepening

The limiter is applied uniformly across the two phases that generate the heaviest LLM traffic: diverging branches and deepening ideas.

### Capping the Divergence Phase

During divergence, the engine spawns one branch per frame. In [`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts) at lines 54-61, each `divergeBranch` call is wrapped inside `limit(async () => …)`. Even though the code uses `Promise.all` over the full array of frames, the underlying limiter ensures that at most `concurrency` branches are active simultaneously.

```ts
const limit = pLimit(concurrency); // ← creates the bounded pool

// Divergence: each frame is scheduled through the limiter
const branches = await Promise.all(
  frames.map((f) =>
    limit(async () => await divergeBranch(problem, context, f, ideasPerFrame, model))
  )
);

```

### Capping the Deepening Phase

After the initial ideas are ranked, the top-K candidates enter a deepening phase. In [`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts) at lines 98-107, these calls are also routed through the same `limit` function. Reusing the identical limiter instance means the deepening phase respects the same global ceiling, preventing bursts of concurrent LLM calls when many ideas qualify for elaboration.

```ts
// Deepening: top‑K ideas are also limited
const deepened = await Promise.all(
  toDeepen.map((idea) =>
    limit(async () => await deepenIdea(problem, idea, allIdeas, model))
  )
);

```

## Practical Example: Configuring Concurrency

You can tune the limit when invoking the engine programmatically. The following example raises the cap to `8` simultaneous LLM calls:

```ts
// Example: run with a custom concurrency of 8
import { run } from "./engine.js";

await run({
  problem: "How can we improve cache consistency?",
  concurrency: 8,          // ← allows up to 8 simultaneous LLM calls
  framesPerRun: 6,
  ideasPerFrame: 5,
});

```

By adjusting `concurrency`, you balance throughput against API quotas. A higher value reduces total wall-clock time for large batches, while a lower value minimizes the risk of rate-limit errors on restricted endpoints.

## Summary

- ADHD uses a single `pLimit` instance in [`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts) to enforce a hard ceiling on parallel LLM requests.
- The `--concurrency` CLI flag is parsed in [`src/cli.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/cli.ts) and typed in [`src/types.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/types.ts), defaulting to `4` when omitted.
- During the **divergence** phase, each `divergeBranch` call is wrapped by the limiter so that no more than `concurrency` frames run at once.
- During the **deepening** phase, the same limiter caps concurrent `deepenIdea` calls, protecting rate limits across both high-traffic stages.
- Because every LLM request flows through the same bounded promise pool, resource consumption stays predictable regardless of how many frames or ideas are scheduled.

## Frequently Asked Questions

### What happens if I do not specify a concurrency value?

If the `concurrency` property is missing from the options, the `run()` function in [`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts) falls back to a default of `4` at lines 19-22. This ensures the engine never runs with an unbounded pool.

### Does the concurrency limit apply to both divergence and deepening?

Yes. The engine creates one shared `pLimit` instance and applies it to both the `divergeBranch` calls in the divergence phase and the `deepenIdea` calls in the deepening phase. This global limit keeps total in-flight requests under the configured cap across all stages.

### Why does ADHD use `p-limit` instead of a simple loop?

`p-limit` provides a bounded promise pool that works cleanly with `Promise.all`, allowing the engine to schedule all frames or ideas concurrently while automatically throttling execution to the desired concurrency level. This avoids manual batching logic and keeps the code in [`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts) concise and maintainable.

### Can I set concurrency higher than the number of frames or ideas?

Yes. Setting a higher value does not cause errors; it simply raises the ceiling. If the number of scheduled tasks is lower than the limit, all tasks run in parallel. If it is higher, the excess tasks wait in the `p-limit` queue until an active slot frees up.