How the ADHD Concurrency Parameter Limits Parallel LLM Calls
The concurrency parameter in ADHD caps parallel LLM calls by routing every request through a single p-limit instance, ensuring that no more than the configured number of promises run simultaneously during divergence and deepening phases.
The ADHD repository by UditAkhourii is an open-source reasoning engine that generates multiple solution branches by executing numerous LLM prompts in parallel. To prevent API rate-limit errors and network congestion, the engine exposes a --concurrency flag that controls exactly how many LLM requests can be in flight at once. Understanding how this concurrency parameter limits parallel LLM calls is essential for tuning performance and staying within provider quotas.
CLI Parsing and Type Definitions
The concurrency journey begins at the command line and flows through the type system before reaching the execution engine.
Parsing the --concurrency Flag in cli.ts
In src/cli.ts, the CLI reads the --concurrency flag and stores it inside the RunOptions object that drives each run. The flag is defined at line 53, making it available to users without touching the source code.
Declaring concurrency in types.ts
The shape of the configuration is declared in src/types.ts at line 55, where the RunOptions interface documents the concurrency property and its default value. This type safety ensures that every downstream consumer expects a numeric limit.
Engine Setup and the Promise Limiter
Once the options reach the core engine, ADHD translates the numeric concurrency value into a hard ceiling on parallel LLM requests.
Destructuring RunOptions with a Fallback
Inside src/engine.ts, the run() function destructures the incoming options and applies a fallback of 4 when no value is supplied. This logic appears at lines 19-22, guaranteeing that the engine always has a bounded concurrency setting even if the caller omits the flag.
Creating the Limiter with p-limit
At line 49 of src/engine.ts, the engine instantiates a single shared limiter by calling pLimit(concurrency). This utility from the p-limit library creates a promise pool that permits only concurrency number of active promises at any moment. All LLM-bound work is then funneled through this one instance.
Applying the Limit to Divergence and Deepening
The limiter is applied uniformly across the two phases that generate the heaviest LLM traffic: diverging branches and deepening ideas.
Capping the Divergence Phase
During divergence, the engine spawns one branch per frame. In src/engine.ts at lines 54-61, each divergeBranch call is wrapped inside limit(async () => …). Even though the code uses Promise.all over the full array of frames, the underlying limiter ensures that at most concurrency branches are active simultaneously.
const limit = pLimit(concurrency); // ← creates the bounded pool
// Divergence: each frame is scheduled through the limiter
const branches = await Promise.all(
frames.map((f) =>
limit(async () => await divergeBranch(problem, context, f, ideasPerFrame, model))
)
);
Capping the Deepening Phase
After the initial ideas are ranked, the top-K candidates enter a deepening phase. In src/engine.ts at lines 98-107, these calls are also routed through the same limit function. Reusing the identical limiter instance means the deepening phase respects the same global ceiling, preventing bursts of concurrent LLM calls when many ideas qualify for elaboration.
// Deepening: top‑K ideas are also limited
const deepened = await Promise.all(
toDeepen.map((idea) =>
limit(async () => await deepenIdea(problem, idea, allIdeas, model))
)
);
Practical Example: Configuring Concurrency
You can tune the limit when invoking the engine programmatically. The following example raises the cap to 8 simultaneous LLM calls:
// Example: run with a custom concurrency of 8
import { run } from "./engine.js";
await run({
problem: "How can we improve cache consistency?",
concurrency: 8, // ← allows up to 8 simultaneous LLM calls
framesPerRun: 6,
ideasPerFrame: 5,
});
By adjusting concurrency, you balance throughput against API quotas. A higher value reduces total wall-clock time for large batches, while a lower value minimizes the risk of rate-limit errors on restricted endpoints.
Summary
- ADHD uses a single
pLimitinstance insrc/engine.tsto enforce a hard ceiling on parallel LLM requests. - The
--concurrencyCLI flag is parsed insrc/cli.tsand typed insrc/types.ts, defaulting to4when omitted. - During the divergence phase, each
divergeBranchcall is wrapped by the limiter so that no more thanconcurrencyframes run at once. - During the deepening phase, the same limiter caps concurrent
deepenIdeacalls, protecting rate limits across both high-traffic stages. - Because every LLM request flows through the same bounded promise pool, resource consumption stays predictable regardless of how many frames or ideas are scheduled.
Frequently Asked Questions
What happens if I do not specify a concurrency value?
If the concurrency property is missing from the options, the run() function in src/engine.ts falls back to a default of 4 at lines 19-22. This ensures the engine never runs with an unbounded pool.
Does the concurrency limit apply to both divergence and deepening?
Yes. The engine creates one shared pLimit instance and applies it to both the divergeBranch calls in the divergence phase and the deepenIdea calls in the deepening phase. This global limit keeps total in-flight requests under the configured cap across all stages.
Why does ADHD use p-limit instead of a simple loop?
p-limit provides a bounded promise pool that works cleanly with Promise.all, allowing the engine to schedule all frames or ideas concurrently while automatically throttling execution to the desired concurrency level. This avoids manual batching logic and keeps the code in src/engine.ts concise and maintainable.
Can I set concurrency higher than the number of frames or ideas?
Yes. Setting a higher value does not cause errors; it simply raises the ceiling. If the number of scheduled tasks is lower than the limit, all tasks run in parallel. If it is higher, the excess tasks wait in the p-limit queue until an active slot frees up.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →