Performance Impact of Running N Isolated Reasoning Processes in ADHD
Running N isolated reasoning processes in ADHD scales linearly—O(N × per_branch) tokens and O(N / concurrency) wall‑clock time—because each branch uses a fresh, stateless Claude Agent SDK session with no shared KV‑cache.
The ADHD ("Automatic Divergent‑Thinking with Heuristic‑Driven pruning") framework generates parallel reasoning branches to explore solution spaces efficiently. Unlike Tree‑of‑Thought approaches that accumulate context, ADHD's branches remain completely isolated. This design choice has specific performance implications that every user should understand before scaling N (the number of reasoning processes).
Linear Token and Compute Cost
Because each branch operates as an independent LLM call through the Claude Agent SDK, there is no shared KV‑cache or message history between them. The architecture overview in documentation/how-it-works.md states:
Each branch consumes
per_branchtokens, so total cost is O(N × per_branch).
This linear scaling applies to both token consumption and underlying inference costs. Whether you run 4 branches or 40, each branch pays the full token price for its reasoning path.
Concurrency Limits and Runtime Scaling
The implementation in src/engine.ts uses a semaphore to control resource utilization:
// src/engine.ts lines 19-21, 49-55
const semaphore = new Semaphore(config.concurrency ?? DEFAULT_CONCURRENCY);
// Each divergeBranch call acquires the semaphore before executing
await semaphore.acquire();
const result = await divergeBranch(problem, config);
semaphore.release();
The default concurrency value is 4. This means:
- At most 4 branches execute simultaneously regardless of total N
- Additional branches queue until a slot opens
- Wall‑clock time follows:
total_time ≈ (N / concurrency) × time_per_branch
Raising concurrency reduces latency if your hardware supports higher parallelism, but increases instantaneous GPU/CPU load and memory pressure.
Memory Characteristics
Memory usage in ADHD is proportional to concurrent sessions, not total branches. With default settings, only 4 isolated sessions occupy memory at any moment. Each session is stateless—created fresh by the wrapper in src/llm.ts—so memory is released immediately after a branch completes.
This contrasts sharply with approaches that maintain all branch contexts simultaneously, which can exhaust memory at high N.
Avoiding Quadratic Cost
ADHD deliberately avoids the quadratic cost seen in in‑context Tree‑of‑Thought methods. As noted in the architecture documentation:
Branches never exchange context during divergence.
Since no information flows between branches until the final aggregation phase, the system cannot suffer from the combinatorial context explosion that plagues interdependent reasoning approaches.
Practical Configuration Examples
TypeScript API
// Run 8 isolated branches with 6 concurrent LLM calls
import { run } from "./src/engine";
await run({
problem: "Design a low‑latency cache for edge services",
framesPerRun: 8, // N = 8 isolated reasoning processes
concurrency: 6, // Parallelism limit
ideasPerFrame: 5,
topK: 3,
});
CLI Usage
# 12 branches, 8 parallel calls
adhd run "Improve database replication latency" \
--frames-per-run 12 \
--concurrency 8
The CLI parser in src/cli.ts (lines 20‑54) validates these flags and passes them to the engine.
Key Implementation Files
| File | Performance‑Critical Role |
|---|---|
src/engine.ts |
Spawns N parallel divergeBranch calls; applies semaphore‑based concurrency limiting |
src/cli.ts |
Exposes --frames-per-run and --concurrency configuration flags |
src/llm.ts |
Creates fresh, stateless Claude Agent SDK sessions per branch |
documentation/how-it-works.md |
Documents linear token scaling and isolation guarantees |
Tuning Guidelines for N Isolated Reasoning Processes
- Token budget fixed? Reduce N directly—no mitigating factors exist.
- Latency‑sensitive? Increase
concurrencyup to hardware limits (monitor GPU utilization). - Memory‑constrained? Lower
concurrencyto cap simultaneous sessions; N can remain high. - Cost‑optimization? Balance N against expected solution quality; more branches improve coverage linearly at linear cost.
Summary
- Token cost scales linearly with N due to isolation—no shared KV‑cache between branches.
- Wall‑clock time scales as N / concurrency; default concurrency of 4 queues excess branches.
- Memory usage depends on
concurrency, not N, because sessions are stateless and short‑lived. - Quadratic cost is eliminated by design—branches never share context during reasoning.
- Tune
framesPerRun(API) or--frames-per-run(CLI) for coverage; tuneconcurrencyfor resource limits.
Frequently Asked Questions
Does increasing N always increase total token cost?
Yes. Each isolated reasoning branch in ADHD consumes independent tokens. According to the source architecture, total token cost is strictly O(N × per_branch) with no reduction mechanisms. If your API budget is constrained, N is the primary control variable.
Why does ADHD use a semaphore instead of running all branches in parallel?
The semaphore in src/engine.ts prevents GPU/CPU oversubscription. Without it, launching 100+ simultaneous LLM calls would exhaust resources and likely trigger rate limits or out‑of‑memory errors. The default concurrency of 4 provides a conservative balance; adjust based on your hardware and API tier.
Can I share context between branches to reduce cost?
No—and this is intentional. The design in src/llm.ts creates fresh Claude Agent SDK sessions for each branch. While shared context could theoretically reduce token usage, it would introduce dependencies that compromise the divergent‑thinking goal and risk quadratic cost growth from context accumulation.
How does ADHD's performance compare to Tree-of-Thought?
ADHD trades inter‑branch communication for predictable linear scaling. Tree‑of‑Thought approaches can incur quadratic costs as contexts grow and merge. ADHD maintains strict isolation, making performance at any N straightforward to predict using the formula total_time ≈ (N / concurrency) × time_per_branch.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →