ADHD Performance Characteristics: Latency, Token Usage, and Concurrency Impact
ADHD typically completes a run in 30–90 seconds using 4 concurrent LLM branches, but token costs scale linearly with frame count because each isolated branch reloads the full system prompt substrate.
The UditAkhourii/adhd repository implements a divergent thinking engine for LLMs that trades computational overhead for reasoning quality. Understanding its performance characteristics—specifically how latency, token usage, and concurrency interact—is essential for cost estimation and production deployment. The architecture intentionally isolates reasoning branches to guarantee divergent output, which fundamentally shapes its resource consumption model.
How ADHD Manages Parallel Execution
The ADHD engine orchestrates multiple independent LLM sessions to explore a problem space simultaneously. Each branch operates as a clean-slate conversation, ensuring that ideas develop without cross-contamination from sibling branches.
The Branch Isolation Model
In src/engine.ts, the orchestration logic spawns isolated LLM calls through a concurrency semaphore. Because every branch is an independent session, the first token of any response cannot be generated until the entire system prompt (referred to as the "substrate") has been transmitted. This architectural decision means latency is dominated by the number of parallel branches rather than the length of the generated ideas themselves. A default run issues approximately 10 LLM calls: 5 for divergence, 1 for scoring, 1 for clustering, and 3 for deepening.
Default Concurrency Settings
The default concurrency limit is 4, defined in src/types.ts within the LLMOptions interface and exposed via the CLI in src/cli.ts. This setting controls the maximum number of simultaneous callLLM invocations allowed by the semaphore. The implementation in src/engine.ts manages this queue, ensuring that exceeding the limit waits for in-flight requests to complete before spawning new ones.
Latency Characteristics in ADHD
Wall-clock latency for a standard run ranges from 30 to 90 seconds when using the default concurrency of 4. This duration encompasses the entire pipeline: divergent generation across multiple frames, scoring of candidates, clustering of similar ideas, and deepening of the top survivors.
Latency scales inversely with concurrency but linearly with frame count. With concurrency = 1, execution stretches to roughly 2–3× longer because branches must execute sequentially. Conversely, increasing concurrency beyond 4 yields diminishing returns due to API rate limits and network bandwidth constraints, offering only marginal latency improvements while increasing the risk of throttling.
Token Usage and Cost Model
Call count is a misleading metric for ADHD expenses. The critical cost driver is the substrate multiplier: each branch reloads the complete base context including system prompts, host-session state, and tool definitions.
If the base substrate measures approximately 26,000 tokens, then N frames consume N × 26,000 tokens before any novel content is generated. The total token formula is:
total_tokens ≈ N × base_context_tokens + Σ(idea_tokens)
Here, idea_tokens represent the actual generated content and typically constitute a small fraction of the total cost. When ADHD operates as a skill within a larger agent, the host's context is appended to each branch, further increasing the multiplier.
Calculating Total Token Costs
To estimate costs programmatically, measure your substrate size and apply the linear model:
const baseTokens = 26_000; // Size of system prompt substrate
const N = 5; // Number of frames (default)
const ideaTokens = 200; // Average tokens per generated idea
const total = N * baseTokens + (N * ideaTokens);
console.log(`≈ ${total.toLocaleString()} tokens`);
// Output: ≈ 131,000 tokens for default settings
Concurrency Impact on Performance
Concurrency controls parallelism without altering the total token expenditure. Each branch still pays the full substrate cost regardless of execution order.
| Setting | Wall-Clock Latency | Token Cost |
|---|---|---|
concurrency = 4 (default) |
30–90 seconds | N × base_context + idea_tokens |
concurrency = 1 |
~2–3× longer | Identical to default |
concurrency = 8 |
Slightly faster (API limits apply) | Identical to default |
The trade-off is strictly between speed and API load. Higher concurrency reduces latency but increases instantaneous API request volume, while lower concurrency stretches execution time while keeping the identical token budget.
Measuring and Configuring Performance
CLI Configuration
Control concurrency directly from the command line using the flag defined in src/cli.ts:
# Default execution: 5 frames, concurrency 4
adhd "design a retry/timeout strategy for an LLM-calling CLI"
# Reduced concurrency: slower execution, identical token cost
adhd --concurrency 2 "design a retry/timeout strategy for an LLM-calling CLI"
Programmatic Control
For TypeScript implementations, pass performance parameters to the run function:
import { run } from "adhd";
await run({
problem: "design a retry/timeout strategy for an LLM-calling CLI",
frames: 5, // Number of parallel divergent branches (N)
ideasPerFrame: 6, // Ideas generated per frame (k)
topK: 3, // Survivors selected for deepening (K)
concurrency: 8, // Max parallel LLM calls
});
Benchmarking Latency
Measure execution time using standard Node.js timing methods:
console.time("ADHD");
await run({
problem: "Optimize database query performance",
concurrency: 4
});
console.timeEnd("ADHD");
// Prints: ADHD: 45.234s (typical range 30-90s)
For detailed implementation of the parallel execution logic, refer to src/engine.ts. The token cost model and isolation rationale are documented in documentation/when-to-use.md and documentation/how-it-works.md, while generated API documentation appears in docs/index.html.
Summary
- Latency: Default runs complete in 30–90 seconds with concurrency set to 4, scaling linearly with frame count and inversely with concurrency limits.
- Token Usage: Costs follow the formula
N × base_context_tokens, making the substrate multiplier the dominant expense factor rather than generated idea length. - Concurrency: Adjusting the concurrency parameter in
src/types.tsor via CLI changes wall-clock latency but leaves total token consumption unchanged. - Architecture: Branch isolation in
src/engine.tsguarantees divergent reasoning by reloading the full system prompt for every parallel branch.
Frequently Asked Questions
Does increasing concurrency reduce token costs in ADHD?
No. Increasing concurrency reduces wall-clock latency by allowing more branches to execute simultaneously, but the total token count remains identical. Each branch still transmits the full substrate context regardless of execution order, so the cost model N × base_context_tokens stays constant.
Why does ADHD reload the full system prompt for every branch?
The architecture enforces branch isolation to guarantee truly divergent reasoning. By treating each frame as an independent LLM session with no shared context beyond the initial substrate, the system prevents early ideas from biasing subsequent generation. This isolation is implemented in src/engine.ts and documented in documentation/how-it-works.md.
How can I estimate token usage before running ADHD?
Calculate the substrate size of your system prompt (including tool definitions and host context), then multiply by the number of frames (N). Add a small buffer for generated idea tokens. For example, with a 26,000-token substrate and 5 frames, expect approximately 130,000 input tokens plus output tokens before execution.
What is the optimal concurrency setting for ADHD?
The default value of 4 balances latency against API rate limits for most use cases. Settings above 4 offer minimal latency improvements due to network bandwidth and provider throttling, while settings below 4 linearly increase execution time. Adjust based on your API tier's rate limits and your tolerance for wall-clock duration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →