Cost and Latency Implications of Running 5 Frames × 6 Ideas with Critic Passes in ADHD

Running 5 frames with 6 ideas each and full critic passes in ADHD consumes roughly 3S + K LLM calls per run, driving linear token cost and wall-clock latency that typically stays within 2–3× a single LLM call under default concurrency.

The ADHD repository by UditAkhourii/adhd orchestrates divergent reasoning through parallel cognitive frames, and every production deployment must budget for the token and time footprint of a full run. When you execute a configuration with 5 frames, 6 ideas per frame, and the default three-phase critic pipeline—scoring, clustering, and deepening—the framework follows a predictable resource model. Understanding the exact cost and latency implications of running 5 frames × 6 ideas with critic passes ensures you can tune parameters before scaling.

How the ADHD Pipeline Consumes LLM Calls

The divergent phase in src/engine.ts launches S independent cognitive frames. For each frame, the generator produces 6 ideas in a single LLM call via the diverge routine. After generation, three distinct critic passes run per frame:

  • Score — assigns a numeric quality score to each idea (1 call per frame via engine.scoreIdeas)
  • Cluster — groups similar ideas and prunes traps (1 call per frame via engine.clusterIdeas)
  • Deepen — re-queries the top K ideas, defaulting to 3, to elaborate them (K calls via engine.deepenIdeas)

That yields a total call count of:

Total calls = S (diverge) + S (score) + S (cluster) + K (deepen)
            = 3S + K

With the default deepen pool of K = 3, a standard run issues 3S + 3 LLM calls.

Token Cost Breakdown

According to documentation/when-to-use.md, ADHD models divergence cost as:

cost ≈ N × (base_context + branch_output)

In this formula, N equals S, the number of frames. Because each frame reloads the full base context and appends its own output tokens, the 6-idea payload directly inflates the branch_output term. The critic passes add a roughly equivalent linear factor per frame; they reuse the same frame prompt as the base context and emit only a small scoring or clustering payload. The deepen stage adds another linear term scaled by K, since each selected idea is submitted as a separate follow-up call.

Consequences of this model include:

  • Token spend grows linearly with frame count.
  • ideasPerFrame increases output tokens per branch but does not multiply the number of network calls.
  • The dominant cost driver is the number of parallel branches (frames), not the critic pass count alone.

Latency and Parallel Execution

As documented in documentation/how-it-works.md, ADHD executes divergent and critic calls under a configurable semaphore. The default concurrency limit is 4, meaning the runtime parallelizes up to four simultaneous LLM requests.

The expected latency follows this batching pattern:

latency ≈ (⌈S / C⌉ + ⌈S / C⌉ + ⌈K / C⌉) × average_llm_latency

Here C is the concurrency limit. The first term covers the initial diverge wave, the second covers the combined critic passes, and the third covers the deepen stage. In practice, with the default settings—S = 5, K = 3, and C = 4—a full run completes in approximately 2–3× the latency of a single LLM call, because the small critic payloads and pipeline overlap keep the wall-clock time well below the raw call count.

Configuring a Run in Code

You can parameterize a run directly through the main entry point. The snippet below sets up S = 5 frames with 6 ideas each and the default critic pipeline:

import { run } from "./src/index.js";

const result = await run({
  problem: "How can we make our API more resilient?",
  framesPerRun: 5,      // S frames
  ideasPerFrame: 6,     // 6 ideas per frame
  topK: 3,              // ideas to deepen (K)
  // Optional: override the model used for critic passes
  // criticModel: "claude-2.1",
});

Internally, this invokes engine.diverge for the initial generation wave, followed by engine.scoreIdeas, engine.clusterIdeas, and finally engine.deepenIdeas. Each stage is traced through src/llm.ts, which handles token accounting and request dispatch.

Key Source Files

For readers who want to verify the behavior in source, the following files define the pipeline:

Summary

  • A full run with S frames, 6 ideas, and critic passes issues 3S + K LLM calls, where K defaults to 3.
  • Token cost scales linearly with frame count according to the formula in documentation/when-to-use.md.
  • ADHD parallelizes frames and critics via a semaphore with default concurrency = 4.
  • Default wall-clock latency for S = 5 and K = 3 stays near 2–3× a single LLM call.
  • You can tune framesPerRun, ideasPerFrame, topK, and concurrency to balance quality against budget.

Frequently Asked Questions

How many LLM calls does one full run generate?

A run generates 3S + K calls, where S is framesPerRun and K is topK. With the default configuration of 5 frames and a deepen pool of 3, that equals 18 LLM calls per run.

Does increasing ideasPerFrame raise token cost?

Yes. Raising ideasPerFrame increases the branch_output tokens emitted during each diverge call, which directly increases the per-frame token spend. However, it does not increase the number of network calls.

Can I reduce latency without lowering the frame count?

Yes. You can raise the concurrency limit in src/engine.ts or via configuration to allow more simultaneous LLM calls. Alternatively, you can skip the deepen stage or reduce topK to cut the final batch of calls.

Where is the cost model officially documented?

The linear cost model is documented in documentation/when-to-use.md, which states cost ≈ N × (base_context + branch_output). The parallel execution model that governs latency is described in documentation/how-it-works.md.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →