Performance Tradeoffs of Running More Parallel Frames in ADHD: Latency, Cost, and Quality Analysis

Increasing parallel frames in ADHD boosts idea diversity but linearly increases API costs and latency while risking rate limits, with diminishing returns beyond the default configuration.

The ADHD repository implements a fan-out orchestration engine that generates creative solutions by running multiple independent cognitive "frames" in parallel. Understanding the performance tradeoffs of running more parallel frames in ADHD is essential for optimizing the balance between idea quality and computational resources. The framesPerRun parameter directly controls how many frames execute simultaneously, affecting everything from wall-clock latency to your LLM API bill.

How Parallel Frame Execution Works in ADHD

ADHD's architecture treats each "frame" as an independent cognitive perspective that generates ideas through LLM calls. The system fans out these frames using JavaScript's asynchronous concurrency patterns.

Frame Selection and Shuffling

In src/frames.ts, the selectFrames function builds a randomized pool of cognitive frames before execution. It filters the global FRAMES catalog based on mode (code vs. general), shuffles the results, and guarantees at least one "wild-card" frame for creative diversity.

export function selectFrames(n: number, codeMode = true): Frame[] {
  const pool = codeMode
    ? FRAMES.filter((f) => f.tags.includes("code") || f.tags.includes("design"))
    : [...FRAMES];
  const wild = FRAMES.filter((f) => f.tags.includes("wild"));
  const shuffled = shuffle(pool);
  const picked = shuffled.slice(0, Math.max(1, n - 1));
  const wildPick = wild[Math.floor(Math.random() * wild.length)];
  if (!picked.find((f) => f.id === wildPick.id)) picked.push(wildPick);
  return picked.slice(0, n);
}

Increasing the n parameter (which maps to framesPerRun) simply returns a larger slice from this shuffled pool, directly scaling the number of parallel tasks submitted to the execution engine.

Concurrency Control with Semaphores

The actual parallel execution occurs in src/engine.ts using pLimit to cap simultaneous LLM requests. The engine maps each selected frame to a promise respecting the concurrency semaphore (default 4).

const frames = selectFrames(framesPerRun, codeMode);
const limit = pLimit(concurrency);
const branches = await Promise.all(
  frames.map((f) =>
    limit(async () => {
      onEvent?.({ kind: "frame:start", frameId: f.id, frameLabel: f.label });
      const b = await divergeBranch(divergeProblem, context, f, ideasPerFrame, model);
      onEvent?.({ kind: "frame:done", frameId: f.id, count: b.ideas.length });
      return b;
    })
  )
);

The Five Critical Performance Tradeoffs

Running more parallel frames creates tension across five operational dimensions. Each additional frame beyond the default (5) amplifies these effects according to the source code implementation.

Throughput and Idea Diversity

More frames increase raw idea throughput. Each frame generates multiple ideas via divergeBranch (default ideasPerFrame = 6), so total LLM calls equal framesPerRun × ideasPerFrame. At defaults, this yields 30 calls per run; doubling frames to 10 yields 60 calls.

This fan-out improves coverage of the solution space by introducing more cognitive perspectives. The random shuffle in selectFrames ensures diverse frame combinations across runs, beneficial for ambiguous problems requiring broad exploration.

Latency and Wall-Clock Time

Latency increases if concurrency remains fixed. The Promise.all waits for all frames to complete, but the semaphore (concurrency = 4) queues excess frames. With default settings, 5 frames execute as 4 parallel + 1 queued, but setting framesPerRun = 12 creates a significant backlog.

Each frame executes a full LLM request-response cycle for every idea generation. Without raising the concurrency limit, wall-clock time grows linearly with the queue depth.

API Cost and Resource Consumption

Costs scale linearly with frame count. Each frame invocation triggers divergeBranch, which makes multiple LLM API calls. Memory consumption also rises to store intermediate ideas from all branches before the scoring phase.

At maximum throughput, high framesPerRun values can consume substantial API quota quickly. The total token count includes not just generation costs but also the context window overhead for each frame's prompt variations.

Diminishing Returns in Final Quality

Extra frames often add noise rather than signal. After generation, the engine runs a scoring and clustering phase that selects a topK shortlist from the complete idea pool. Once the pool reaches sufficient diversity, additional frames produce overlapping or low-scoring ideas that get filtered out during clustering.

The optimal frame count depends on problem ambiguity. Highly constrained problems saturate early, while open-ended "wild" problems benefit from broader exploration.

Rate Limit and Throttling Risks

Burst requests trigger provider limits. While the concurrency parameter caps simultaneous connections to 4 by default, the total request volume still scales with framesPerRun. LLM providers typically implement per-second or per-minute rate limits.

A configuration of framesPerRun = 20 with ideasPerFrame = 6 generates 120 API calls in rapid succession. Without exponential backoff or request spacing, this burst pattern risks HTTP 429 errors or temporary bans from API providers.

Practical Tuning Guidelines

Adjust framesPerRun based on operational constraints rather than maximizing diversity.

  • Increase frames when exploring highly ambiguous problems with generous API quotas. Values of 8–12 provide meaningful coverage gains for complex design challenges.
  • Maintain defaults (5 frames) for time-sensitive executions or budget-constrained environments. The scoring/clustering pipeline already filters effectively, making excessive frames redundant.
  • Scale concurrency proportionally when raising framesPerRun. If increasing to 10 or 12 frames, raise concurrency to 8 or 12 to prevent queuing delays, but verify your provider's rate limits first.

Configuration Example

The following configuration optimizes for high-throughput exploration while managing latency through increased concurrency:

import { run } from "./engine.js";

await run({
  problem: "Design a low‑latency chat overlay for live video.",
  context: "",
  framesPerRun: 12,   // ↑ from default 5
  ideasPerFrame: 5,
  topK: 4,
  concurrency: 8,     // ↑ to keep latency reasonable
  model: "gpt-4o-mini",
  criticModel: "gpt-4o",
});

Summary

  • Frame scaling: The framesPerRun parameter in src/engine.ts controls parallel execution breadth, defaulting to 5 frames per run.
  • Concurrency bottleneck: The pLimit semaphore caps simultaneous LLM calls at 4 by default; exceeding this creates serial queuing that increases latency.
  • Cost formula: Total API calls equal framesPerRun × ideasPerFrame (default: 30 calls), scaling linearly with frame count.
  • Quality ceiling: The scoring and clustering pipeline filters ideas to topK, causing diminishing returns as frames add overlapping suggestions.
  • Rate limit exposure: High frame counts generate API request bursts that risk throttling from LLM providers despite concurrency controls.

Frequently Asked Questions

How does increasing framesPerRun affect API costs in ADHD?

API costs scale linearly with framesPerRun because each frame invokes divergeBranch, which makes ideasPerFrame LLM calls (default 6). Raising frames from 5 to 10 doubles the total requests from 30 to 60 per execution. Token costs include both generation and the context window overhead for each frame's prompt variations.

What is the optimal concurrency setting when raising framesPerRun?

When increasing framesPerRun above the default 5, set concurrency to at least half the frame count to prevent excessive queuing. For framesPerRun = 12, use concurrency = 8 or 12. However, verify your LLM provider's per-second rate limits first, as higher concurrency increases burst request density and throttling risk.

Why do additional frames show diminishing returns for idea quality?

Additional frames yield diminishing returns because the run function implements a robust filtering pipeline. After parallel generation, the engine scores and clusters all ideas before selecting a topK shortlist. Once the idea pool reaches sufficient diversity, extra frames contribute overlapping or low-scoring suggestions that get filtered out, adding noise rather than signal to the final output.

Where is parallel frame execution implemented in the ADHD codebase?

Parallel execution is implemented in src/engine.ts between lines 52-61, where the engine uses Promise.all combined with pLimit(concurrency) to manage asynchronous frame tasks. The selectFrames function in src/frames.ts (lines 34-37) prepares the randomized frame pool, while divergeBranch handles the per-frame LLM calls.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →