# Performance Tradeoffs of Running More Parallel Frames in ADHD: Latency, Cost, and Quality Analysis

> Explore the performance tradeoffs of running more parallel frames in ADHD. Analyze latency, cost, and quality impacts to optimize your setup for better idea diversity and avoid rate limits.

- Repository: [Udit Akhouri/adhd](https://github.com/UditAkhourii/adhd)
- Tags: performance
- Published: 2026-07-30

---

**Increasing parallel frames in ADHD boosts idea diversity but linearly increases API costs and latency while risking rate limits, with diminishing returns beyond the default configuration.**

The ADHD repository implements a fan-out orchestration engine that generates creative solutions by running multiple independent cognitive "frames" in parallel. Understanding the performance tradeoffs of running more parallel frames in ADHD is essential for optimizing the balance between idea quality and computational resources. The `framesPerRun` parameter directly controls how many frames execute simultaneously, affecting everything from wall-clock latency to your LLM API bill.

## How Parallel Frame Execution Works in ADHD

ADHD's architecture treats each "frame" as an independent cognitive perspective that generates ideas through LLM calls. The system fans out these frames using JavaScript's asynchronous concurrency patterns.

### Frame Selection and Shuffling

In [`src/frames.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/frames.ts), the `selectFrames` function builds a randomized pool of cognitive frames before execution. It filters the global `FRAMES` catalog based on mode (code vs. general), shuffles the results, and guarantees at least one "wild-card" frame for creative diversity.

```typescript
export function selectFrames(n: number, codeMode = true): Frame[] {
  const pool = codeMode
    ? FRAMES.filter((f) => f.tags.includes("code") || f.tags.includes("design"))
    : [...FRAMES];
  const wild = FRAMES.filter((f) => f.tags.includes("wild"));
  const shuffled = shuffle(pool);
  const picked = shuffled.slice(0, Math.max(1, n - 1));
  const wildPick = wild[Math.floor(Math.random() * wild.length)];
  if (!picked.find((f) => f.id === wildPick.id)) picked.push(wildPick);
  return picked.slice(0, n);
}

```

Increasing the `n` parameter (which maps to `framesPerRun`) simply returns a larger slice from this shuffled pool, directly scaling the number of parallel tasks submitted to the execution engine.

### Concurrency Control with Semaphores

The actual parallel execution occurs in [`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts) using `pLimit` to cap simultaneous LLM requests. The engine maps each selected frame to a promise respecting the `concurrency` semaphore (default `4`).

```typescript
const frames = selectFrames(framesPerRun, codeMode);
const limit = pLimit(concurrency);
const branches = await Promise.all(
  frames.map((f) =>
    limit(async () => {
      onEvent?.({ kind: "frame:start", frameId: f.id, frameLabel: f.label });
      const b = await divergeBranch(divergeProblem, context, f, ideasPerFrame, model);
      onEvent?.({ kind: "frame:done", frameId: f.id, count: b.ideas.length });
      return b;
    })
  )
);

```

## The Five Critical Performance Tradeoffs

Running more parallel frames creates tension across five operational dimensions. Each additional frame beyond the default (`5`) amplifies these effects according to the source code implementation.

### Throughput and Idea Diversity

**More frames increase raw idea throughput.** Each frame generates multiple ideas via `divergeBranch` (default `ideasPerFrame = 6`), so total LLM calls equal `framesPerRun × ideasPerFrame`. At defaults, this yields 30 calls per run; doubling frames to 10 yields 60 calls.

This fan-out improves coverage of the solution space by introducing more cognitive perspectives. The random shuffle in `selectFrames` ensures diverse frame combinations across runs, beneficial for ambiguous problems requiring broad exploration.

### Latency and Wall-Clock Time

**Latency increases if concurrency remains fixed.** The `Promise.all` waits for all frames to complete, but the semaphore (`concurrency = 4`) queues excess frames. With default settings, 5 frames execute as 4 parallel + 1 queued, but setting `framesPerRun = 12` creates a significant backlog.

Each frame executes a full LLM request-response cycle for every idea generation. Without raising the `concurrency` limit, wall-clock time grows linearly with the queue depth.

### API Cost and Resource Consumption

**Costs scale linearly with frame count.** Each frame invocation triggers `divergeBranch`, which makes multiple LLM API calls. Memory consumption also rises to store intermediate ideas from all branches before the scoring phase.

At maximum throughput, high `framesPerRun` values can consume substantial API quota quickly. The total token count includes not just generation costs but also the context window overhead for each frame's prompt variations.

### Diminishing Returns in Final Quality

**Extra frames often add noise rather than signal.** After generation, the engine runs a **scoring** and **clustering** phase that selects a `topK` shortlist from the complete idea pool. Once the pool reaches sufficient diversity, additional frames produce overlapping or low-scoring ideas that get filtered out during clustering.

The optimal frame count depends on problem ambiguity. Highly constrained problems saturate early, while open-ended "wild" problems benefit from broader exploration.

### Rate Limit and Throttling Risks

**Burst requests trigger provider limits.** While the `concurrency` parameter caps simultaneous connections to `4` by default, the total request volume still scales with `framesPerRun`. LLM providers typically implement per-second or per-minute rate limits.

A configuration of `framesPerRun = 20` with `ideasPerFrame = 6` generates 120 API calls in rapid succession. Without exponential backoff or request spacing, this burst pattern risks HTTP 429 errors or temporary bans from API providers.

## Practical Tuning Guidelines

Adjust `framesPerRun` based on operational constraints rather than maximizing diversity.

- **Increase frames** when exploring highly ambiguous problems with generous API quotas. Values of `8`–`12` provide meaningful coverage gains for complex design challenges.
- **Maintain defaults** (`5` frames) for time-sensitive executions or budget-constrained environments. The scoring/clustering pipeline already filters effectively, making excessive frames redundant.
- **Scale concurrency proportionally** when raising `framesPerRun`. If increasing to `10` or `12` frames, raise `concurrency` to `8` or `12` to prevent queuing delays, but verify your provider's rate limits first.

### Configuration Example

The following configuration optimizes for high-throughput exploration while managing latency through increased concurrency:

```typescript
import { run } from "./engine.js";

await run({
  problem: "Design a low‑latency chat overlay for live video.",
  context: "",
  framesPerRun: 12,   // ↑ from default 5
  ideasPerFrame: 5,
  topK: 4,
  concurrency: 8,     // ↑ to keep latency reasonable
  model: "gpt-4o-mini",
  criticModel: "gpt-4o",
});

```

## Summary

- **Frame scaling**: The `framesPerRun` parameter in [`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts) controls parallel execution breadth, defaulting to `5` frames per run.
- **Concurrency bottleneck**: The `pLimit` semaphore caps simultaneous LLM calls at `4` by default; exceeding this creates serial queuing that increases latency.
- **Cost formula**: Total API calls equal `framesPerRun × ideasPerFrame` (default: 30 calls), scaling linearly with frame count.
- **Quality ceiling**: The scoring and clustering pipeline filters ideas to `topK`, causing diminishing returns as frames add overlapping suggestions.
- **Rate limit exposure**: High frame counts generate API request bursts that risk throttling from LLM providers despite concurrency controls.

## Frequently Asked Questions

### How does increasing framesPerRun affect API costs in ADHD?

API costs scale linearly with `framesPerRun` because each frame invokes `divergeBranch`, which makes `ideasPerFrame` LLM calls (default `6`). Raising frames from `5` to `10` doubles the total requests from 30 to 60 per execution. Token costs include both generation and the context window overhead for each frame's prompt variations.

### What is the optimal concurrency setting when raising framesPerRun?

When increasing `framesPerRun` above the default `5`, set `concurrency` to at least half the frame count to prevent excessive queuing. For `framesPerRun = 12`, use `concurrency = 8` or `12`. However, verify your LLM provider's per-second rate limits first, as higher concurrency increases burst request density and throttling risk.

### Why do additional frames show diminishing returns for idea quality?

Additional frames yield diminishing returns because the `run` function implements a robust filtering pipeline. After parallel generation, the engine **scores** and **clusters** all ideas before selecting a `topK` shortlist. Once the idea pool reaches sufficient diversity, extra frames contribute overlapping or low-scoring suggestions that get filtered out, adding noise rather than signal to the final output.

### Where is parallel frame execution implemented in the ADHD codebase?

Parallel execution is implemented in [`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts) between lines 52-61, where the engine uses `Promise.all` combined with `pLimit(concurrency)` to manage asynchronous frame tasks. The `selectFrames` function in [`src/frames.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/frames.ts) (lines 34-37) prepares the randomized frame pool, while `divergeBranch` handles the per-frame LLM calls.