# Performance Considerations for UditAkhourii/adhd: Optimizing Tree-of-Thought LLM Orchestration

> Explore performance considerations for UditAkhourii/adhd. Optimize your Tree-of-Thought LLM orchestration by focusing on concurrency limits, frame configuration, and memory management for peak performance.

- Repository: [Udit Akhouri/adhd](https://github.com/UditAkhourii/adhd)
- Tags: performance
- Published: 2026-07-30

---

**The `adhd` repository implements a Tree-of-Thought engine that fans out dozens of parallel LLM calls, making performance dependent on concurrency limits, frame configuration, and memory management rather than just algorithmic complexity.**

The `adhd` repository by UditAkhourii implements a sophisticated **Tree-of-Thought (ToT)** orchestration engine for large language models. Because each run fans out into multiple divergent branches, scoring passes, and clustering operations, understanding the performance considerations for UditAkhourii/adhd is critical before running production workloads. This guide examines the specific bottlenecks in the codebase—from the `callLLM` wrapper in [`src/llm.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/llm.ts) to the frame selection logic in [`src/frames.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/frames.ts)—and provides concrete tuning strategies.

## Managing LLM Call Concurrency with p-limit

The engine generates parallelism through **p-limit**, imported in [`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts) as `import pLimit from "p-limit"`. Every divergent branch, scoring pass, clustering operation, and deepening step creates separate LLM requests via the `callLLM` function in [`src/llm.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/llm.ts).

Uncontrolled parallelism can overwhelm LLM providers or trigger rate limits. The codebase caps concurrent requests using `pLimit(concurrency)`, configured via the `RunOptions.maxParallel` flag exposed in the public API in [`src/index.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/index.ts).

To limit parallel requests:

```typescript
import { run } from "adhd";

await run({
  problem: "How to auto-scale my web service?",
  maxParallel: 4,            // Caps concurrent LLM requests
  ideasPerFrame: 5,
  frames: 3,
  model: "claude-instant",   // Cheaper, faster model
});

```

## Optimizing Frame and Idea Counts

The `selectFrames` function selects a configurable number of frames (defaulting to 5) and requests `ideasPerFrame` ideas per frame (defaulting to 8). This generates dozens of LLM calls per run, creating quadratic growth in both network time and memory usage.

More frames multiplied by more ideas equals exponential growth in LLM calls, storage requirements for `Idea` and `Branch` objects, and downstream processing time. For large problems, reduce `RunOptions.frames` or `RunOptions.ideasPerFrame`, or run multiple smaller passes and merge results.

## Memory Management for In-Memory Data Structures

All ideas, scores, clusters, and deepened concepts are stored as plain JavaScript objects (defined in [`src/types.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/types.ts)) including `Idea`, `Branch`, and `Score` types. With many branches, the process can consume significant RAM, particularly if each idea contains long rationales or sketches.

Mitigate memory pressure by streaming intermediate results to disk (e.g., JSONL format) or pruning aggressively after each scoring step using `topK` selection to retain only the most promising branches.

```typescript
import { run, renderText } from "adhd";
import fs from "fs/promises";

const result = await run({ problem: "..." });
await fs.writeFile("ideas.jsonl", JSON.stringify(result.branches));
console.log(renderText(result.finalIdea));

```

## Reducing Validation and Parsing Overhead

Every LLM response undergoes **Zod** schema validation via the `parseJSON` function in [`src/llm.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/llm.ts). While this ensures type safety, strict validation adds CPU overhead and can abort runs if the schema rejects malformed responses.

For exploratory runs, disable strict validation by setting `RunOptions.strict = false`:

```typescript
await run({
  problem: "Design a low-latency cache",
  strict: false,               // Disables Zod validation for speed
  maxParallel: 6,
});

```

## UUID Generation and Shuffling Costs

Each idea receives a UUID generated via `randomUUID()`. While cheap, UUID generation can become a bottleneck when creating millions of ideas in massive runs. Consider switching to sequential counters for short-lived runs or batch-generating UUIDs.

Additionally, the `shuffle` function (Fisher-Yates algorithm) executes on the entire frame pool each run in [`src/frames.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/frames.ts). This O(N) operation is negligible for small frame sets but grows linearly with larger pools. Cache shuffled orders across runs if the frame set remains static.

## Model Selection and Network Latency

The `model` field passes directly to Anthropic's Claude SDK. Larger models like `claude-2` incur higher latency and token costs compared to lightweight alternatives. Latency directly adds to overall runtime since the engine waits for network responses.

For exploratory passes, configure `RunOptions.model` to use `"claude-instant"` or similar lightweight models. Reserve heavyweight models only for final deepening steps where quality is paramount.

## Profiling and Baseline Configuration

The `run` function exported from [`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts) returns a `RunResult` object containing timing metrics for each phase. Use these metrics to identify the slowest stage before scaling up.

**Start small:** Begin with low `maxParallel`, minimal frames, and low-cost models to verify correctness. Scale gradually by increasing parallelism and frame count only after establishing baseline performance.

## Summary

- **Limit concurrency** using `maxParallel` in `RunOptions` to prevent rate limiting and provider throttling.
- **Reduce memory footprint** by lowering `frames` and `ideasPerFrame`, or stream results to disk instead of holding all branches in RAM.
- **Disable strict validation** with `strict: false` for faster exploratory runs.
- **Choose models strategically**—use lightweight models like `claude-instant` for initial exploration and heavyweight models only for final output.
- **Profile before scaling** by examining the timing metrics in `RunResult` to identify bottlenecks in the divergence, scoring, or clustering phases.

## Frequently Asked Questions

### How does the adhd repository handle concurrent LLM requests?

The engine uses the **p-limit** library in [`src/engine.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/engine.ts) to wrap all async LLM calls made via `callLLM` in [`src/llm.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/llm.ts). By default, it respects the `RunOptions.maxParallel` parameter, which caps the number of simultaneous requests to prevent overwhelming the LLM provider or hitting rate limits. Adjust this value based on your API tier and network capacity.

### What causes memory issues in adhd runs and how can I prevent them?

Memory consumption scales with the number of branches because all `Idea`, `Branch`, and `Score` objects (defined in [`src/types.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/types.ts)) remain in RAM as plain JavaScript objects. To prevent issues, reduce `frames` and `ideasPerFrame`, or implement `topK` pruning after each scoring step to discard low-quality branches. For massive runs, stream intermediate results to JSONL files on disk instead of maintaining the entire tree in memory.

### Can I disable validation to speed up adhd execution?

Yes. By setting `strict: false` in `RunOptions`, you disable Zod schema validation in the `parseJSON` function located in [`src/llm.ts`](https://github.com/UditAkhourii/adhd/blob/main/src/llm.ts). This removes CPU overhead from parsing and prevents run abortions due to schema mismatches, though it trades safety for speed. Use this only for exploratory runs where data integrity is less critical.

### How do I choose the right model for performance versus quality?

The `model` parameter passes directly to the Anthropic SDK. For rapid iteration and cost savings, specify lightweight models like `"claude-instant"` in `RunOptions`. Reserve larger models such as `claude-2` for final deepening steps where output quality matters most. This tiered approach balances speed and accuracy according to the analysis of the codebase.