Performance Considerations for UditAkhourii/adhd: Optimizing Tree-of-Thought LLM Orchestration
The adhd repository implements a Tree-of-Thought engine that fans out dozens of parallel LLM calls, making performance dependent on concurrency limits, frame configuration, and memory management rather than just algorithmic complexity.
The adhd repository by UditAkhourii implements a sophisticated Tree-of-Thought (ToT) orchestration engine for large language models. Because each run fans out into multiple divergent branches, scoring passes, and clustering operations, understanding the performance considerations for UditAkhourii/adhd is critical before running production workloads. This guide examines the specific bottlenecks in the codebase—from the callLLM wrapper in src/llm.ts to the frame selection logic in src/frames.ts—and provides concrete tuning strategies.
Managing LLM Call Concurrency with p-limit
The engine generates parallelism through p-limit, imported in src/engine.ts as import pLimit from "p-limit". Every divergent branch, scoring pass, clustering operation, and deepening step creates separate LLM requests via the callLLM function in src/llm.ts.
Uncontrolled parallelism can overwhelm LLM providers or trigger rate limits. The codebase caps concurrent requests using pLimit(concurrency), configured via the RunOptions.maxParallel flag exposed in the public API in src/index.ts.
To limit parallel requests:
import { run } from "adhd";
await run({
problem: "How to auto-scale my web service?",
maxParallel: 4, // Caps concurrent LLM requests
ideasPerFrame: 5,
frames: 3,
model: "claude-instant", // Cheaper, faster model
});
Optimizing Frame and Idea Counts
The selectFrames function selects a configurable number of frames (defaulting to 5) and requests ideasPerFrame ideas per frame (defaulting to 8). This generates dozens of LLM calls per run, creating quadratic growth in both network time and memory usage.
More frames multiplied by more ideas equals exponential growth in LLM calls, storage requirements for Idea and Branch objects, and downstream processing time. For large problems, reduce RunOptions.frames or RunOptions.ideasPerFrame, or run multiple smaller passes and merge results.
Memory Management for In-Memory Data Structures
All ideas, scores, clusters, and deepened concepts are stored as plain JavaScript objects (defined in src/types.ts) including Idea, Branch, and Score types. With many branches, the process can consume significant RAM, particularly if each idea contains long rationales or sketches.
Mitigate memory pressure by streaming intermediate results to disk (e.g., JSONL format) or pruning aggressively after each scoring step using topK selection to retain only the most promising branches.
import { run, renderText } from "adhd";
import fs from "fs/promises";
const result = await run({ problem: "..." });
await fs.writeFile("ideas.jsonl", JSON.stringify(result.branches));
console.log(renderText(result.finalIdea));
Reducing Validation and Parsing Overhead
Every LLM response undergoes Zod schema validation via the parseJSON function in src/llm.ts. While this ensures type safety, strict validation adds CPU overhead and can abort runs if the schema rejects malformed responses.
For exploratory runs, disable strict validation by setting RunOptions.strict = false:
await run({
problem: "Design a low-latency cache",
strict: false, // Disables Zod validation for speed
maxParallel: 6,
});
UUID Generation and Shuffling Costs
Each idea receives a UUID generated via randomUUID(). While cheap, UUID generation can become a bottleneck when creating millions of ideas in massive runs. Consider switching to sequential counters for short-lived runs or batch-generating UUIDs.
Additionally, the shuffle function (Fisher-Yates algorithm) executes on the entire frame pool each run in src/frames.ts. This O(N) operation is negligible for small frame sets but grows linearly with larger pools. Cache shuffled orders across runs if the frame set remains static.
Model Selection and Network Latency
The model field passes directly to Anthropic's Claude SDK. Larger models like claude-2 incur higher latency and token costs compared to lightweight alternatives. Latency directly adds to overall runtime since the engine waits for network responses.
For exploratory passes, configure RunOptions.model to use "claude-instant" or similar lightweight models. Reserve heavyweight models only for final deepening steps where quality is paramount.
Profiling and Baseline Configuration
The run function exported from src/engine.ts returns a RunResult object containing timing metrics for each phase. Use these metrics to identify the slowest stage before scaling up.
Start small: Begin with low maxParallel, minimal frames, and low-cost models to verify correctness. Scale gradually by increasing parallelism and frame count only after establishing baseline performance.
Summary
- Limit concurrency using
maxParallelinRunOptionsto prevent rate limiting and provider throttling. - Reduce memory footprint by lowering
framesandideasPerFrame, or stream results to disk instead of holding all branches in RAM. - Disable strict validation with
strict: falsefor faster exploratory runs. - Choose models strategically—use lightweight models like
claude-instantfor initial exploration and heavyweight models only for final output. - Profile before scaling by examining the timing metrics in
RunResultto identify bottlenecks in the divergence, scoring, or clustering phases.
Frequently Asked Questions
How does the adhd repository handle concurrent LLM requests?
The engine uses the p-limit library in src/engine.ts to wrap all async LLM calls made via callLLM in src/llm.ts. By default, it respects the RunOptions.maxParallel parameter, which caps the number of simultaneous requests to prevent overwhelming the LLM provider or hitting rate limits. Adjust this value based on your API tier and network capacity.
What causes memory issues in adhd runs and how can I prevent them?
Memory consumption scales with the number of branches because all Idea, Branch, and Score objects (defined in src/types.ts) remain in RAM as plain JavaScript objects. To prevent issues, reduce frames and ideasPerFrame, or implement topK pruning after each scoring step to discard low-quality branches. For massive runs, stream intermediate results to JSONL files on disk instead of maintaining the entire tree in memory.
Can I disable validation to speed up adhd execution?
Yes. By setting strict: false in RunOptions, you disable Zod schema validation in the parseJSON function located in src/llm.ts. This removes CPU overhead from parsing and prevents run abortions due to schema mismatches, though it trades safety for speed. Use this only for exploratory runs where data integrity is less critical.
How do I choose the right model for performance versus quality?
The model parameter passes directly to the Anthropic SDK. For rapid iteration and cost savings, specify lightweight models like "claude-instant" in RunOptions. Reserve larger models such as claude-2 for final deepening steps where output quality matters most. This tiered approach balances speed and accuracy according to the analysis of the codebase.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →