How the p-queue Library Manages Concurrent Requests in deepwiki-mcp
The p-queue library caps simultaneous HTTP requests at 5 by default, queues pending fetches, and provides real-time metrics for monitoring crawl progress in deepwiki-mcp.
The deepwiki-mcp project implements a breadth-first web crawler that must balance speed against server load. To manage this, the project leverages the p-queue library in src/lib/httpCrawler.ts to orchestrate concurrent network requests without overwhelming target hosts.
Why Concurrency Control Matters for Web Crawling
Uncontrolled parallel HTTP requests can exhaust system resources and trigger rate limits on target servers. The deepwiki-mcp crawler addresses this by treating each URL fetch as a discrete task that must compete for limited execution slots. This pattern ensures politeness toward remote hosts while maximizing throughput within safe bounds.
How p-queue Integrates with deepwiki-mcp
The integration centers on a single PQueue instance configured at module initialization and shared across the crawl lifecycle.
Configuring the Concurrency Limit
In src/lib/httpCrawler.ts, the crawler instantiates PQueue with a configurable concurrency ceiling. The default value is 5 simultaneous requests, though operators can override this via the DEEPWIKI_CONCURRENCY environment variable.
import PQueue from 'p-queue'
const MAX_CONCURRENCY = Number(process.env.DEEPWIKI_CONCURRENCY ?? 5)
const queue = new PQueue({ concurrency: MAX_CONCURRENCY })
Scheduling Fetch Tasks with queue.add()
Each discovered URL becomes a promise-wrapped task submitted to the queue. The crawler uses queue.add() to schedule the actual fetch operation, HTML parsing, and link extraction. This method returns a promise that resolves when the task completes, allowing the crawler to handle results asynchronously.
queue.add(async () => {
const response = await fetch(url, { dispatcher: agent })
if (!response.headers.get('content-type')?.includes('text/html')) {
return
}
const html = await response.text()
// Parse HTML and extract links for further crawling
})
Monitoring Back-Pressure and Progress
The p-queue instance exposes queue.size (tasks waiting to start) and queue.pending (tasks currently executing). The crawler emits these metrics through progress events, giving callers real-time visibility into system load.
emit({
type: 'progress',
url: url.href,
queued: queue.size + queue.pending,
// Additional crawl statistics...
})
Graceful Shutdown with onIdle()
To ensure all pending requests complete before the process exits, the crawler awaits queue.onIdle(). This promise resolves only when the queue is empty and no tasks remain in flight, preventing data loss and ensuring clean resource cleanup.
await queue.onIdle()
Implementation Example: Complete Crawler Setup
The following pattern demonstrates how deepwiki-mcp orchestrates the entire flow:
import PQueue from 'p-queue'
const MAX_CONCURRENCY = Number(process.env.DEEPWIKI_CONCURRENCY ?? 5)
const queue = new PQueue({ concurrency: MAX_CONCURRENCY })
// Seed the queue with initial URLs
queue.add(async () => {
// Fetch and parse logic
})
// Wait for completion
await queue.onIdle()
Summary
- p-queue provides the concurrency primitive that prevents deepwiki-mcp from overwhelming target servers.
- The crawler configures a default limit of 5 simultaneous requests, configurable via
DEEPWIKI_CONCURRENCY. - Each URL fetch is wrapped in
queue.add(), enabling automatic scheduling and back-pressure handling. - Real-time metrics (
queue.size,queue.pending) feed progress reporting throughout the crawl lifecycle. queue.onIdle()ensures graceful shutdown by waiting for all in-flight requests to complete.
Frequently Asked Questions
How does p-queue prevent rate limiting in deepwiki-mcp?
By capping concurrent execution to 5 simultaneous requests by default, p-queue ensures the crawler never opens more connections than configured. This throttling prevents the rapid-fire request patterns that typically trigger server-side rate limits or IP bans.
Can I adjust the concurrency limit without modifying source code?
Yes. Set the DEEPWIKI_CONCURRENCY environment variable before starting the crawler. The httpCrawler.ts module reads this value at initialization and passes it to the PQueue constructor, overriding the default value of 5.
What happens to pending requests when the crawler finishes?
The crawler calls await queue.onIdle() before exiting, which blocks until both queue.size (waiting tasks) and queue.pending (running tasks) reach zero. This guarantees that every scheduled fetch completes and that no promises are left unresolved.
How does the crawler report current queue load?
During execution, the crawler emits progress events containing queue.size + queue.pending to indicate total backlog. These metrics allow monitoring tools to visualize back-pressure and verify that the concurrency limit is effectively throttling the crawl rate.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →