How the deepwiki-mcp Crawler Implements Rate Limiting and Respects robots.txt
The deepwiki-mcp crawler uses p-queue to limit concurrent requests to 5 by default, implements exponential backoff retries up to 3 attempts, and enforces robots.txt rules using robots-parser to filter disallowed URLs before fetching.
The regenrek/deepwiki-mcp repository provides a Model Context Protocol (MCP) server designed to crawl documentation sites and convert them to Markdown. Understanding how this tool manages request rates and respects site policies is essential for ethical web scraping. This article examines the implementation details in src/lib/httpCrawler.ts to explain exactly how the deepwiki-mcp crawler handles rate limiting and robots.txt compliance.
Core Rate Limiting Mechanisms in deepwiki-mcp
Concurrency Control with p-queue
The primary rate limiting mechanism relies on the p-queue library to cap simultaneous connections. In src/lib/httpCrawler.ts, the crawler initializes a queue with a default MAX_CONCURRENCY of 5:
const queue = new PQueue({ concurrency: MAX_CONCURRENCY })
Every URL fetch operation is wrapped in queue.add(), ensuring that no more than five HTTP requests execute simultaneously. The queue guarantees that excess requests wait until slots become available, throttling the overall request rate to prevent overwhelming target servers.
Configurable Concurrency via Environment Variables
Users can adjust crawling speed without modifying source code. The crawler checks for the DEEPWIKI_CONCURRENCY environment variable at startup (src/lib/httpCrawler.ts:L10-L12):
export DEEPWIKI_CONCURRENCY=10
When set, this value overrides the default limit of 5, allowing up to ten simultaneous requests. This flexibility enables adaptation to different network conditions or server capabilities.
Exponential Backoff for Retry Logic
Transient failures trigger an exponential backoff strategy rather than immediate retries. The implementation defines RETRY_LIMIT as 3 attempts and BACKOFF_BASE_MS as 250 milliseconds. When a request fails, the crawler waits according to the formula (src/lib/httpCrawler.ts:L78-L82):
await setTimeout(BACKOFF_BASE_MS * 2 ** (retries - 1))
This yields wait times of 250ms, 500ms, and 1000ms for successive retries, mitigating temporary spikes or transient failures without flooding the target host with aggressive retry attempts.
robots.txt Compliance Implementation
Fetching and Parsing robots.txt
The crawler enforces site policies by pre-fetching and parsing the root domain's robots.txt file. Using the robots-parser library, it constructs a ruleset before crawling begins (src/lib/httpCrawler.ts:L42-L52):
const robotsUrl = new URL('/robots.txt', root)
const res = await fetch(robotsUrl)
robots = robotsParser(robotsUrl.href, body)
This parser instance evaluates crawl rules against user-agent strings and path patterns defined by the site owner, creating a filter for subsequent URL processing.
URL Filtering Against Crawl Rules
Every candidate URL undergoes a permission check before entering the fetch queue. The crawler validates URLs using the robots.isAllowed() method with a wildcard user-agent (src/lib/httpCrawler.ts:L30-L32):
if (robots && !robots.isAllowed(url.href, '*')) return
If robots.txt explicitly disallows the path—such as Disallow: /private/ or Disallow: /admin/—the crawler immediately excludes the URL from the queue, ensuring compliance with site restrictions.
Fail-Open Safety Mechanism
When robots.txt cannot be fetched due to network errors, 404 responses, or server failures, the crawler implements a fail-open strategy. It proceeds with crawling without rule enforcement rather than blocking all access. This behavior follows standard web crawler conventions, allowing documentation gathering to continue when robots.txt is temporarily unavailable while maintaining safety through other rate limiting mechanisms.
Practical Usage Examples
Basic crawl with default rate limiting
import { crawl } from '@/lib/httpCrawler'
import { URL } from 'node:url'
await crawl({
root: new URL('https://deepwiki.com/owner/repo'),
maxDepth: 1,
emit: () => {}, // progress callback (optional)
verbose: true,
})
This configuration runs with the default concurrency of 5 requests and automatically fetches and respects the site's robots.txt rules.
Override concurrency via environment variable
export DEEPWIKI_CONCURRENCY=10
node myScript.js
Setting DEEPWIKI_CONCURRENCY to 10 increases the limit to ten simultaneous requests, as read from process.env.DEEPWIKI_CONCURRENCY (src/lib/httpCrawler.ts:L10-L12).
Inspect robots.txt handling
import { crawl } from '@/lib/httpCrawler'
await crawl({
root: new URL('https://deepwiki.com/some/project'),
maxDepth: 1,
emit: () => {},
verbose: false,
})
The crawler first fetches https://deepwiki.com/robots.txt. If that file contains Disallow: /private/, the crawler automatically excludes any URLs under that path from the crawl queue due to the robots.isAllowed guard.
Summary
The deepwiki-mcp crawler implements a robust, polite crawling strategy through three integrated mechanisms:
- Concurrency throttling via
p-queuewith a default limit of 5 simultaneous requests, configurable through theDEEPWIKI_CONCURRENCYenvironment variable. - Exponential backoff retry logic that waits 250ms, 500ms, and 1000ms between three retry attempts to handle transient failures gracefully.
- Automatic robots.txt compliance using
robots-parserto fetch, parse, and enforce crawl rules, with a fail-open approach when the file is unavailable.
These features ensure that documentation gathering remains respectful of target server resources and adheres to standard web crawling etiquette.
Frequently Asked Questions
What is the default request concurrency in deepwiki-mcp?
The default concurrency limit is 5 simultaneous requests. This is defined by the MAX_CONCURRENCY constant in src/lib/httpCrawler.ts and enforced by the p-queue library, which queues excess requests until slots become available.
How can I adjust the crawling speed without modifying code?
Set the DEEPWIKI_CONCURRENCY environment variable before running the crawler. For example, export DEEPWIKI_CONCURRENCY=10 increases the limit to 10 concurrent requests. The crawler reads this variable at startup (src/lib/httpCrawler.ts:L10-L12) and passes the value to the queue configuration.
Does deepwiki-mcp skip URLs blocked by robots.txt?
Yes. The crawler fetches /robots.txt from the root domain and uses the robots-parser library to evaluate every candidate URL. If robots.isAllowed(url.href, '*') returns false, the URL is immediately excluded from the crawl queue, ensuring compliance with site restrictions like Disallow: /private/.
What happens if the robots.txt file is missing or unreachable?
The crawler implements a fail-open strategy. If fetching robots.txt results in a network error, 404, or server failure, the crawler proceeds without rule enforcement rather than blocking all access. This follows standard web crawler conventions, allowing documentation gathering to continue while relying on other rate limiting mechanisms to prevent server overload.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →