# How the deepwiki-mcp Crawler Implements Rate Limiting and Respects robots.txt

> Discover how the deepwiki-mcp crawler applies rate limiting with p-queue, exponential backoff retries, and respects robots.txt rules using robots-parser for efficient crawling.

- Repository: [Kevin Kern/deepwiki-mcp](https://github.com/regenrek/deepwiki-mcp)
- Tags: internals
- Published: 2026-02-16

---

**The deepwiki-mcp crawler uses p-queue to limit concurrent requests to 5 by default, implements exponential backoff retries up to 3 attempts, and enforces robots.txt rules using robots-parser to filter disallowed URLs before fetching.**

The `regenrek/deepwiki-mcp` repository provides a Model Context Protocol (MCP) server designed to crawl documentation sites and convert them to Markdown. Understanding how this tool manages request rates and respects site policies is essential for ethical web scraping. This article examines the implementation details in [`src/lib/httpCrawler.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/lib/httpCrawler.ts) to explain exactly how the deepwiki-mcp crawler handles rate limiting and robots.txt compliance.

## Core Rate Limiting Mechanisms in deepwiki-mcp

### Concurrency Control with p-queue

The primary rate limiting mechanism relies on the `p-queue` library to cap simultaneous connections. In [`src/lib/httpCrawler.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/lib/httpCrawler.ts), the crawler initializes a queue with a default **MAX_CONCURRENCY** of 5:

```ts
const queue = new PQueue({ concurrency: MAX_CONCURRENCY })

```

Every URL fetch operation is wrapped in `queue.add()`, ensuring that no more than five HTTP requests execute simultaneously. The queue guarantees that excess requests wait until slots become available, throttling the overall request rate to prevent overwhelming target servers.

### Configurable Concurrency via Environment Variables

Users can adjust crawling speed without modifying source code. The crawler checks for the **DEEPWIKI_CONCURRENCY** environment variable at startup (`src/lib/httpCrawler.ts:L10-L12`):

```bash
export DEEPWIKI_CONCURRENCY=10

```

When set, this value overrides the default limit of 5, allowing up to ten simultaneous requests. This flexibility enables adaptation to different network conditions or server capabilities.

### Exponential Backoff for Retry Logic

Transient failures trigger an exponential backoff strategy rather than immediate retries. The implementation defines **RETRY_LIMIT** as 3 attempts and **BACKOFF_BASE_MS** as 250 milliseconds. When a request fails, the crawler waits according to the formula (`src/lib/httpCrawler.ts:L78-L82`):

```ts
await setTimeout(BACKOFF_BASE_MS * 2 ** (retries - 1))

```

This yields wait times of 250ms, 500ms, and 1000ms for successive retries, mitigating temporary spikes or transient failures without flooding the target host with aggressive retry attempts.

## robots.txt Compliance Implementation

### Fetching and Parsing robots.txt

The crawler enforces site policies by pre-fetching and parsing the root domain's robots.txt file. Using the **robots-parser** library, it constructs a ruleset before crawling begins (`src/lib/httpCrawler.ts:L42-L52`):

```ts
const robotsUrl = new URL('/robots.txt', root)
const res = await fetch(robotsUrl)
robots = robotsParser(robotsUrl.href, body)

```

This parser instance evaluates crawl rules against user-agent strings and path patterns defined by the site owner, creating a filter for subsequent URL processing.

### URL Filtering Against Crawl Rules

Every candidate URL undergoes a permission check before entering the fetch queue. The crawler validates URLs using the `robots.isAllowed()` method with a wildcard user-agent (`src/lib/httpCrawler.ts:L30-L32`):

```ts
if (robots && !robots.isAllowed(url.href, '*')) return

```

If robots.txt explicitly disallows the path—such as `Disallow: /private/` or `Disallow: /admin/`—the crawler immediately excludes the URL from the queue, ensuring compliance with site restrictions.

### Fail-Open Safety Mechanism

When robots.txt cannot be fetched due to network errors, 404 responses, or server failures, the crawler implements a **fail-open** strategy. It proceeds with crawling without rule enforcement rather than blocking all access. This behavior follows standard web crawler conventions, allowing documentation gathering to continue when robots.txt is temporarily unavailable while maintaining safety through other rate limiting mechanisms.

## Practical Usage Examples

### Basic crawl with default rate limiting

```ts
import { crawl } from '@/lib/httpCrawler'
import { URL } from 'node:url'

await crawl({
  root: new URL('https://deepwiki.com/owner/repo'),
  maxDepth: 1,
  emit: () => {},               // progress callback (optional)
  verbose: true,
})

```

This configuration runs with the default concurrency of 5 requests and automatically fetches and respects the site's robots.txt rules.

### Override concurrency via environment variable

```bash
export DEEPWIKI_CONCURRENCY=10
node myScript.js

```

Setting `DEEPWIKI_CONCURRENCY` to 10 increases the limit to ten simultaneous requests, as read from `process.env.DEEPWIKI_CONCURRENCY` (`src/lib/httpCrawler.ts:L10-L12`).

### Inspect robots.txt handling

```ts
import { crawl } from '@/lib/httpCrawler'

await crawl({
  root: new URL('https://deepwiki.com/some/project'),
  maxDepth: 1,
  emit: () => {},
  verbose: false,
})

```

The crawler first fetches `https://deepwiki.com/robots.txt`. If that file contains `Disallow: /private/`, the crawler automatically excludes any URLs under that path from the crawl queue due to the `robots.isAllowed` guard.

## Summary

The deepwiki-mcp crawler implements a robust, polite crawling strategy through three integrated mechanisms:

- **Concurrency throttling** via `p-queue` with a default limit of 5 simultaneous requests, configurable through the `DEEPWIKI_CONCURRENCY` environment variable.
- **Exponential backoff retry logic** that waits 250ms, 500ms, and 1000ms between three retry attempts to handle transient failures gracefully.
- **Automatic robots.txt compliance** using `robots-parser` to fetch, parse, and enforce crawl rules, with a fail-open approach when the file is unavailable.

These features ensure that documentation gathering remains respectful of target server resources and adheres to standard web crawling etiquette.

## Frequently Asked Questions

### What is the default request concurrency in deepwiki-mcp?

The default concurrency limit is **5 simultaneous requests**. This is defined by the `MAX_CONCURRENCY` constant in [`src/lib/httpCrawler.ts`](https://github.com/regenrek/deepwiki-mcp/blob/main/src/lib/httpCrawler.ts) and enforced by the `p-queue` library, which queues excess requests until slots become available.

### How can I adjust the crawling speed without modifying code?

Set the `DEEPWIKI_CONCURRENCY` environment variable before running the crawler. For example, `export DEEPWIKI_CONCURRENCY=10` increases the limit to 10 concurrent requests. The crawler reads this variable at startup (`src/lib/httpCrawler.ts:L10-L12`) and passes the value to the queue configuration.

### Does deepwiki-mcp skip URLs blocked by robots.txt?

Yes. The crawler fetches [`/robots.txt`](https://github.com/regenrek/deepwiki-mcp/blob/main//robots.txt) from the root domain and uses the `robots-parser` library to evaluate every candidate URL. If `robots.isAllowed(url.href, '*')` returns false, the URL is immediately excluded from the crawl queue, ensuring compliance with site restrictions like `Disallow: /private/`.

### What happens if the robots.txt file is missing or unreachable?

The crawler implements a **fail-open** strategy. If fetching [`robots.txt`](https://github.com/regenrek/deepwiki-mcp/blob/main/robots.txt) results in a network error, 404, or server failure, the crawler proceeds without rule enforcement rather than blocking all access. This follows standard web crawler conventions, allowing documentation gathering to continue while relying on other rate limiting mechanisms to prevent server overload.