# How Wigolo Handles Rate Limiting, robots.txt Compliance, and Polite Crawling

> Learn how Wigolo ensures polite crawling with robots.txt compliance, Crawl-delay parsing, and strict rate limiting for efficient web scraping and data collection.

- Repository: [Towhid Khan/wigolo](https://github.com/KnockOutEZ/wigolo)
- Tags: how-to-guide
- Published: 2026-07-19

---

**Wigolo enforces polite crawling by parsing robots.txt directives, extracting Crawl-delay values, and applying the stricter of global or per-domain rate limits before every HTTP request.**

Wigolo is an open-source web crawler designed for research-grade data extraction that prioritizes server-friendly behavior. The implementation splits **rate limiting**, **robots.txt compliance**, and **polite crawling** logic across three core TypeScript modules. This architecture ensures automatic adherence to site policies while maintaining deterministic crawl performance.

## Robots.txt Parsing and Policy Extraction

The `RobotsParser` class in [`src/crawl/robots.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/crawl/robots.ts) handles the initial policy extraction. It fetches and parses the text returned from [`/robots.txt`](https://github.com/KnockOutEZ/wigolo/blob/main//robots.txt), extracting `Disallow`, `Allow`, and `Crawl-delay` rules for the `*` user-agent. When a `Crawl-delay` directive is present, the parser converts the value to milliseconds and stores it for later enforcement.

The parser's logic is validated by [`tests/unit/crawl/robots.test.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/tests/unit/crawl/robots.test.ts), which verifies allow/deny pattern matching and crawl-delay handling.

## Per-Domain Rate Limiting

The `RateLimiter` class in [`src/crawl/rate-limiter.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/crawl/rate-limiter.ts) maintains a private map of domain-specific delays (`robotsDelays`). When scheduling a request, the limiter compares the global delay configured in `wigolo.config` against any domain-specific delay retrieved from robots.txt. The **effective delay** is always the larger of the two values, ensuring the crawler never exceeds the stricter policy. If a subsequent robots.txt fetch returns a higher delay value, the limiter updates its stored mapping immediately.

All HTTP requests are gated through [`src/tools/fetch.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/tools/fetch.ts), which respects the rate limiter’s constraints before executing network calls.

## Sitemap Discovery via robots.txt

Before attempting to fetch [`/sitemap.xml`](https://github.com/KnockOutEZ/wigolo/blob/main//sitemap.xml), Wigolo first requests `<origin>/robots.txt` to check for explicit sitemap directives. The `extractSitemapUrlFromRobots` function in [`src/crawl/sitemap.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/crawl/sitemap.ts) parses any `Sitemap:` URLs listed in the file. If no directive exists, the crawler falls back to the conventional [`/sitemap.xml`](https://github.com/KnockOutEZ/wigolo/blob/main//sitemap.xml) location.

This workflow is orchestrated by [`src/crawl/sitemap-first.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/crawl/sitemap-first.ts) and integrated into the main crawl loop via [`src/crawl/mapper.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/crawl/mapper.ts), ensuring site-specific sitemap declarations take precedence over default paths.

## Configuration and Opt-Out Behavior

The `RESPECT_ROBOTS_TXT` configuration flag (default `true`) controls whether the crawler enforces these policies. Documented in [`docs/configuration.md`](https://github.com/KnockOutEZ/wigolo/blob/main/docs/configuration.md), this setting propagates to both the `RobotsParser` and `RateLimiter` instances. Users can disable compliance for special-purpose crawls by setting this flag to `false` or using the `--no-robots` CLI option.

## Implementation Examples

### CLI Usage

Crawl a site while automatically obeying robots.txt and per-domain rate limits:

```bash
wigolo crawl https://example.com --depth 3

```

Disable robots.txt handling (not recommended for polite crawling):

```bash
wigolo crawl https://example.com --depth 3 --no-robots

```

### Programmatic SDK Usage

Using the JavaScript SDK with custom configuration:

```javascript
import { crawl } from '@wigolo/sdk';

const cfg = { RESPECT_ROBOTS_TXT: true, GLOBAL_DELAY_MS: 200 };

crawl('https://example.com', { depth: 3, config: cfg })
  .then(result => console.log('Crawled pages:', result.pages))
  .catch(err => console.error('Crawl failed:', err));

```

### Inspecting Rate Limiter State

For debugging purposes, you can inspect the current delay mappings:

```javascript
import { RateLimiter } from 'wigolo/src/crawl/rate-limiter';

const limiter = new RateLimiter({ delayMs: 100 });
await limiter.waitFor('example.com');
console.log('Current delay map:', limiter.robotsDelays);

```

## Summary

- Wigolo parses robots.txt using the `RobotsParser` class in [`src/crawl/robots.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/crawl/robots.ts), extracting `Disallow`, `Allow`, and `Crawl-delay` directives for the default user-agent.
- The `RateLimiter` in [`src/crawl/rate-limiter.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/crawl/rate-limiter.ts) computes the **effective delay** as the maximum of the global delay and any per-domain robots.txt delay, storing values in the `robotsDelays` map.
- Sitemap discovery prioritizes robots.txt declarations via [`src/crawl/sitemap-first.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/crawl/sitemap-first.ts) before falling back to conventional [`/sitemap.xml`](https://github.com/KnockOutEZ/wigolo/blob/main//sitemap.xml) paths.
- Compliance can be disabled via the `RESPECT_ROBOTS_TXT` flag or `--no-robots` CLI option for specialized crawl scenarios.
- All behaviors are validated by unit tests in [`tests/unit/crawl/robots.test.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/tests/unit/crawl/robots.test.ts) and [`tests/unit/crawl/rate-limiter.test.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/tests/unit/crawl/rate-limiter.test.ts).

## Frequently Asked Questions

### How does Wigolo determine the delay between requests?

The `RateLimiter` class compares the global delay setting against any domain-specific `Crawl-delay` found in robots.txt. It applies the larger value, ensuring the crawler respects the stricter policy. This logic is implemented in [`src/crawl/rate-limiter.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/crawl/rate-limiter.ts) and enforced before each HTTP request through [`src/tools/fetch.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/tools/fetch.ts).

### Can I disable robots.txt compliance when using Wigolo?

Yes. Set the `RESPECT_ROBOTS_TXT` configuration flag to `false` or pass the `--no-robots` flag via CLI. This bypasses the `RobotsParser` and removes per-domain delay constraints, though this is not recommended for production crawling.

### What happens if a website does not specify a sitemap in robots.txt?

Wigolo falls back to requesting the conventional [`/sitemap.xml`](https://github.com/KnockOutEZ/wigolo/blob/main//sitemap.xml) endpoint. The [`src/crawl/sitemap-first.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/crawl/sitemap-first.ts) module orchestrates this fallback behavior after checking for `Sitemap:` directives in the robots.txt file via `extractSitemapUrlFromRobots`.

### How does Wigolo validate URLs against robots.txt rules?

After parsing robots.txt, the crawler rejects any URLs that match `Disallow` patterns before they are queued for fetching. The `RobotsParser` class processes both `Allow` and `Disallow` directives for the `*` user-agent, with the specific matching logic validated in [`tests/unit/crawl/robots.test.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/tests/unit/crawl/robots.test.ts).