How Wigolo Handles Rate Limiting, robots.txt Compliance, and Polite Crawling

Wigolo enforces polite crawling by parsing robots.txt directives, extracting Crawl-delay values, and applying the stricter of global or per-domain rate limits before every HTTP request.

Wigolo is an open-source web crawler designed for research-grade data extraction that prioritizes server-friendly behavior. The implementation splits rate limiting, robots.txt compliance, and polite crawling logic across three core TypeScript modules. This architecture ensures automatic adherence to site policies while maintaining deterministic crawl performance.

Robots.txt Parsing and Policy Extraction

The RobotsParser class in src/crawl/robots.ts handles the initial policy extraction. It fetches and parses the text returned from /robots.txt, extracting Disallow, Allow, and Crawl-delay rules for the * user-agent. When a Crawl-delay directive is present, the parser converts the value to milliseconds and stores it for later enforcement.

The parser's logic is validated by tests/unit/crawl/robots.test.ts, which verifies allow/deny pattern matching and crawl-delay handling.

Per-Domain Rate Limiting

The RateLimiter class in src/crawl/rate-limiter.ts maintains a private map of domain-specific delays (robotsDelays). When scheduling a request, the limiter compares the global delay configured in wigolo.config against any domain-specific delay retrieved from robots.txt. The effective delay is always the larger of the two values, ensuring the crawler never exceeds the stricter policy. If a subsequent robots.txt fetch returns a higher delay value, the limiter updates its stored mapping immediately.

All HTTP requests are gated through src/tools/fetch.ts, which respects the rate limiter’s constraints before executing network calls.

Sitemap Discovery via robots.txt

Before attempting to fetch /sitemap.xml, Wigolo first requests <origin>/robots.txt to check for explicit sitemap directives. The extractSitemapUrlFromRobots function in src/crawl/sitemap.ts parses any Sitemap: URLs listed in the file. If no directive exists, the crawler falls back to the conventional /sitemap.xml location.

This workflow is orchestrated by src/crawl/sitemap-first.ts and integrated into the main crawl loop via src/crawl/mapper.ts, ensuring site-specific sitemap declarations take precedence over default paths.

Configuration and Opt-Out Behavior

The RESPECT_ROBOTS_TXT configuration flag (default true) controls whether the crawler enforces these policies. Documented in docs/configuration.md, this setting propagates to both the RobotsParser and RateLimiter instances. Users can disable compliance for special-purpose crawls by setting this flag to false or using the --no-robots CLI option.

Implementation Examples

CLI Usage

Crawl a site while automatically obeying robots.txt and per-domain rate limits:

wigolo crawl https://example.com --depth 3

Disable robots.txt handling (not recommended for polite crawling):

wigolo crawl https://example.com --depth 3 --no-robots

Programmatic SDK Usage

Using the JavaScript SDK with custom configuration:

import { crawl } from '@wigolo/sdk';

const cfg = { RESPECT_ROBOTS_TXT: true, GLOBAL_DELAY_MS: 200 };

crawl('https://example.com', { depth: 3, config: cfg })
  .then(result => console.log('Crawled pages:', result.pages))
  .catch(err => console.error('Crawl failed:', err));

Inspecting Rate Limiter State

For debugging purposes, you can inspect the current delay mappings:

import { RateLimiter } from 'wigolo/src/crawl/rate-limiter';

const limiter = new RateLimiter({ delayMs: 100 });
await limiter.waitFor('example.com');
console.log('Current delay map:', limiter.robotsDelays);

Summary

Frequently Asked Questions

How does Wigolo determine the delay between requests?

The RateLimiter class compares the global delay setting against any domain-specific Crawl-delay found in robots.txt. It applies the larger value, ensuring the crawler respects the stricter policy. This logic is implemented in src/crawl/rate-limiter.ts and enforced before each HTTP request through src/tools/fetch.ts.

Can I disable robots.txt compliance when using Wigolo?

Yes. Set the RESPECT_ROBOTS_TXT configuration flag to false or pass the --no-robots flag via CLI. This bypasses the RobotsParser and removes per-domain delay constraints, though this is not recommended for production crawling.

What happens if a website does not specify a sitemap in robots.txt?

Wigolo falls back to requesting the conventional /sitemap.xml endpoint. The src/crawl/sitemap-first.ts module orchestrates this fallback behavior after checking for Sitemap: directives in the robots.txt file via extractSitemapUrlFromRobots.

How does Wigolo validate URLs against robots.txt rules?

After parsing robots.txt, the crawler rejects any URLs that match Disallow patterns before they are queued for fetching. The RobotsParser class processes both Allow and Disallow directives for the * user-agent, with the specific matching logic validated in tests/unit/crawl/robots.test.ts.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →