How Wigolo's Crawl Tool Manages BFS/DFS Traversal, Sitemaps, and Robots.txt Compliance

Wigolo's crawl tool implements configurable graph traversal through a queue-based system that switches between BFS (queue.shift()) and DFS (queue.pop()) algorithms, while automatically discovering sitemaps and enforcing robots.txt rules via a per-domain RateLimiter.

Wigolo's crawl tool provides a robust web crawling solution that balances aggressive discovery with polite bot behavior. According to the KnockOutEZ/wigolo source code, the tool coordinates traversal strategies, sitemap parsing, and robots.txt compliance through a centralized Crawler class in the src/crawl package. The implementation supports multiple strategies including automatic sitemap detection, breadth-first exploration, and depth-first drilling.

Architecture and Entry Points

The crawling pipeline begins at handleCrawl in src/tools/crawl.ts, which receives a CrawlInput object specifying the target URL and configuration. When a user invokes the crawl tool via CLI or SDK, the system defaults to BFS traversal if no strategy is specified (lines 34-38).

The Crawler class in src/crawl/crawler.ts orchestrates the actual execution. It first validates the seed URL through SSRF protection (guardFetchUrl), then initializes the requested strategy before entering the main traversal loop.

Configurable Traversal Strategies

Wigolo supports five distinct crawling strategies that determine how the discovery graph is explored:

  • bfs – Explores pages level by level using a FIFO queue
  • dfs – Drills deep into branches using a LIFO stack approach
  • auto – Probes for sitemaps first, falling back to BFS if none exist
  • sitemap – Crawls only URLs explicitly listed in discovered sitemaps
  • map – Lightweight URL-only discovery without full page fetching

Strategy selection occurs in the crawl method (lines 33-63), which delegates to specialized handlers based on the input parameter.

How BFS Traversal Works

When configured for breadth-first search, the crawlTraversal method (lines 86-108) maintains a queue of [url, depth] tuples. The algorithm removes the oldest entry using queue.shift(), processes the page, then appends newly discovered links to the queue's end. This ensures all URLs at depth n are processed before any URLs at depth n+1.

How DFS Traversal Works

For depth-first search, the same crawlTraversal method uses queue.pop() instead, removing the most recently added URL. This creates a LIFO behavior that follows each link chain to its maximum depth before backtracking. Both modes respect the maxDepth and maxPages limits configured in CrawlInput.

Sitemap Discovery and Processing

The auto and sitemap strategies rely on probeSitemap in src/crawl/sitemap-first.ts. This helper checks the standard /sitemap.xml location and parses any sitemap URLs advertised within robots.txt.

The discoverSitemapUrls method (lines 300-346 in crawler.ts) handles XML parsing, including support for sitemap indexes that reference multiple sub-sitemaps. The function returns a sorted list of entry URLs, which crawlFromExplicitUrls processes sequentially when operating in sitemap-only mode.

Robots.txt Compliance and Rate Limiting

Before traversing begins, fetchRobots (lines 67-78) retrieves <origin>/robots.txt when the global respectRobotsTxt config is enabled. The RobotsParser class in src/crawl/robots.ts (lines 10-14) parses allow/disallow rules and extracts optional Crawl-Delay directives.

Per-request compliance occurs in the traversal loop (lines 110-114), where robotsParser.isAllowed(path) validates each URL before fetching. Disallowed URLs are skipped and logged without generating network requests.

Rate limiting is enforced through the RateLimiter class in src/crawl/rate-limiter.ts. The acquire(url) method returns a release function that delays subsequent requests to the same domain according to the crawl-delay specified in that domain's robots.txt. This prevents overloading target servers.

Additional link filtering in filterLinks (lines 82-91) removes off-origin links, already-visited URLs, and patterns matching user-specified include/exclude rules before queuing.

Practical Usage Examples

Command Line Interface


# BFS crawl (default) – fetch up to 10 pages, depth 2

wigolo crawl https://example.com --max-pages 10 --max-depth 2

# DFS crawl – explore deeper first

wigolo crawl https://example.com --strategy dfs --max-pages 15

# Auto strategy – try sitemap first, otherwise BFS

wigolo crawl https://example.com --strategy auto

# Sitemap-only crawl – ignore graph traversal, just fetch sitemap entries

wigolo crawl https://example.com --strategy sitemap --max-pages 30

JavaScript/TypeScript API

import { handleCrawl } from 'wigolo/src/tools/crawl.js';
import { createRouter } from 'wigolo/src/fetch/router.js';

const router = createRouter();                      // SmartRouter instance
const input = {
  url: 'https://example.com',
  strategy: 'dfs',          // 'bfs' | 'auto' | 'sitemap' | 'map'
  max_pages: 20,
  max_depth: 3,
  extract_links: true,
};

const result = await handleCrawl(input, router);
console.log('Crawled pages:', result.crawled);
console.log('Discovered URLs:', result.total_found);

Verifying Robots.txt Effects

// Assuming config.respectRobotsTxt = true
await handleCrawl({ url: 'https://example.com' }, router);
// If example.com/robots.txt disallows /private/*, those pages will never be queued.

Summary

  • Queue-based traversal in crawlTraversal switches between BFS (shift()) and DFS (pop()) through simple array operations.
  • Automatic sitemap detection via probeSitemap and discoverSitemapUrls enables efficient seeding without manual URL lists.
  • Robots.txt parsing by RobotsParser extracts both access rules and crawl delays before traversal begins.
  • Per-domain rate limiting through RateLimiter.acquire enforces politeness policies derived from robots.txt.
  • Multi-layer filtering combines robots compliance, URL deduplication, and pattern matching before each fetch.

Frequently Asked Questions

How does Wigolo decide between BFS and DFS?

The crawlTraversal method uses the same queue structure for both strategies. When configured for BFS, it calls queue.shift() to process URLs in FIFO order. For DFS, it calls queue.pop() to process the most recently discovered URLs first, creating a depth-first exploration pattern. This logic resides in src/crawl/crawler.ts within the traversal loop (lines 86-108).

What happens if a site has no sitemap but I use --strategy auto?

The auto strategy first calls probeSitemap to check /sitemap.xml and robots.txt references. If no sitemap is discovered, the crawler automatically falls back to BFS traversal starting from the seed URL. This ensures robust crawling even when sitemap metadata is absent.

Does Wigolo cache robots.txt results between requests?

Yes, the Crawler class fetches robots.txt once per domain during initialization (fetchRobots, lines 67-78) and stores the parsed RobotsParser instance. This parser is reused throughout the session to check isAllowed(path) before every page fetch without redundant network requests.

Can Wigolo respect custom crawl delays not specified in robots.txt?

While Wigolo automatically extracts Crawl-Delay from robots.txt via the RateLimiter, the current implementation primarily respects delays advertised by the target site. For custom throttling, you would need to modify the RateLimiter class in src/crawl/rate-limiter.ts to accept additional user-defined delays alongside the domain-specific values.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →