# How Wigolo's Crawl Tool Manages BFS/DFS Traversal, Sitemaps, and Robots.txt Compliance

> Discover how Wigolo's crawl tool manages BFS DFS traversal sitemaps and robots txt compliance. Learn about its configurable queue system and RateLimiter for efficient web crawling.

- Repository: [Towhid Khan/wigolo](https://github.com/KnockOutEZ/wigolo)
- Tags: how-to-guide
- Published: 2026-07-29

---

**Wigolo's crawl tool implements configurable graph traversal through a queue-based system that switches between BFS (`queue.shift()`) and DFS (`queue.pop()`) algorithms, while automatically discovering sitemaps and enforcing robots.txt rules via a per-domain RateLimiter.**

Wigolo's crawl tool provides a robust web crawling solution that balances aggressive discovery with polite bot behavior. According to the KnockOutEZ/wigolo source code, the tool coordinates traversal strategies, sitemap parsing, and robots.txt compliance through a centralized `Crawler` class in the `src/crawl` package. The implementation supports multiple strategies including automatic sitemap detection, breadth-first exploration, and depth-first drilling.

## Architecture and Entry Points

The crawling pipeline begins at `handleCrawl` in [`src/tools/crawl.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/tools/crawl.ts), which receives a `CrawlInput` object specifying the target URL and configuration. When a user invokes the **crawl** tool via CLI or SDK, the system defaults to **BFS** traversal if no strategy is specified (lines 34-38).

The `Crawler` class in [`src/crawl/crawler.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/crawl/crawler.ts) orchestrates the actual execution. It first validates the seed URL through SSRF protection (`guardFetchUrl`), then initializes the requested strategy before entering the main traversal loop.

## Configurable Traversal Strategies

Wigolo supports five distinct crawling strategies that determine how the discovery graph is explored:

- **`bfs`** – Explores pages level by level using a FIFO queue
- **`dfs`** – Drills deep into branches using a LIFO stack approach  
- **`auto`** – Probes for sitemaps first, falling back to BFS if none exist
- **`sitemap`** – Crawls only URLs explicitly listed in discovered sitemaps
- **`map`** – Lightweight URL-only discovery without full page fetching

Strategy selection occurs in the `crawl` method (lines 33-63), which delegates to specialized handlers based on the input parameter.

### How BFS Traversal Works

When configured for breadth-first search, the `crawlTraversal` method (lines 86-108) maintains a queue of `[url, depth]` tuples. The algorithm removes the oldest entry using `queue.shift()`, processes the page, then appends newly discovered links to the queue's end. This ensures all URLs at depth *n* are processed before any URLs at depth *n+1*.

### How DFS Traversal Works

For depth-first search, the same `crawlTraversal` method uses `queue.pop()` instead, removing the most recently added URL. This creates a LIFO behavior that follows each link chain to its maximum depth before backtracking. Both modes respect the `maxDepth` and `maxPages` limits configured in `CrawlInput`.

## Sitemap Discovery and Processing

The **auto** and **sitemap** strategies rely on `probeSitemap` in [`src/crawl/sitemap-first.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/crawl/sitemap-first.ts). This helper checks the standard [`/sitemap.xml`](https://github.com/KnockOutEZ/wigolo/blob/main//sitemap.xml) location and parses any sitemap URLs advertised within [`robots.txt`](https://github.com/KnockOutEZ/wigolo/blob/main/robots.txt).

The `discoverSitemapUrls` method (lines 300-346 in [`crawler.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/crawler.ts)) handles XML parsing, including support for sitemap indexes that reference multiple sub-sitemaps. The function returns a sorted list of entry URLs, which `crawlFromExplicitUrls` processes sequentially when operating in sitemap-only mode.

## Robots.txt Compliance and Rate Limiting

Before traversing begins, `fetchRobots` (lines 67-78) retrieves `<origin>/robots.txt` when the global `respectRobotsTxt` config is enabled. The `RobotsParser` class in [`src/crawl/robots.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/crawl/robots.ts) (lines 10-14) parses allow/disallow rules and extracts optional `Crawl-Delay` directives.

**Per-request compliance** occurs in the traversal loop (lines 110-114), where `robotsParser.isAllowed(path)` validates each URL before fetching. Disallowed URLs are skipped and logged without generating network requests.

**Rate limiting** is enforced through the `RateLimiter` class in [`src/crawl/rate-limiter.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/crawl/rate-limiter.ts). The `acquire(url)` method returns a release function that delays subsequent requests to the same domain according to the crawl-delay specified in that domain's robots.txt. This prevents overloading target servers.

Additional link filtering in `filterLinks` (lines 82-91) removes off-origin links, already-visited URLs, and patterns matching user-specified `include/exclude` rules before queuing.

## Practical Usage Examples

### Command Line Interface

```bash

# BFS crawl (default) – fetch up to 10 pages, depth 2

wigolo crawl https://example.com --max-pages 10 --max-depth 2

# DFS crawl – explore deeper first

wigolo crawl https://example.com --strategy dfs --max-pages 15

# Auto strategy – try sitemap first, otherwise BFS

wigolo crawl https://example.com --strategy auto

# Sitemap-only crawl – ignore graph traversal, just fetch sitemap entries

wigolo crawl https://example.com --strategy sitemap --max-pages 30

```

### JavaScript/TypeScript API

```typescript
import { handleCrawl } from 'wigolo/src/tools/crawl.js';
import { createRouter } from 'wigolo/src/fetch/router.js';

const router = createRouter();                      // SmartRouter instance
const input = {
  url: 'https://example.com',
  strategy: 'dfs',          // 'bfs' | 'auto' | 'sitemap' | 'map'
  max_pages: 20,
  max_depth: 3,
  extract_links: true,
};

const result = await handleCrawl(input, router);
console.log('Crawled pages:', result.crawled);
console.log('Discovered URLs:', result.total_found);

```

### Verifying Robots.txt Effects

```typescript
// Assuming config.respectRobotsTxt = true
await handleCrawl({ url: 'https://example.com' }, router);
// If example.com/robots.txt disallows /private/*, those pages will never be queued.

```

## Summary

- **Queue-based traversal** in `crawlTraversal` switches between BFS (`shift()`) and DFS (`pop()`) through simple array operations.
- **Automatic sitemap detection** via `probeSitemap` and `discoverSitemapUrls` enables efficient seeding without manual URL lists.
- **Robots.txt parsing** by `RobotsParser` extracts both access rules and crawl delays before traversal begins.
- **Per-domain rate limiting** through `RateLimiter.acquire` enforces politeness policies derived from robots.txt.
- **Multi-layer filtering** combines robots compliance, URL deduplication, and pattern matching before each fetch.

## Frequently Asked Questions

### How does Wigolo decide between BFS and DFS?

The `crawlTraversal` method uses the same queue structure for both strategies. When configured for **BFS**, it calls `queue.shift()` to process URLs in FIFO order. For **DFS**, it calls `queue.pop()` to process the most recently discovered URLs first, creating a depth-first exploration pattern. This logic resides in [`src/crawl/crawler.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/crawl/crawler.ts) within the traversal loop (lines 86-108).

### What happens if a site has no sitemap but I use `--strategy auto`?

The **auto** strategy first calls `probeSitemap` to check [`/sitemap.xml`](https://github.com/KnockOutEZ/wigolo/blob/main//sitemap.xml) and robots.txt references. If no sitemap is discovered, the crawler automatically falls back to BFS traversal starting from the seed URL. This ensures robust crawling even when sitemap metadata is absent.

### Does Wigolo cache robots.txt results between requests?

Yes, the `Crawler` class fetches robots.txt once per domain during initialization (`fetchRobots`, lines 67-78) and stores the parsed `RobotsParser` instance. This parser is reused throughout the session to check `isAllowed(path)` before every page fetch without redundant network requests.

### Can Wigolo respect custom crawl delays not specified in robots.txt?

While Wigolo automatically extracts `Crawl-Delay` from robots.txt via the `RateLimiter`, the current implementation primarily respects delays advertised by the target site. For custom throttling, you would need to modify the `RateLimiter` class in [`src/crawl/rate-limiter.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/crawl/rate-limiter.ts) to accept additional user-defined delays alongside the domain-specific values.