# Deep Crawling Strategies in crawl4ai: BFS vs DFS vs Best-First Explained

> Master deep crawling strategies BFS DFS and Best-First in crawl4ai. Understand how each method discovers URLs to optimize your web crawling and data extraction.

- Repository: [UncleCode/crawl4ai](https://github.com/unclecode/crawl4ai)
- Tags: deep-dive
- Published: 2026-03-05

---

**crawl4ai provides three built-in deep crawling strategies—BFSDeepCrawlStrategy for level-order traversal, DFSDeepCrawlStrategy for branch-deep exploration, and BestFirstCrawlingStrategy for score-prioritized fetching—that determine how URLs are discovered and fetched while respecting depth limits, page caps, and filtering rules.**

When scraping large websites, the order in which you discover pages dramatically impacts efficiency and data quality. The `crawl4ai` open-source repository (available at `unclecode/crawl4ai`) ships with three interchangeable **deep crawling strategies** that control traversal behavior while sharing a unified configuration interface for filters, scoring, and resource limits.

## What Are Deep Crawling Strategies in crawl4ai?

All three strategies inherit from the abstract `DeepCrawlStrategy` class defined in [`crawl4ai/deep_crawling/__init__.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/deep_crawling/__init__.py). They share an identical constructor signature that configures crawling boundaries:

```python
strategy = StrategyClass(
    max_depth=2,
    filter_chain=FilterChain(),
    url_scorer=URLScorer(),      # Optional for BFS/DFS, required for Best-First

    include_external=False,
    score_threshold=-float('inf'),
    max_pages=100,
    logger=my_logger,
    resume_state=None,           # Crash-recovery checkpoint

    on_state_change=callback,    # Async callback after each page

)

```

- **`max_depth`** – Hard limit on link depth from the start URL.
- **`max_pages`** – Global ceiling; discovery stops once reached.
- **`filter_chain`** – Pipeline of `Filter` objects (e.g., domain whitelists) that reject URLs before they enter the frontier.
- **`url_scorer`** – Assigns numeric scores to URLs; used for threshold filtering and ordering.

Each strategy implements `link_discovery()` to extract links from `result.links`, normalize them via `normalize_url_for_deep_crawl` (found in [`crawl4ai/utils.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/utils.py)), apply filters, and insert valid URLs into their specific traversal data structure.

## BFSDeepCrawlStrategy: Breadth-First Level Traversal

**BFSDeepCrawlStrategy** implements a classic breadth-first search using a **queue** composed of two lists: `current_level` and `next_level`. As implemented in [`crawl4ai/deep_crawling/bfs_strategy.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/deep_crawling/bfs_strategy.py), this strategy guarantees that all URLs at depth *d* are fetched before any URL at depth *d + 1*.

The algorithm maintains a global `visited` set to prevent duplicate crawling. When discovery yields new URLs, they are added to `next_level` only if they pass the filter chain and exceed `score_threshold`. If a `URLScorer` is provided and the number of discovered URLs exceeds the remaining `max_pages` capacity, the strategy sorts links **descending by score** before truncation; otherwise, it preserves pure level-order traversal.

```python
from crawl4ai import BFSDeepCrawlStrategy, FilterChain, URLScorer, AsyncCrawler

class BlogScorer(URLScorer):
    def score(self, url: str) -> float:
        return 1.0 if "blog" in url else 0.0

strategy = BFSDeepCrawlStrategy(
    max_depth=3,
    filter_chain=FilterChain(),
    url_scorer=BlogScorer(),
    max_pages=50,
)

async def run():
    crawler = AsyncCrawler()
    results = await crawler.run(
        start_url="https://example.com",
        deep_crawl_strategy=strategy,
    )
    for r in results:
        print(r.url, r.metadata.get("depth"))

```

Use **BFS** when you need predictable, exhaustive coverage where shallow pages are prioritized over deep branches.

## DFSDeepCrawlStrategy: Depth-First Branch Exploration

**DFSDeepCrawlStrategy**, located in [`crawl4ai/deep_crawling/dfs_strategy.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/deep_crawling/dfs_strategy.py), uses a **stack** (`self.stack`) to explore one branch as deeply as possible before backtracking. This approach is ideal for drilling down into specific site sections quickly.

Unlike BFS, DFS maintains a separate `_dfs_seen` set (lines 37–43 of [`dfs_strategy.py`](https://github.com/unclecode/crawl4ai/blob/main/dfs_strategy.py)) that tracks URLs during the discovery phase without marking them globally visited. This prevents duplicate insertion into the stack while allowing the same URL to be rediscovered via different paths if necessary. To preserve discovery order, the strategy pushes newly found URLs onto the stack **in reverse** order so the first discovered link is processed next.

Scoring behavior mirrors BFS: if `max_pages` capacity is exceeded during discovery, URLs are sorted by score before being added to the stack.

```python
from crawl4ai import DFSDeepCrawlStrategy, FilterChain, AsyncCrawler

strategy = DFSDeepCrawlStrategy(
    max_depth=4,
    filter_chain=FilterChain(),
    max_pages=30,
)

async def run():
    crawler = AsyncCrawler()
    results = await crawler.run(
        start_url="https://example.org",
        deep_crawl_strategy=strategy,
    )
    for r in results:
        print(r.url, r.metadata["depth"])

```

The printed depths will show characteristic stack behavior (e.g., 0 → 1 → 2 → 3 → 2 → 1 → 0), confirming deep-branch exploration.

## BestFirstCrawlingStrategy: Score-Driven Priority Crawling

**BestFirstCrawlingStrategy** (implemented in [`crawl4ai/deep_crawling/bff_strategy.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/deep_crawling/bff_strategy.py)) abandons depth-based ordering in favor of a **priority queue** (heap) that always extracts the highest-scoring URL next. This strategy requires a `URLScorer` and is optimal for targeting high-value pages early, regardless of their position in the site hierarchy.

The strategy uses the same discovery logic as BFS/DFS but inserts each candidate into a heap keyed by its numeric score. The heap structure guarantees that the next call to `pop()` yields the URL with the largest score, enabling aggressive prioritization of product pages, articles, or other valuable content.

```python
from crawl4ai import BestFirstCrawlingStrategy, FilterChain, URLScorer, AsyncCrawler

class ProductScorer(URLScorer):
    def score(self, url: str) -> float:
        return 10.0 if "/product/" in url else 1.0

strategy = BestFirstCrawlingStrategy(
    max_depth=5,
    filter_chain=FilterChain(),
    url_scorer=ProductScorer(),
    max_pages=40,
)

async def run():
    crawler = AsyncCrawler()
    results = await crawler.run(
        start_url="https://shop.example.com",
        deep_crawl_strategy=strategy,
    )
    for r in results:
        print(f"{r.url} (score={r.metadata.get('score')})")

```

Results appear in descending score order, ensuring product pages surface immediately even if they reside at depth 4 while navigation pages at depth 1 are deferred.

## Shared Infrastructure and Configuration

All three **deep crawling strategies** share a common execution flow defined in the base class:

1. **Link Extraction** – Retrieve internal (and optionally external) links from the crawled page.
2. **Normalization** – Convert relative URLs to absolute canonical forms using `normalize_url_for_deep_crawl`.
3. **Filtering** – Apply the `filter_chain` to reject invalid URLs.
4. **Scoring** – Calculate scores via `url_scorer.score()` and drop URLs below `score_threshold`.
5. **Frontier Insertion** – Add valid URLs to the strategy-specific data structure (queue, stack, or heap).

This unified interface allows you to swap strategies with a single line of code without modifying your crawling pipeline. All strategies respect `max_depth` during discovery and halt immediately when `max_pages` is reached, preventing runaway crawls.

## Summary

- **BFSDeepCrawlStrategy** uses a queue for level-order traversal, ensuring shallow pages are exhausted before exploring deeper levels; best for exhaustive site mapping.
- **DFSDeepCrawlStrategy** employs a stack with a separate `_dfs_seen` set for depth-first exploration; ideal for drilling into specific content branches quickly.
- **BestFirstCrawlingStrategy** leverages a heap for score-prioritized crawling; optimal when page value matters more than depth or discovery order.
- All strategies inherit from `DeepCrawlStrategy` in [`crawl4ai/deep_crawling/__init__.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/deep_crawling/__init__.py) and share common parameters: `max_depth`, `max_pages`, `filter_chain`, `url_scorer`, and `score_threshold`.

## Frequently Asked Questions

### What is the difference between BFS and DFS in crawl4ai?

**BFSDeepCrawlStrategy** explores the web graph level by level using a queue (`current_level` and `next_level`), ensuring all pages at depth 1 are processed before depth 2. **DFSDeepCrawlStrategy** uses a stack to follow a single branch to its maximum depth before backtracking, utilizing a `_dfs_seen` set to manage duplicates without interfering with stack ordering. Choose BFS for broad coverage and DFS for deep-section analysis.

### When should I use BestFirstCrawlingStrategy over BFS or DFS?

Use **BestFirstCrawlingStrategy** when your crawling goal prioritizes content quality or specific page types (e.g., product pages, articles) over systematic exploration. Because it uses a priority queue (heap) ordered by `URLScorer` output, it surfaces high-scoring pages immediately regardless of depth, making it superior for SEO audits or targeted content harvesting where time-to-value is critical.

### How does crawl4ai handle duplicate URLs across different strategies?

All strategies use deduplication to prevent redundant fetching. **BFS** relies on a global `visited` set checked before queue insertion. **DFS** adds a separate `_dfs_seen` set to filter duplicates during stack insertion without marking them permanently crawled. **Best-First** uses a similar pre-insertion set to avoid duplicate heap entries. All strategies normalize URLs first via `normalize_url_for_deep_crawl` to ensure canonical comparison.

### Can I combine URL scoring with BFS or DFS strategies?

Yes. While **BestFirstCrawlingStrategy** requires a `URLScorer`, both **BFS** and **DFS** accept optional scorers. When provided, these strategies use the score only for threshold filtering (dropping URLs below `score_threshold`) and for sorting discovered links when the number of new URLs exceeds the remaining `max_pages` budget. This allows score-based prioritization without abandoning the underlying traversal order.