Deep Crawling Strategies in crawl4ai: BFS vs DFS vs Best-First Explained
crawl4ai provides three built-in deep crawling strategies—BFSDeepCrawlStrategy for level-order traversal, DFSDeepCrawlStrategy for branch-deep exploration, and BestFirstCrawlingStrategy for score-prioritized fetching—that determine how URLs are discovered and fetched while respecting depth limits, page caps, and filtering rules.
When scraping large websites, the order in which you discover pages dramatically impacts efficiency and data quality. The crawl4ai open-source repository (available at unclecode/crawl4ai) ships with three interchangeable deep crawling strategies that control traversal behavior while sharing a unified configuration interface for filters, scoring, and resource limits.
What Are Deep Crawling Strategies in crawl4ai?
All three strategies inherit from the abstract DeepCrawlStrategy class defined in crawl4ai/deep_crawling/__init__.py. They share an identical constructor signature that configures crawling boundaries:
strategy = StrategyClass(
max_depth=2,
filter_chain=FilterChain(),
url_scorer=URLScorer(), # Optional for BFS/DFS, required for Best-First
include_external=False,
score_threshold=-float('inf'),
max_pages=100,
logger=my_logger,
resume_state=None, # Crash-recovery checkpoint
on_state_change=callback, # Async callback after each page
)
max_depth– Hard limit on link depth from the start URL.max_pages– Global ceiling; discovery stops once reached.filter_chain– Pipeline ofFilterobjects (e.g., domain whitelists) that reject URLs before they enter the frontier.url_scorer– Assigns numeric scores to URLs; used for threshold filtering and ordering.
Each strategy implements link_discovery() to extract links from result.links, normalize them via normalize_url_for_deep_crawl (found in crawl4ai/utils.py), apply filters, and insert valid URLs into their specific traversal data structure.
BFSDeepCrawlStrategy: Breadth-First Level Traversal
BFSDeepCrawlStrategy implements a classic breadth-first search using a queue composed of two lists: current_level and next_level. As implemented in crawl4ai/deep_crawling/bfs_strategy.py, this strategy guarantees that all URLs at depth d are fetched before any URL at depth d + 1.
The algorithm maintains a global visited set to prevent duplicate crawling. When discovery yields new URLs, they are added to next_level only if they pass the filter chain and exceed score_threshold. If a URLScorer is provided and the number of discovered URLs exceeds the remaining max_pages capacity, the strategy sorts links descending by score before truncation; otherwise, it preserves pure level-order traversal.
from crawl4ai import BFSDeepCrawlStrategy, FilterChain, URLScorer, AsyncCrawler
class BlogScorer(URLScorer):
def score(self, url: str) -> float:
return 1.0 if "blog" in url else 0.0
strategy = BFSDeepCrawlStrategy(
max_depth=3,
filter_chain=FilterChain(),
url_scorer=BlogScorer(),
max_pages=50,
)
async def run():
crawler = AsyncCrawler()
results = await crawler.run(
start_url="https://example.com",
deep_crawl_strategy=strategy,
)
for r in results:
print(r.url, r.metadata.get("depth"))
Use BFS when you need predictable, exhaustive coverage where shallow pages are prioritized over deep branches.
DFSDeepCrawlStrategy: Depth-First Branch Exploration
DFSDeepCrawlStrategy, located in crawl4ai/deep_crawling/dfs_strategy.py, uses a stack (self.stack) to explore one branch as deeply as possible before backtracking. This approach is ideal for drilling down into specific site sections quickly.
Unlike BFS, DFS maintains a separate _dfs_seen set (lines 37–43 of dfs_strategy.py) that tracks URLs during the discovery phase without marking them globally visited. This prevents duplicate insertion into the stack while allowing the same URL to be rediscovered via different paths if necessary. To preserve discovery order, the strategy pushes newly found URLs onto the stack in reverse order so the first discovered link is processed next.
Scoring behavior mirrors BFS: if max_pages capacity is exceeded during discovery, URLs are sorted by score before being added to the stack.
from crawl4ai import DFSDeepCrawlStrategy, FilterChain, AsyncCrawler
strategy = DFSDeepCrawlStrategy(
max_depth=4,
filter_chain=FilterChain(),
max_pages=30,
)
async def run():
crawler = AsyncCrawler()
results = await crawler.run(
start_url="https://example.org",
deep_crawl_strategy=strategy,
)
for r in results:
print(r.url, r.metadata["depth"])
The printed depths will show characteristic stack behavior (e.g., 0 → 1 → 2 → 3 → 2 → 1 → 0), confirming deep-branch exploration.
BestFirstCrawlingStrategy: Score-Driven Priority Crawling
BestFirstCrawlingStrategy (implemented in crawl4ai/deep_crawling/bff_strategy.py) abandons depth-based ordering in favor of a priority queue (heap) that always extracts the highest-scoring URL next. This strategy requires a URLScorer and is optimal for targeting high-value pages early, regardless of their position in the site hierarchy.
The strategy uses the same discovery logic as BFS/DFS but inserts each candidate into a heap keyed by its numeric score. The heap structure guarantees that the next call to pop() yields the URL with the largest score, enabling aggressive prioritization of product pages, articles, or other valuable content.
from crawl4ai import BestFirstCrawlingStrategy, FilterChain, URLScorer, AsyncCrawler
class ProductScorer(URLScorer):
def score(self, url: str) -> float:
return 10.0 if "/product/" in url else 1.0
strategy = BestFirstCrawlingStrategy(
max_depth=5,
filter_chain=FilterChain(),
url_scorer=ProductScorer(),
max_pages=40,
)
async def run():
crawler = AsyncCrawler()
results = await crawler.run(
start_url="https://shop.example.com",
deep_crawl_strategy=strategy,
)
for r in results:
print(f"{r.url} (score={r.metadata.get('score')})")
Results appear in descending score order, ensuring product pages surface immediately even if they reside at depth 4 while navigation pages at depth 1 are deferred.
Shared Infrastructure and Configuration
All three deep crawling strategies share a common execution flow defined in the base class:
- Link Extraction – Retrieve internal (and optionally external) links from the crawled page.
- Normalization – Convert relative URLs to absolute canonical forms using
normalize_url_for_deep_crawl. - Filtering – Apply the
filter_chainto reject invalid URLs. - Scoring – Calculate scores via
url_scorer.score()and drop URLs belowscore_threshold. - Frontier Insertion – Add valid URLs to the strategy-specific data structure (queue, stack, or heap).
This unified interface allows you to swap strategies with a single line of code without modifying your crawling pipeline. All strategies respect max_depth during discovery and halt immediately when max_pages is reached, preventing runaway crawls.
Summary
- BFSDeepCrawlStrategy uses a queue for level-order traversal, ensuring shallow pages are exhausted before exploring deeper levels; best for exhaustive site mapping.
- DFSDeepCrawlStrategy employs a stack with a separate
_dfs_seenset for depth-first exploration; ideal for drilling into specific content branches quickly. - BestFirstCrawlingStrategy leverages a heap for score-prioritized crawling; optimal when page value matters more than depth or discovery order.
- All strategies inherit from
DeepCrawlStrategyincrawl4ai/deep_crawling/__init__.pyand share common parameters:max_depth,max_pages,filter_chain,url_scorer, andscore_threshold.
Frequently Asked Questions
What is the difference between BFS and DFS in crawl4ai?
BFSDeepCrawlStrategy explores the web graph level by level using a queue (current_level and next_level), ensuring all pages at depth 1 are processed before depth 2. DFSDeepCrawlStrategy uses a stack to follow a single branch to its maximum depth before backtracking, utilizing a _dfs_seen set to manage duplicates without interfering with stack ordering. Choose BFS for broad coverage and DFS for deep-section analysis.
When should I use BestFirstCrawlingStrategy over BFS or DFS?
Use BestFirstCrawlingStrategy when your crawling goal prioritizes content quality or specific page types (e.g., product pages, articles) over systematic exploration. Because it uses a priority queue (heap) ordered by URLScorer output, it surfaces high-scoring pages immediately regardless of depth, making it superior for SEO audits or targeted content harvesting where time-to-value is critical.
How does crawl4ai handle duplicate URLs across different strategies?
All strategies use deduplication to prevent redundant fetching. BFS relies on a global visited set checked before queue insertion. DFS adds a separate _dfs_seen set to filter duplicates during stack insertion without marking them permanently crawled. Best-First uses a similar pre-insertion set to avoid duplicate heap entries. All strategies normalize URLs first via normalize_url_for_deep_crawl to ensure canonical comparison.
Can I combine URL scoring with BFS or DFS strategies?
Yes. While BestFirstCrawlingStrategy requires a URLScorer, both BFS and DFS accept optional scorers. When provided, these strategies use the score only for threshold filtering (dropping URLs below score_threshold) and for sorting discovered links when the number of new URLs exceeds the remaining max_pages budget. This allows score-based prioritization without abandoning the underlying traversal order.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →