How to Implement Crash Recovery with resume_state for Long Crawls in crawl4ai

Use the resume_state parameter in deep-crawling strategies (BFS, DFS, BFF) or resume_from in AdaptiveCrawler to persist and restore crawl progress, preventing data loss during crashes.

The unclecode/crawl4ai library provides built-in crash-recovery mechanisms that allow long-running crawls to survive process interruptions. By leveraging state serialization in BFSDeepCrawlStrategy or the CrawlState class for adaptive crawling, you can resume exactly where the crawler left off without reprocessing visited URLs.

How Resume State Works in Deep Crawling Strategies

The deep-crawling architecture in crawl4ai maintains internal traversal queues, visited sets, and depth counters. When you provide a resume_state dictionary to the strategy constructor, the crawler hydrates these structures before processing begins.

State Hydration in BFS and DFS

In crawl4ai/deep_crawling/bfs_strategy.py, the _arun_batch method (lines 65-73) checks for self._resume_state at startup. If present, it reconstructs the visited set, pending list, depths map, and pages_crawled counter. The DFS implementation in crawl4ai/deep_crawling/dfs_strategy.py follows an identical pattern but restores a stack instead of a queue (lines 41-50).


# Simplified internal logic from bfs_strategy.py

if self._resume_state:
    visited = set(self._resume_state.get("visited", []))
    pending = list(self._resume_state.get("pending", []))
    depths = self._resume_state.get("depths", {})
    pages_crawled = self._resume_state.get("pages_crawled", 0)

The State Capture Callback

After each URL completes, strategies invoke the on_state_change callback with a state dictionary. This callback, referenced in bfs_strategy.py (lines 13-23), receives the current visited, pending, depths, and pages_crawled values. You persist this dictionary externally—typically to JSON—to create a restore point.

Step-by-Step Implementation for Long Crawls

Implementing crash recovery requires two components: a persistence callback to save state after every page, and logic to load that state when initializing the crawler.

Creating an Async State Persistence Callback

Write an asynchronous function that atomically writes the state dictionary to disk. Using aiofiles prevents blocking the event loop during I/O operations.

import json
import aiofiles
from pathlib import Path

async def persist_state(state: dict, path: Path = Path("crawl_state.json")):
    """Atomically save crawl state to prevent corruption during crashes."""
    tmp = path.with_suffix(".tmp")
    async with aiofiles.open(tmp, "w") as f:
        await f.write(json.dumps(state, indent=2))
    await aiofiles.os.rename(tmp, path)

Resuming a BFS Deep Crawl

Load the saved state from disk and pass it to the BFSDeepCrawlStrategy constructor via the resume_state parameter. Assign the persist_state function to on_state_change to enable automatic state updates.

import asyncio
import json
from pathlib import Path
from crawl4ai import AsyncWebCrawler
from crawl4ai.deep_crawling.bfs_strategy import BFSDeepCrawlStrategy
from crawl4ai.async_configs import CrawlerRunConfig

async def main():
    state_path = Path("crawl_state.json")
    
    # Load previous state if resuming

    resume_state = None
    if state_path.exists():
        resume_state = json.loads(state_path.read_text())
    
    # Initialize strategy with resume capability

    strategy = BFSDeepCrawlStrategy(
        max_depth=5,
        max_pages=2000,
        resume_state=resume_state,
        on_state_change=lambda s: persist_state(s, state_path)
    )
    
    async with AsyncWebCrawler() as crawler:
        config = CrawlerRunConfig(timeout=30)
        results = await strategy.arun("https://example.com", crawler, config)
        print(f"Retrieved {len(results)} pages")

if __name__ == "__main__":
    asyncio.run(main())

If the process crashes, restart the script to reload crawl_state.json and continue from the last persisted URL.

Crash Recovery for Adaptive Crawling

The AdaptiveCrawler class uses a different mechanism based on the CrawlState class defined in crawl4ai/adaptive_crawler.py (lines 53-88). Instead of a dictionary callback, you save the entire session state to JSON using CrawlState.save(), then restore it via the resume_from parameter in the constructor (lines 1311-1315).

from crawl4ai import AdaptiveCrawler, AdaptiveConfig

# First run - auto-saves state

config = AdaptiveConfig(
    max_pages=500,
    save_state=True,
    state_path="adaptive_state.json"
)
crawler = AdaptiveCrawler(query="machine learning", config=config)
await crawler.run()

# Resume after crash

resumed_crawler = AdaptiveCrawler(
    query="machine learning", 
    resume_from="adaptive_state.json"
)
await resumed_crawler.run()  # Continues from saved CrawlState

The CrawlState object serializes crawled URLs, knowledge base data, embeddings, and metrics, providing a complete snapshot of the adaptive crawling session.

Summary

  • Deep-crawling strategies accept a resume_state dictionary to restore visited, pending, and depths structures from crawl4ai/deep_crawling/bfs_strategy.py or dfs_strategy.py.
  • The on_state_change callback triggers after each URL, emitting a state dictionary suitable for external persistence.
  • AdaptiveCrawler utilizes the CrawlState class with save_state=True and resume_from="path.json" for full session recovery.
  • Always use atomic file writes (write to temp, then rename) to prevent state corruption during crashes.
  • State files contain only metadata (URLs and counters), not page content, ensuring lightweight serialization.

Frequently Asked Questions

What is the difference between resume_state and resume_from in crawl4ai?

resume_state is a dictionary parameter used by deep-crawling strategies (BFSDeepCrawlStrategy, DFSDeepCrawlStrategy) to restore traversal queues and visited sets immediately upon instantiation. resume_from is a file path string used by AdaptiveCrawler to load a complete CrawlState object from JSON, restoring the entire adaptive crawling session including embeddings and metrics.

How does BFSDeepCrawlStrategy handle state hydration?

The strategy's _arun_batch method checks for self._resume_state at initialization (lines 65-73 in crawl4ai/deep_crawling/bfs_strategy.py). If present, it reconstructs the internal visited set, pending list, depth map, and page counter from the dictionary values before beginning the crawl loop.

Can I change crawl parameters like max_depth when resuming?

Yes. The resume_state restores progress tracking (which URLs were visited and which are pending), but new limits such as max_depth or max_pages passed to the constructor take effect immediately. The crawler respects the new boundaries while preserving the existing queue state.

Is the state file format compatible across different crawl strategies?

No. The resume_state dictionaries differ between BFS, DFS, and BFF strategies due to their distinct internal structures (queue vs. stack vs. priority queue). Additionally, AdaptiveCrawler uses a separate CrawlState JSON schema. Always resume using the same strategy class that generated the state file.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →