# How to Implement Crash Recovery with resume_state for Long Crawls in crawl4ai

> Implement crash recovery for long crawls in crawl4ai using resume_state. Prevent data loss and restore crawl progress seamlessly with our guide.

- Repository: [UncleCode/crawl4ai](https://github.com/unclecode/crawl4ai)
- Tags: how-to-guide
- Published: 2026-03-05

---

**Use the `resume_state` parameter in deep-crawling strategies (BFS, DFS, BFF) or `resume_from` in AdaptiveCrawler to persist and restore crawl progress, preventing data loss during crashes.**

The `unclecode/crawl4ai` library provides built-in crash-recovery mechanisms that allow long-running crawls to survive process interruptions. By leveraging state serialization in `BFSDeepCrawlStrategy` or the `CrawlState` class for adaptive crawling, you can resume exactly where the crawler left off without reprocessing visited URLs.

## How Resume State Works in Deep Crawling Strategies

The deep-crawling architecture in `crawl4ai` maintains internal traversal queues, visited sets, and depth counters. When you provide a `resume_state` dictionary to the strategy constructor, the crawler hydrates these structures before processing begins.

### State Hydration in BFS and DFS

In [`crawl4ai/deep_crawling/bfs_strategy.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/deep_crawling/bfs_strategy.py), the `_arun_batch` method (lines 65-73) checks for `self._resume_state` at startup. If present, it reconstructs the `visited` set, `pending` list, `depths` map, and `pages_crawled` counter. The DFS implementation in [`crawl4ai/deep_crawling/dfs_strategy.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/deep_crawling/dfs_strategy.py) follows an identical pattern but restores a `stack` instead of a queue (lines 41-50).

```python

# Simplified internal logic from bfs_strategy.py

if self._resume_state:
    visited = set(self._resume_state.get("visited", []))
    pending = list(self._resume_state.get("pending", []))
    depths = self._resume_state.get("depths", {})
    pages_crawled = self._resume_state.get("pages_crawled", 0)

```

### The State Capture Callback

After each URL completes, strategies invoke the `on_state_change` callback with a state dictionary. This callback, referenced in [`bfs_strategy.py`](https://github.com/unclecode/crawl4ai/blob/main/bfs_strategy.py) (lines 13-23), receives the current `visited`, `pending`, `depths`, and `pages_crawled` values. You persist this dictionary externally—typically to JSON—to create a restore point.

## Step-by-Step Implementation for Long Crawls

Implementing crash recovery requires two components: a persistence callback to save state after every page, and logic to load that state when initializing the crawler.

### Creating an Async State Persistence Callback

Write an asynchronous function that atomically writes the state dictionary to disk. Using `aiofiles` prevents blocking the event loop during I/O operations.

```python
import json
import aiofiles
from pathlib import Path

async def persist_state(state: dict, path: Path = Path("crawl_state.json")):
    """Atomically save crawl state to prevent corruption during crashes."""
    tmp = path.with_suffix(".tmp")
    async with aiofiles.open(tmp, "w") as f:
        await f.write(json.dumps(state, indent=2))
    await aiofiles.os.rename(tmp, path)

```

### Resuming a BFS Deep Crawl

Load the saved state from disk and pass it to the `BFSDeepCrawlStrategy` constructor via the `resume_state` parameter. Assign the `persist_state` function to `on_state_change` to enable automatic state updates.

```python
import asyncio
import json
from pathlib import Path
from crawl4ai import AsyncWebCrawler
from crawl4ai.deep_crawling.bfs_strategy import BFSDeepCrawlStrategy
from crawl4ai.async_configs import CrawlerRunConfig

async def main():
    state_path = Path("crawl_state.json")
    
    # Load previous state if resuming

    resume_state = None
    if state_path.exists():
        resume_state = json.loads(state_path.read_text())
    
    # Initialize strategy with resume capability

    strategy = BFSDeepCrawlStrategy(
        max_depth=5,
        max_pages=2000,
        resume_state=resume_state,
        on_state_change=lambda s: persist_state(s, state_path)
    )
    
    async with AsyncWebCrawler() as crawler:
        config = CrawlerRunConfig(timeout=30)
        results = await strategy.arun("https://example.com", crawler, config)
        print(f"Retrieved {len(results)} pages")

if __name__ == "__main__":
    asyncio.run(main())

```

If the process crashes, restart the script to reload [`crawl_state.json`](https://github.com/unclecode/crawl4ai/blob/main/crawl_state.json) and continue from the last persisted URL.

## Crash Recovery for Adaptive Crawling

The `AdaptiveCrawler` class uses a different mechanism based on the `CrawlState` class defined in [`crawl4ai/adaptive_crawler.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/adaptive_crawler.py) (lines 53-88). Instead of a dictionary callback, you save the entire session state to JSON using `CrawlState.save()`, then restore it via the `resume_from` parameter in the constructor (lines 1311-1315).

```python
from crawl4ai import AdaptiveCrawler, AdaptiveConfig

# First run - auto-saves state

config = AdaptiveConfig(
    max_pages=500,
    save_state=True,
    state_path="adaptive_state.json"
)
crawler = AdaptiveCrawler(query="machine learning", config=config)
await crawler.run()

# Resume after crash

resumed_crawler = AdaptiveCrawler(
    query="machine learning", 
    resume_from="adaptive_state.json"
)
await resumed_crawler.run()  # Continues from saved CrawlState

```

The `CrawlState` object serializes crawled URLs, knowledge base data, embeddings, and metrics, providing a complete snapshot of the adaptive crawling session.

## Summary

- **Deep-crawling strategies** accept a `resume_state` dictionary to restore `visited`, `pending`, and `depths` structures from [`crawl4ai/deep_crawling/bfs_strategy.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/deep_crawling/bfs_strategy.py) or [`dfs_strategy.py`](https://github.com/unclecode/crawl4ai/blob/main/dfs_strategy.py).
- The **`on_state_change` callback** triggers after each URL, emitting a state dictionary suitable for external persistence.
- **AdaptiveCrawler** utilizes the `CrawlState` class with `save_state=True` and `resume_from="path.json"` for full session recovery.
- Always use **atomic file writes** (write to temp, then rename) to prevent state corruption during crashes.
- State files contain only metadata (URLs and counters), not page content, ensuring lightweight serialization.

## Frequently Asked Questions

### What is the difference between resume_state and resume_from in crawl4ai?

**`resume_state`** is a dictionary parameter used by deep-crawling strategies (`BFSDeepCrawlStrategy`, `DFSDeepCrawlStrategy`) to restore traversal queues and visited sets immediately upon instantiation. **`resume_from`** is a file path string used by `AdaptiveCrawler` to load a complete `CrawlState` object from JSON, restoring the entire adaptive crawling session including embeddings and metrics.

### How does BFSDeepCrawlStrategy handle state hydration?

The strategy's `_arun_batch` method checks for `self._resume_state` at initialization (lines 65-73 in [`crawl4ai/deep_crawling/bfs_strategy.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/deep_crawling/bfs_strategy.py)). If present, it reconstructs the internal `visited` set, `pending` list, depth map, and page counter from the dictionary values before beginning the crawl loop.

### Can I change crawl parameters like max_depth when resuming?

Yes. The `resume_state` restores progress tracking (which URLs were visited and which are pending), but new limits such as `max_depth` or `max_pages` passed to the constructor take effect immediately. The crawler respects the new boundaries while preserving the existing queue state.

### Is the state file format compatible across different crawl strategies?

No. The `resume_state` dictionaries differ between BFS, DFS, and BFF strategies due to their distinct internal structures (queue vs. stack vs. priority queue). Additionally, `AdaptiveCrawler` uses a separate `CrawlState` JSON schema. Always resume using the same strategy class that generated the state file.