# How Adaptive Crawling Learns Site Patterns Automatically in crawl4ai

> Discover how crawl4ai's adaptive crawling automatically learns site patterns. It scores links for relevance, novelty, and semantic gap coverage to optimize your crawls efficiently.

- Repository: [UncleCode/crawl4ai](https://github.com/unclecode/crawl4ai)
- Tags: deep-dive
- Published: 2026-03-05

---

**crawl4ai's adaptive crawler automatically learns a website's most informative URL patterns by iteratively scoring links based on relevance, novelty, and semantic gap coverage, halting when confidence thresholds indicate sufficient knowledge has been gathered.**

The open-source crawl4ai library includes an intelligent adaptive crawler that eliminates manual sitemap configuration by discovering high-value pages through continuous learning. Instead of following static rules, the system uses statistical or embedding-based strategies to identify which site sections—such as `/docs/` or `/api/` endpoints—contain the most relevant information for a given query.

## The Five-Stage Learning Loop

The adaptive crawling process operates as a feedback loop that refines its understanding of a site's structure with every page visited. According to the source code in [`crawl4ai/adaptive_crawler.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/adaptive_crawler.py), the learning cycle consists of:

1. **Initial query expansion** – The crawler creates a semantic cloud of related queries to broaden the search scope.
2. **Content extraction and term statistics** – Each crawled page updates term-frequency, document-frequency, and novelty counters stored in the crawl state.
3. **Confidence estimation** – A multi-metric score evaluates how much of the information space is already covered, using either statistical coverage or embedding-based similarity.
4. **Link ranking** – Pending links are scored by relevance, novelty, and authority, with the embedding mode further refining scores by measuring how well a link fills semantic gaps in the query cloud.
5. **Stopping decision** – When confidence exceeds a threshold, saturation is high, or resource limits are hit, the crawler automatically terminates.

This process repeats until the crawler either reaches the confidence target or exhausts useful links, implicitly learning which URL patterns are most promising for the given query.

## AdaptiveCrawler Orchestration

The top-level class driving this loop resides in [`crawl4ai/adaptive_crawler.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/adaptive_crawler.py). The `AdaptiveCrawler` class manages the entire lifecycle of the learning process:

- **Initialization** (lines 1271‑1286): Sets up a `CrawlState`, chooses a strategy (`statistical` or `embedding`), and validates the configuration.
- **`digest` method** (lines 1312‑1344): The entry point that loads or creates a fresh `CrawlState`, expands the query space for the embedding strategy, and runs the main adaptive loop.
- **Main adaptive loop** (lines 1365‑1449): Each iteration calculates confidence, checks stopping conditions, ranks links, crawls the top-K links, and updates the state via `self.strategy.update_state`.
- **State persistence** (lines 1426‑1429): Optionally saves the crawl state after each iteration, enabling resume functionality.

## CrawlState: The Learning Memory

The `CrawlState` class (lines 25‑45 in [`adaptive_crawler.py`](https://github.com/unclecode/crawl4ai/blob/main/adaptive_crawler.py)) serves as the persistent memory for everything the crawler learns:

- `term_frequencies`, `document_frequencies`, `documents_with_terms` – Classic TF/IDF statistics for the statistical strategy.
- `new_terms_history`, `crawl_order` – Track novelty over time, used for calculating saturation metrics.
- `kb_embeddings`, `query_embeddings` – Vector representations for the embedding strategy.

The state is serializable via `save` and `load` methods, allowing the crawler to resume learning from a previous session without re-crawling known pages.

## Learning Strategies

crawl4ai implements two distinct strategies for learning site patterns, selectable via `AdaptiveConfig`.

### StatisticalStrategy (TF/IDF-Based)

The `StatisticalStrategy` class implements classic information retrieval metrics to guide crawling decisions:

- **Coverage** (lines 112‑128): Measures how many query terms appear across documents using `_calculate_coverage`.
- **Consistency**: Calculates Jaccard overlap of term sets between documents via `_calculate_consistency`.
- **Saturation**: Tracks the rate of new term discovery using `_calculate_saturation`.

These three metrics are combined with weights (0.4 coverage + 0.3 consistency + 0.3 saturation) to produce a composite confidence score. For link ranking, the strategy uses BM25-like relevance (`_calculate_relevance`) and novelty scores (`_calculate_novelty`) at lines 142‑150.

### EmbeddingStrategy (Semantic-Based)

The `EmbeddingStrategy` learns a semantic map of the query space and crawled knowledge base:

- **Query expansion**: The `map_query_semantic_space` method calls an LLM to generate synthetic query variations, then embeds them via `_get_embeddings`.
- **Coverage shape** (lines 83‑89): `compute_coverage_shape` builds a centroid and radius model for the query cloud.
- **Gap detection**: `find_coverage_gaps` computes minimum cosine distances from query embeddings to current knowledge-base embeddings using vectorized `np.dot` operations.
- **Link selection**: `select_links_for_expansion` scores candidates by how much they would shrink semantic gaps while penalizing redundancy via `embedding_overlap_threshold`.
- **Confidence calculation** (lines 67‑84): `calculate_confidence` reports the mean best similarity between query embeddings and knowledge-base embeddings.

This strategy learns site patterns implicitly by repeatedly selecting links whose content fills uncovered semantic regions, causing the crawler to gravitate toward URL structures that host missing information.

## Automatic Stopping Criteria

Both strategies expose a `should_stop` method that automatically terminates crawling when learning objectives are met:

- **Statistical** (lines 290‑298): Stops when confidence ≥ `confidence_threshold`, maximum pages reached, or saturation ≥ `config.saturation_threshold`.
- **Embedding** (lines 138‑155): Adds a minimum relevance guard (`embedding_min_confidence_threshold`) and checks for stagnation in the confidence history.

These criteria ensure the crawler halts once it has learned enough about the site to answer the original query with the desired confidence level, preventing redundant crawling of similar pages.

## Practical Implementation

To use the adaptive crawler with the default statistical strategy:

```python
from crawl4ai import AdaptiveCrawler, AsyncWebCrawler

# Create an adaptive crawler (default uses statistical strategy)

adaptive = AdaptiveCrawler(
    crawler=AsyncWebCrawler(),                # optional – created automatically if None

    config=None                               # None → defaults (confidence 0.7, max_pages 20)

)

# Run adaptive crawling on a site

state = await adaptive.digest(
    start_url="https://example.com/docs/intro",
    query="how to authenticate API requests"
)

print(f"Pages crawled: {len(state.crawled_urls)}")
print(f"Final confidence: {state.metrics['confidence']:.2%}")
print("Top relevant pages:")
for doc in state.knowledge_base[:3]:
    print(f"- {doc.url}")

```

To switch to the embedding strategy for semantic pattern learning:

```python
from crawl4ai import AdaptiveCrawler, AdaptiveConfig

embed_cfg = AdaptiveConfig(
    strategy="embedding",
    confidence_threshold=0.75,
    max_pages=30,
    embedding_model="sentence-transformers/all-MiniLM-L6-v2"
)

adaptive = AdaptiveCrawler(config=embed_cfg)
state = await adaptive.digest(
    start_url="https://example.com/api",
    query="list all available endpoints"
)

```

The embedding-based run automatically generates related queries (e.g., "list users endpoint", "retrieve order details") and focuses on URLs that close semantic gaps, thereby learning the site's URL patterns that host the needed information without explicit configuration.

## Summary

- **crawl4ai** implements adaptive crawling through the `AdaptiveCrawler` class in [`crawl4ai/adaptive_crawler.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/adaptive_crawler.py), which orchestrates a continuous learning loop.
- The system stores learned patterns in `CrawlState`, tracking term statistics or embeddings depending on the selected strategy.
- **StatisticalStrategy** uses TF/IDF metrics (coverage, consistency, saturation) and BM25 relevance scoring to identify informative pages.
- **EmbeddingStrategy** generates query variations and uses cosine similarity to detect and fill semantic gaps in the knowledge base.
- Links are ranked by their ability to provide novel, relevant information, automatically directing the crawler toward high-value URL patterns like `/docs/` or `/api/` sections.
- Crawling stops automatically when confidence thresholds, saturation limits, or resource constraints are met.

## Frequently Asked Questions

### What is the difference between statistical and embedding strategies in crawl4ai?

The **statistical strategy** relies on term-frequency statistics (TF/IDF) to measure coverage, consistency, and saturation of query terms across crawled documents. The **embedding strategy** uses vector embeddings to map the semantic space of queries and content, identifying gaps by measuring cosine distances between query embeddings and the knowledge base. The statistical approach is lighter and faster, while the embedding approach provides deeper semantic understanding for complex queries.

### How does crawl4ai determine which URLs to crawl next?

After each batch of pages, the system ranks pending links using a composite score that weighs **relevance** (BM25 or semantic similarity), **novelty** (new terms or distant embeddings), and **authority**. In embedding mode, links are specifically scored by how much they would reduce semantic gaps in the query cloud. The top-K highest-scoring links are selected for the next iteration, causing the crawler to automatically favor URL patterns that consistently provide high-value information.

### Can the adaptive crawler resume from a previous session?

Yes. The `CrawlState` object (defined at lines 25‑45 in [`adaptive_crawler.py`](https://github.com/unclecode/crawl4ai/blob/main/adaptive_crawler.py)) is serializable via its `save` and `load` methods. When calling `adaptive.digest()`, you can pass a `state` parameter to resume from a previous crawl, or the system will automatically persist state between iterations if configured to do so (lines 1426‑1429). This prevents re-crawling known pages and preserves learned term statistics or embeddings.

### What signals indicate that the adaptive crawler has learned enough about a site?

The crawler monitors three primary stopping signals: **confidence score** (when the weighted coverage metrics exceed `confidence_threshold`), **saturation** (when the rate of new term discovery drops below `saturation_threshold`), and **resource limits** (maximum page count or depth). In embedding mode, an additional **minimum relevance** guard ensures the semantic similarity between queries and the knowledge base meets `embedding_min_confidence_threshold` before stopping.