How Adaptive Crawling Learns Site Patterns Automatically in crawl4ai

crawl4ai's adaptive crawler automatically learns a website's most informative URL patterns by iteratively scoring links based on relevance, novelty, and semantic gap coverage, halting when confidence thresholds indicate sufficient knowledge has been gathered.

The open-source crawl4ai library includes an intelligent adaptive crawler that eliminates manual sitemap configuration by discovering high-value pages through continuous learning. Instead of following static rules, the system uses statistical or embedding-based strategies to identify which site sections—such as /docs/ or /api/ endpoints—contain the most relevant information for a given query.

The Five-Stage Learning Loop

The adaptive crawling process operates as a feedback loop that refines its understanding of a site's structure with every page visited. According to the source code in crawl4ai/adaptive_crawler.py, the learning cycle consists of:

  1. Initial query expansion – The crawler creates a semantic cloud of related queries to broaden the search scope.
  2. Content extraction and term statistics – Each crawled page updates term-frequency, document-frequency, and novelty counters stored in the crawl state.
  3. Confidence estimation – A multi-metric score evaluates how much of the information space is already covered, using either statistical coverage or embedding-based similarity.
  4. Link ranking – Pending links are scored by relevance, novelty, and authority, with the embedding mode further refining scores by measuring how well a link fills semantic gaps in the query cloud.
  5. Stopping decision – When confidence exceeds a threshold, saturation is high, or resource limits are hit, the crawler automatically terminates.

This process repeats until the crawler either reaches the confidence target or exhausts useful links, implicitly learning which URL patterns are most promising for the given query.

AdaptiveCrawler Orchestration

The top-level class driving this loop resides in crawl4ai/adaptive_crawler.py. The AdaptiveCrawler class manages the entire lifecycle of the learning process:

  • Initialization (lines 1271‑1286): Sets up a CrawlState, chooses a strategy (statistical or embedding), and validates the configuration.
  • digest method (lines 1312‑1344): The entry point that loads or creates a fresh CrawlState, expands the query space for the embedding strategy, and runs the main adaptive loop.
  • Main adaptive loop (lines 1365‑1449): Each iteration calculates confidence, checks stopping conditions, ranks links, crawls the top-K links, and updates the state via self.strategy.update_state.
  • State persistence (lines 1426‑1429): Optionally saves the crawl state after each iteration, enabling resume functionality.

CrawlState: The Learning Memory

The CrawlState class (lines 25‑45 in adaptive_crawler.py) serves as the persistent memory for everything the crawler learns:

  • term_frequencies, document_frequencies, documents_with_terms – Classic TF/IDF statistics for the statistical strategy.
  • new_terms_history, crawl_order – Track novelty over time, used for calculating saturation metrics.
  • kb_embeddings, query_embeddings – Vector representations for the embedding strategy.

The state is serializable via save and load methods, allowing the crawler to resume learning from a previous session without re-crawling known pages.

Learning Strategies

crawl4ai implements two distinct strategies for learning site patterns, selectable via AdaptiveConfig.

StatisticalStrategy (TF/IDF-Based)

The StatisticalStrategy class implements classic information retrieval metrics to guide crawling decisions:

  • Coverage (lines 112‑128): Measures how many query terms appear across documents using _calculate_coverage.
  • Consistency: Calculates Jaccard overlap of term sets between documents via _calculate_consistency.
  • Saturation: Tracks the rate of new term discovery using _calculate_saturation.

These three metrics are combined with weights (0.4 coverage + 0.3 consistency + 0.3 saturation) to produce a composite confidence score. For link ranking, the strategy uses BM25-like relevance (_calculate_relevance) and novelty scores (_calculate_novelty) at lines 142‑150.

EmbeddingStrategy (Semantic-Based)

The EmbeddingStrategy learns a semantic map of the query space and crawled knowledge base:

  • Query expansion: The map_query_semantic_space method calls an LLM to generate synthetic query variations, then embeds them via _get_embeddings.
  • Coverage shape (lines 83‑89): compute_coverage_shape builds a centroid and radius model for the query cloud.
  • Gap detection: find_coverage_gaps computes minimum cosine distances from query embeddings to current knowledge-base embeddings using vectorized np.dot operations.
  • Link selection: select_links_for_expansion scores candidates by how much they would shrink semantic gaps while penalizing redundancy via embedding_overlap_threshold.
  • Confidence calculation (lines 67‑84): calculate_confidence reports the mean best similarity between query embeddings and knowledge-base embeddings.

This strategy learns site patterns implicitly by repeatedly selecting links whose content fills uncovered semantic regions, causing the crawler to gravitate toward URL structures that host missing information.

Automatic Stopping Criteria

Both strategies expose a should_stop method that automatically terminates crawling when learning objectives are met:

  • Statistical (lines 290‑298): Stops when confidence ≥ confidence_threshold, maximum pages reached, or saturation ≥ config.saturation_threshold.
  • Embedding (lines 138‑155): Adds a minimum relevance guard (embedding_min_confidence_threshold) and checks for stagnation in the confidence history.

These criteria ensure the crawler halts once it has learned enough about the site to answer the original query with the desired confidence level, preventing redundant crawling of similar pages.

Practical Implementation

To use the adaptive crawler with the default statistical strategy:

from crawl4ai import AdaptiveCrawler, AsyncWebCrawler

# Create an adaptive crawler (default uses statistical strategy)

adaptive = AdaptiveCrawler(
    crawler=AsyncWebCrawler(),                # optional – created automatically if None

    config=None                               # None → defaults (confidence 0.7, max_pages 20)

)

# Run adaptive crawling on a site

state = await adaptive.digest(
    start_url="https://example.com/docs/intro",
    query="how to authenticate API requests"
)

print(f"Pages crawled: {len(state.crawled_urls)}")
print(f"Final confidence: {state.metrics['confidence']:.2%}")
print("Top relevant pages:")
for doc in state.knowledge_base[:3]:
    print(f"- {doc.url}")

To switch to the embedding strategy for semantic pattern learning:

from crawl4ai import AdaptiveCrawler, AdaptiveConfig

embed_cfg = AdaptiveConfig(
    strategy="embedding",
    confidence_threshold=0.75,
    max_pages=30,
    embedding_model="sentence-transformers/all-MiniLM-L6-v2"
)

adaptive = AdaptiveCrawler(config=embed_cfg)
state = await adaptive.digest(
    start_url="https://example.com/api",
    query="list all available endpoints"
)

The embedding-based run automatically generates related queries (e.g., "list users endpoint", "retrieve order details") and focuses on URLs that close semantic gaps, thereby learning the site's URL patterns that host the needed information without explicit configuration.

Summary

  • crawl4ai implements adaptive crawling through the AdaptiveCrawler class in crawl4ai/adaptive_crawler.py, which orchestrates a continuous learning loop.
  • The system stores learned patterns in CrawlState, tracking term statistics or embeddings depending on the selected strategy.
  • StatisticalStrategy uses TF/IDF metrics (coverage, consistency, saturation) and BM25 relevance scoring to identify informative pages.
  • EmbeddingStrategy generates query variations and uses cosine similarity to detect and fill semantic gaps in the knowledge base.
  • Links are ranked by their ability to provide novel, relevant information, automatically directing the crawler toward high-value URL patterns like /docs/ or /api/ sections.
  • Crawling stops automatically when confidence thresholds, saturation limits, or resource constraints are met.

Frequently Asked Questions

What is the difference between statistical and embedding strategies in crawl4ai?

The statistical strategy relies on term-frequency statistics (TF/IDF) to measure coverage, consistency, and saturation of query terms across crawled documents. The embedding strategy uses vector embeddings to map the semantic space of queries and content, identifying gaps by measuring cosine distances between query embeddings and the knowledge base. The statistical approach is lighter and faster, while the embedding approach provides deeper semantic understanding for complex queries.

How does crawl4ai determine which URLs to crawl next?

After each batch of pages, the system ranks pending links using a composite score that weighs relevance (BM25 or semantic similarity), novelty (new terms or distant embeddings), and authority. In embedding mode, links are specifically scored by how much they would reduce semantic gaps in the query cloud. The top-K highest-scoring links are selected for the next iteration, causing the crawler to automatically favor URL patterns that consistently provide high-value information.

Can the adaptive crawler resume from a previous session?

Yes. The CrawlState object (defined at lines 25‑45 in adaptive_crawler.py) is serializable via its save and load methods. When calling adaptive.digest(), you can pass a state parameter to resume from a previous crawl, or the system will automatically persist state between iterations if configured to do so (lines 1426‑1429). This prevents re-crawling known pages and preserves learned term statistics or embeddings.

What signals indicate that the adaptive crawler has learned enough about a site?

The crawler monitors three primary stopping signals: confidence score (when the weighted coverage metrics exceed confidence_threshold), saturation (when the rate of new term discovery drops below saturation_threshold), and resource limits (maximum page count or depth). In embedding mode, an additional minimum relevance guard ensures the semantic similarity between queries and the knowledge base meets embedding_min_confidence_threshold before stopping.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →