How to Use URL Seeding and Filtering for Focused Crawling in crawl4ai

Use AsyncUrlSeeder with SeedingConfig to discover URLs from sitemaps or Common Crawl, apply glob patterns and BM25 relevance scoring, and feed filtered results into AsyncWebCrawler for targeted data extraction.

The crawl4ai library provides a high-performance asynchronous framework for web crawling that supports intelligent discovery engines. By leveraging URL seeding and filtering for focused crawling, you can discover and prioritize only the most relevant pages from massive domains without wasting resources on irrelevant content. The implementation centers around the AsyncUrlSeeder class defined in crawl4ai/async_url_seeder.py and its configuration object, enabling precise control over discovery sources, caching, and relevance scoring.

Understanding URL Seeding in crawl4ai

The AsyncUrlSeeder class acts as an asynchronous discovery engine that pulls URLs from two primary public sources. In crawl4ai/async_url_seeder.py, the method _from_sitemaps resolves a domain's sitemap.xml (or sitemap_index.xml), downloads the XML or .gz files, and extracts every <loc> entry while optionally validating <lastmod> timestamps against a local cache. Alternatively, the _from_cc method streams the Common Crawl index (https://index.commoncrawl.org/...) for the requested domain and yields each discovered URL in JSON format.

Both sources can be combined or used individually based on the source field of SeedingConfig. Valid options include "sitemap+cc" (default), "sitemap", or "cc", allowing you to choose between comprehensive coverage and speed.

Configuring the SeedingConfig Class

The SeedingConfig class, located in crawl4ai/async_configs.py, provides every parameter needed to steer the discovery process. You can instantiate it directly or use the .clone() helper method for easy modification of base configurations.

Key parameters include:

  • source – Choose from "sitemap", "cc", or "sitemap+cc" to control discovery sources.
  • pattern – A glob-style filter (e.g., *blog/*, */api/*) applied after URL discovery to retain only topic-specific paths.
  • live_check – When True, performs a HEAD request to filter dead links before returning URLs.
  • extract_head – Retrieves the <head> section to parse titles, meta tags, and JSON-LD; required for BM25 relevance scoring.
  • query and score_threshold – Enable BM25-based relevance scoring against the extracted head content.
  • max_urls – Sets an upper bound on returned URLs (-1 for unlimited).
  • concurrency and hits_per_sec – Control parallel HTTP calls and global rate limiting to maintain polite crawling.
  • force – Bypasses the on-disk cache to re-seed after site updates.
  • filter_nonsense_urls – Automatically drops utility files like robots.txt, sitemap.xml, and ads.txt.
from crawl4ai.async_configs import SeedingConfig

seed_cfg = SeedingConfig(
    source="sitemap+cc",
    pattern="*/blog/*",
    live_check=True,
    extract_head=True,
    query="machine learning",
    score_threshold=0.2,
    max_urls=200,
)

The Internal Seeding Pipeline

When you call await seeder.urls(domain, config), AsyncUrlSeeder executes a fully asynchronous, back-pressured pipeline:

  1. Initialization – Creates an httpx.AsyncClient and resolves the latest Common Crawl index ID via _latest_index.
  2. Source Parsing – Splits the source string (e.g., "sitemap+cc") and triggers the respective async generators (_from_sitemaps or _from_cc).
  3. Cache Handling – Stores sitemap data in ~/.crawl4ai/seeder_cache/ as JSON. The _is_cache_valid method checks TTL and optional <lastmod> validation to decide between cache reads or refetching.
  4. Filtering – Applies the _match function using fnmatch against the URL pattern and host part. If filter_nonsense_urls is enabled, internal utility URLs are discarded.
  5. Live-Check and Head Extraction – The _validate method performs HEAD requests (or small GETs for <head> content), while _parse_head extracts metadata. This data attaches to the result dictionary.
  6. Scoring – When a query is supplied, _apply_bm25_scoring ranks URLs by relevance, filters by score_threshold, and sorts the final output.

The pipeline uses an asyncio.Queue sized to approximately concurrency * 100 to bound memory usage during high-throughput discovery.

Integrating with AsyncWebCrawler

AsyncWebCrawler (in crawl4ai/async_webcrawler.py) automatically creates an internal AsyncUrlSeeder when you supply a SeedingConfig. This enables focused crawling where only filtered, relevance-ranked URLs enter the crawling queue.

import asyncio
from crawl4ai import AsyncWebCrawler, SeedingConfig

async def main():
    seed_cfg = SeedingConfig(
        source="sitemap",
        pattern="*/docs/*",
        live_check=False,
        extract_head=True,
        query="data science",
        score_threshold=0.3,
        max_urls=50,
    )
    crawler = AsyncWebCrawler(seeding_config=seed_cfg)

    data = await crawler.arun("https://example.com")
    print("Discovered URLs:", [u["url"] for u in data.urls])

    await crawler.close()

asyncio.run(main())

Standalone URL Seeding

For scenarios requiring only URL discovery without full page crawling, interact with AsyncUrlSeeder directly:

import asyncio
from crawl4ai.async_url_seeder import AsyncUrlSeeder
from crawl4ai.async_configs import SeedingConfig

async def demo():
    seeder = AsyncUrlSeeder()
    cfg = SeedingConfig(
        source="sitemap",
        pattern="*/news/*",
        live_check=False,
        extract_head=False,
        max_urls=20,
    )
    urls = await seeder.urls("nytimes.com", cfg)
    for u in urls:
        print(u["url"])
    await seeder.close()

asyncio.run(demo())

Optimization Tips for Focused Crawling

Goal Configuration Reason
Isolate blog posts pattern="*/blog/*" Glob patterns filter paths containing specific directory names.
Eliminate dead links live_check=True HEAD requests prune 404/5xx errors before they reach the crawler.
Prioritize topical content extract_head=True, query="deep learning", score_threshold=0.2 BM25 scoring uses page titles and metadata to rank relevance.
Accelerate large domains source="cc", force=False Common Crawl streaming is faster than sitemap parsing; caching prevents redundant downloads.
Control request rates hits_per_sec=3 Prevents accidental aggressive traffic against target servers.
Bound result sets max_urls=100 Guarantees finite memory usage and predictable downstream processing.

Summary

  • AsyncUrlSeeder in crawl4ai/async_url_seeder.py provides asynchronous URL discovery from sitemaps and Common Crawl indexes.
  • SeedingConfig controls filtering via glob patterns, live link validation, BM25 relevance scoring, and rate limiting.
  • The internal pipeline uses caching (~/.crawl4ai/seeder_cache/), _match filtering with fnmatch, and optional _apply_bm25_scoring for semantic ranking.
  • AsyncWebCrawler accepts SeedingConfig to perform focused crawling, or you can use AsyncUrlSeeder standalone for pure discovery tasks.

Frequently Asked Questions

What is the difference between sitemap and Common Crawl sources?

Sitemap discovery downloads the domain's sitemap.xml to extract explicitly listed URLs, while Common Crawl (CC) streams URLs from the public CC index for historical snapshots of the domain. Sitemaps provide current, officially published URLs, whereas CC offers broader coverage of indexed content that may no longer be in the current sitemap. You can combine both with source="sitemap+cc" for maximum completeness.

How does BM25 relevance scoring work in crawl4ai?

BM25 scoring requires extract_head=True in your SeedingConfig. The seeder fetches each URL's <head> section via _parse_head, then feeds titles and meta descriptions into the _apply_bm25_scoring method. This ranks URLs against your query string; results below score_threshold are filtered out, and the remainder are sorted by relevance before being returned or passed to the crawler.

Can I use URL seeding without crawling the pages immediately?

Yes. Instantiate AsyncUrlSeeder directly and call await seeder.urls(domain, config) to receive a list of filtered, validated URLs without invoking AsyncWebCrawler. This is useful for building URL lists for later batch processing or for analyzing site structure without downloading full page content.

How do I prevent hitting the same sitemap repeatedly?

Enable smart caching by setting force=False (default) and configuring cache_ttl_hours. The seeder stores sitemap data in ~/.crawl4ai/seeder_cache/ and uses _is_cache_valid to check both the TTL and optional <lastmod> timestamps before refetching. This ensures you only download fresh sitemaps when they have actually changed on the server.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →