# How to Use URL Seeding and Filtering for Focused Crawling in crawl4ai

> Master focused crawling in crawl4ai using URL seeding and filtering. Discover, filter, and extract targeted data efficiently with AsyncUrlSeeder and AsyncWebCrawler for precise web scraping.

- Repository: [UncleCode/crawl4ai](https://github.com/unclecode/crawl4ai)
- Tags: how-to-guide
- Published: 2026-03-05

---

**Use `AsyncUrlSeeder` with `SeedingConfig` to discover URLs from sitemaps or Common Crawl, apply glob patterns and BM25 relevance scoring, and feed filtered results into `AsyncWebCrawler` for targeted data extraction.**

The **crawl4ai** library provides a high-performance asynchronous framework for web crawling that supports intelligent discovery engines. By leveraging **URL seeding and filtering for focused crawling**, you can discover and prioritize only the most relevant pages from massive domains without wasting resources on irrelevant content. The implementation centers around the `AsyncUrlSeeder` class defined in [`crawl4ai/async_url_seeder.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/async_url_seeder.py) and its configuration object, enabling precise control over discovery sources, caching, and relevance scoring.

## Understanding URL Seeding in crawl4ai

The `AsyncUrlSeeder` class acts as an asynchronous discovery engine that pulls URLs from two primary public sources. In [`crawl4ai/async_url_seeder.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/async_url_seeder.py), the method `_from_sitemaps` resolves a domain's [`sitemap.xml`](https://github.com/unclecode/crawl4ai/blob/main/sitemap.xml) (or [`sitemap_index.xml`](https://github.com/unclecode/crawl4ai/blob/main/sitemap_index.xml)), downloads the XML or `.gz` files, and extracts every `<loc>` entry while optionally validating `<lastmod>` timestamps against a local cache. Alternatively, the `_from_cc` method streams the Common Crawl index (`https://index.commoncrawl.org/...`) for the requested domain and yields each discovered URL in JSON format.

Both sources can be combined or used individually based on the `source` field of `SeedingConfig`. Valid options include `"sitemap+cc"` (default), `"sitemap"`, or `"cc"`, allowing you to choose between comprehensive coverage and speed.

## Configuring the SeedingConfig Class

The `SeedingConfig` class, located in [`crawl4ai/async_configs.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/async_configs.py), provides every parameter needed to steer the discovery process. You can instantiate it directly or use the `.clone()` helper method for easy modification of base configurations.

Key parameters include:

- **`source`** – Choose from `"sitemap"`, `"cc"`, or `"sitemap+cc"` to control discovery sources.
- **`pattern`** – A glob-style filter (e.g., `*blog/*`, `*/api/*`) applied after URL discovery to retain only topic-specific paths.
- **`live_check`** – When `True`, performs a HEAD request to filter dead links before returning URLs.
- **`extract_head`** – Retrieves the `<head>` section to parse titles, meta tags, and JSON-LD; required for BM25 relevance scoring.
- **`query`** and **`score_threshold`** – Enable BM25-based relevance scoring against the extracted head content.
- **`max_urls`** – Sets an upper bound on returned URLs (`-1` for unlimited).
- **`concurrency`** and **`hits_per_sec`** – Control parallel HTTP calls and global rate limiting to maintain polite crawling.
- **`force`** – Bypasses the on-disk cache to re-seed after site updates.
- **`filter_nonsense_urls`** – Automatically drops utility files like [`robots.txt`](https://github.com/unclecode/crawl4ai/blob/main/robots.txt), [`sitemap.xml`](https://github.com/unclecode/crawl4ai/blob/main/sitemap.xml), and [`ads.txt`](https://github.com/unclecode/crawl4ai/blob/main/ads.txt).

```python
from crawl4ai.async_configs import SeedingConfig

seed_cfg = SeedingConfig(
    source="sitemap+cc",
    pattern="*/blog/*",
    live_check=True,
    extract_head=True,
    query="machine learning",
    score_threshold=0.2,
    max_urls=200,
)

```

## The Internal Seeding Pipeline

When you call `await seeder.urls(domain, config)`, `AsyncUrlSeeder` executes a fully asynchronous, back-pressured pipeline:

1. **Initialization** – Creates an `httpx.AsyncClient` and resolves the latest Common Crawl index ID via `_latest_index`.
2. **Source Parsing** – Splits the `source` string (e.g., `"sitemap+cc"`) and triggers the respective async generators (`_from_sitemaps` or `_from_cc`).
3. **Cache Handling** – Stores sitemap data in `~/.crawl4ai/seeder_cache/` as JSON. The `_is_cache_valid` method checks TTL and optional `<lastmod>` validation to decide between cache reads or refetching.
4. **Filtering** – Applies the `_match` function using `fnmatch` against the URL pattern and host part. If `filter_nonsense_urls` is enabled, internal utility URLs are discarded.
5. **Live-Check and Head Extraction** – The `_validate` method performs HEAD requests (or small GETs for `<head>` content), while `_parse_head` extracts metadata. This data attaches to the result dictionary.
6. **Scoring** – When a `query` is supplied, `_apply_bm25_scoring` ranks URLs by relevance, filters by `score_threshold`, and sorts the final output.

The pipeline uses an `asyncio.Queue` sized to approximately `concurrency * 100` to bound memory usage during high-throughput discovery.

## Integrating with AsyncWebCrawler

`AsyncWebCrawler` (in [`crawl4ai/async_webcrawler.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/async_webcrawler.py)) automatically creates an internal `AsyncUrlSeeder` when you supply a `SeedingConfig`. This enables focused crawling where only filtered, relevance-ranked URLs enter the crawling queue.

```python
import asyncio
from crawl4ai import AsyncWebCrawler, SeedingConfig

async def main():
    seed_cfg = SeedingConfig(
        source="sitemap",
        pattern="*/docs/*",
        live_check=False,
        extract_head=True,
        query="data science",
        score_threshold=0.3,
        max_urls=50,
    )
    crawler = AsyncWebCrawler(seeding_config=seed_cfg)

    data = await crawler.arun("https://example.com")
    print("Discovered URLs:", [u["url"] for u in data.urls])

    await crawler.close()

asyncio.run(main())

```

## Standalone URL Seeding

For scenarios requiring only URL discovery without full page crawling, interact with `AsyncUrlSeeder` directly:

```python
import asyncio
from crawl4ai.async_url_seeder import AsyncUrlSeeder
from crawl4ai.async_configs import SeedingConfig

async def demo():
    seeder = AsyncUrlSeeder()
    cfg = SeedingConfig(
        source="sitemap",
        pattern="*/news/*",
        live_check=False,
        extract_head=False,
        max_urls=20,
    )
    urls = await seeder.urls("nytimes.com", cfg)
    for u in urls:
        print(u["url"])
    await seeder.close()

asyncio.run(demo())

```

## Optimization Tips for Focused Crawling

| Goal | Configuration | Reason |
|------|---------------|--------|
| **Isolate blog posts** | `pattern="*/blog/*"` | Glob patterns filter paths containing specific directory names. |
| **Eliminate dead links** | `live_check=True` | HEAD requests prune 404/5xx errors before they reach the crawler. |
| **Prioritize topical content** | `extract_head=True`, `query="deep learning"`, `score_threshold=0.2` | BM25 scoring uses page titles and metadata to rank relevance. |
| **Accelerate large domains** | `source="cc"`, `force=False` | Common Crawl streaming is faster than sitemap parsing; caching prevents redundant downloads. |
| **Control request rates** | `hits_per_sec=3` | Prevents accidental aggressive traffic against target servers. |
| **Bound result sets** | `max_urls=100` | Guarantees finite memory usage and predictable downstream processing. |

## Summary

- **`AsyncUrlSeeder`** in [`crawl4ai/async_url_seeder.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/async_url_seeder.py) provides asynchronous URL discovery from sitemaps and Common Crawl indexes.
- **`SeedingConfig`** controls filtering via glob patterns, live link validation, BM25 relevance scoring, and rate limiting.
- The internal pipeline uses caching (`~/.crawl4ai/seeder_cache/`), `_match` filtering with `fnmatch`, and optional `_apply_bm25_scoring` for semantic ranking.
- **`AsyncWebCrawler`** accepts `SeedingConfig` to perform focused crawling, or you can use `AsyncUrlSeeder` standalone for pure discovery tasks.

## Frequently Asked Questions

### What is the difference between sitemap and Common Crawl sources?

**Sitemap** discovery downloads the domain's [`sitemap.xml`](https://github.com/unclecode/crawl4ai/blob/main/sitemap.xml) to extract explicitly listed URLs, while **Common Crawl (CC)** streams URLs from the public CC index for historical snapshots of the domain. Sitemaps provide current, officially published URLs, whereas CC offers broader coverage of indexed content that may no longer be in the current sitemap. You can combine both with `source="sitemap+cc"` for maximum completeness.

### How does BM25 relevance scoring work in crawl4ai?

BM25 scoring requires `extract_head=True` in your `SeedingConfig`. The seeder fetches each URL's `<head>` section via `_parse_head`, then feeds titles and meta descriptions into the `_apply_bm25_scoring` method. This ranks URLs against your `query` string; results below `score_threshold` are filtered out, and the remainder are sorted by relevance before being returned or passed to the crawler.

### Can I use URL seeding without crawling the pages immediately?

Yes. Instantiate `AsyncUrlSeeder` directly and call `await seeder.urls(domain, config)` to receive a list of filtered, validated URLs without invoking `AsyncWebCrawler`. This is useful for building URL lists for later batch processing or for analyzing site structure without downloading full page content.

### How do I prevent hitting the same sitemap repeatedly?

Enable smart caching by setting `force=False` (default) and configuring `cache_ttl_hours`. The seeder stores sitemap data in `~/.crawl4ai/seeder_cache/` and uses `_is_cache_valid` to check both the TTL and optional `<lastmod>` timestamps before refetching. This ensures you only download fresh sitemaps when they have actually changed on the server.