How to Use URL Seeding and Filtering for Focused Crawling in crawl4ai
Use AsyncUrlSeeder with SeedingConfig to discover URLs from sitemaps or Common Crawl, apply glob patterns and BM25 relevance scoring, and feed filtered results into AsyncWebCrawler for targeted data extraction.
The crawl4ai library provides a high-performance asynchronous framework for web crawling that supports intelligent discovery engines. By leveraging URL seeding and filtering for focused crawling, you can discover and prioritize only the most relevant pages from massive domains without wasting resources on irrelevant content. The implementation centers around the AsyncUrlSeeder class defined in crawl4ai/async_url_seeder.py and its configuration object, enabling precise control over discovery sources, caching, and relevance scoring.
Understanding URL Seeding in crawl4ai
The AsyncUrlSeeder class acts as an asynchronous discovery engine that pulls URLs from two primary public sources. In crawl4ai/async_url_seeder.py, the method _from_sitemaps resolves a domain's sitemap.xml (or sitemap_index.xml), downloads the XML or .gz files, and extracts every <loc> entry while optionally validating <lastmod> timestamps against a local cache. Alternatively, the _from_cc method streams the Common Crawl index (https://index.commoncrawl.org/...) for the requested domain and yields each discovered URL in JSON format.
Both sources can be combined or used individually based on the source field of SeedingConfig. Valid options include "sitemap+cc" (default), "sitemap", or "cc", allowing you to choose between comprehensive coverage and speed.
Configuring the SeedingConfig Class
The SeedingConfig class, located in crawl4ai/async_configs.py, provides every parameter needed to steer the discovery process. You can instantiate it directly or use the .clone() helper method for easy modification of base configurations.
Key parameters include:
source– Choose from"sitemap","cc", or"sitemap+cc"to control discovery sources.pattern– A glob-style filter (e.g.,*blog/*,*/api/*) applied after URL discovery to retain only topic-specific paths.live_check– WhenTrue, performs a HEAD request to filter dead links before returning URLs.extract_head– Retrieves the<head>section to parse titles, meta tags, and JSON-LD; required for BM25 relevance scoring.queryandscore_threshold– Enable BM25-based relevance scoring against the extracted head content.max_urls– Sets an upper bound on returned URLs (-1for unlimited).concurrencyandhits_per_sec– Control parallel HTTP calls and global rate limiting to maintain polite crawling.force– Bypasses the on-disk cache to re-seed after site updates.filter_nonsense_urls– Automatically drops utility files likerobots.txt,sitemap.xml, andads.txt.
from crawl4ai.async_configs import SeedingConfig
seed_cfg = SeedingConfig(
source="sitemap+cc",
pattern="*/blog/*",
live_check=True,
extract_head=True,
query="machine learning",
score_threshold=0.2,
max_urls=200,
)
The Internal Seeding Pipeline
When you call await seeder.urls(domain, config), AsyncUrlSeeder executes a fully asynchronous, back-pressured pipeline:
- Initialization – Creates an
httpx.AsyncClientand resolves the latest Common Crawl index ID via_latest_index. - Source Parsing – Splits the
sourcestring (e.g.,"sitemap+cc") and triggers the respective async generators (_from_sitemapsor_from_cc). - Cache Handling – Stores sitemap data in
~/.crawl4ai/seeder_cache/as JSON. The_is_cache_validmethod checks TTL and optional<lastmod>validation to decide between cache reads or refetching. - Filtering – Applies the
_matchfunction usingfnmatchagainst the URL pattern and host part. Iffilter_nonsense_urlsis enabled, internal utility URLs are discarded. - Live-Check and Head Extraction – The
_validatemethod performs HEAD requests (or small GETs for<head>content), while_parse_headextracts metadata. This data attaches to the result dictionary. - Scoring – When a
queryis supplied,_apply_bm25_scoringranks URLs by relevance, filters byscore_threshold, and sorts the final output.
The pipeline uses an asyncio.Queue sized to approximately concurrency * 100 to bound memory usage during high-throughput discovery.
Integrating with AsyncWebCrawler
AsyncWebCrawler (in crawl4ai/async_webcrawler.py) automatically creates an internal AsyncUrlSeeder when you supply a SeedingConfig. This enables focused crawling where only filtered, relevance-ranked URLs enter the crawling queue.
import asyncio
from crawl4ai import AsyncWebCrawler, SeedingConfig
async def main():
seed_cfg = SeedingConfig(
source="sitemap",
pattern="*/docs/*",
live_check=False,
extract_head=True,
query="data science",
score_threshold=0.3,
max_urls=50,
)
crawler = AsyncWebCrawler(seeding_config=seed_cfg)
data = await crawler.arun("https://example.com")
print("Discovered URLs:", [u["url"] for u in data.urls])
await crawler.close()
asyncio.run(main())
Standalone URL Seeding
For scenarios requiring only URL discovery without full page crawling, interact with AsyncUrlSeeder directly:
import asyncio
from crawl4ai.async_url_seeder import AsyncUrlSeeder
from crawl4ai.async_configs import SeedingConfig
async def demo():
seeder = AsyncUrlSeeder()
cfg = SeedingConfig(
source="sitemap",
pattern="*/news/*",
live_check=False,
extract_head=False,
max_urls=20,
)
urls = await seeder.urls("nytimes.com", cfg)
for u in urls:
print(u["url"])
await seeder.close()
asyncio.run(demo())
Optimization Tips for Focused Crawling
| Goal | Configuration | Reason |
|---|---|---|
| Isolate blog posts | pattern="*/blog/*" |
Glob patterns filter paths containing specific directory names. |
| Eliminate dead links | live_check=True |
HEAD requests prune 404/5xx errors before they reach the crawler. |
| Prioritize topical content | extract_head=True, query="deep learning", score_threshold=0.2 |
BM25 scoring uses page titles and metadata to rank relevance. |
| Accelerate large domains | source="cc", force=False |
Common Crawl streaming is faster than sitemap parsing; caching prevents redundant downloads. |
| Control request rates | hits_per_sec=3 |
Prevents accidental aggressive traffic against target servers. |
| Bound result sets | max_urls=100 |
Guarantees finite memory usage and predictable downstream processing. |
Summary
AsyncUrlSeederincrawl4ai/async_url_seeder.pyprovides asynchronous URL discovery from sitemaps and Common Crawl indexes.SeedingConfigcontrols filtering via glob patterns, live link validation, BM25 relevance scoring, and rate limiting.- The internal pipeline uses caching (
~/.crawl4ai/seeder_cache/),_matchfiltering withfnmatch, and optional_apply_bm25_scoringfor semantic ranking. AsyncWebCrawleracceptsSeedingConfigto perform focused crawling, or you can useAsyncUrlSeederstandalone for pure discovery tasks.
Frequently Asked Questions
What is the difference between sitemap and Common Crawl sources?
Sitemap discovery downloads the domain's sitemap.xml to extract explicitly listed URLs, while Common Crawl (CC) streams URLs from the public CC index for historical snapshots of the domain. Sitemaps provide current, officially published URLs, whereas CC offers broader coverage of indexed content that may no longer be in the current sitemap. You can combine both with source="sitemap+cc" for maximum completeness.
How does BM25 relevance scoring work in crawl4ai?
BM25 scoring requires extract_head=True in your SeedingConfig. The seeder fetches each URL's <head> section via _parse_head, then feeds titles and meta descriptions into the _apply_bm25_scoring method. This ranks URLs against your query string; results below score_threshold are filtered out, and the remainder are sorted by relevance before being returned or passed to the crawler.
Can I use URL seeding without crawling the pages immediately?
Yes. Instantiate AsyncUrlSeeder directly and call await seeder.urls(domain, config) to receive a list of filtered, validated URLs without invoking AsyncWebCrawler. This is useful for building URL lists for later batch processing or for analyzing site structure without downloading full page content.
How do I prevent hitting the same sitemap repeatedly?
Enable smart caching by setting force=False (default) and configuring cache_ttl_hours. The seeder stores sitemap data in ~/.crawl4ai/seeder_cache/ and uses _is_cache_valid to check both the TTL and optional <lastmod> timestamps before refetching. This ensures you only download fresh sitemaps when they have actually changed on the server.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →