How to Use LinkPreview for URL Metadata Extraction in crawl4ai

LinkPreview is a built-in crawl4ai feature that fetches only the <head> section of discovered links in parallel, applying configurable include/exclude patterns and optional BM25 relevance scoring against a query.

LinkPreview enables high-performance URL metadata extraction without the overhead of full page loads. As implemented in the unclecode/crawl4ai repository, this feature integrates directly with the AdaptiveCrawler to enrich link discovery with title tags, meta descriptions, and relevance scores.

What Is LinkPreview?

The LinkPreview class in crawl4ai/link_preview.py orchestrates lightweight metadata extraction from hyperlink targets. Unlike full page crawling, LinkPreview uses the AsyncUrlSeeder to perform parallel HTTP HEAD requests (or lightweight GETs limited to the head section), extracting only metadata tags.

Key capabilities include:

  • Parallel extraction using configurable concurrency limits
  • Pattern-based filtering with include/exclude glob patterns
  • BM25 relevance scoring when a query string is provided
  • Score thresholding to filter low-relevance links
  • Integration with AdaptiveCrawler via link_preview_config parameter

Configuration Options

All LinkPreview behavior is controlled through LinkPreviewConfig, defined in crawl4ai/async_configs.py. This dataclass exposes the following parameters:

  • include_internal – Fetch metadata for same-domain links
  • include_external – Fetch metadata for external links
  • include_patterns – List of glob patterns to match (e.g., ["*/blog/*", "*/docs/*"])
  • exclude_patterns – List of glob patterns to reject (e.g., ["*/admin/*", "*/login*"])
  • concurrency – Number of parallel requests (default: 8)
  • timeout – Request timeout in seconds
  • max_links – Maximum links to process per crawl
  • query – Optional BM25 query string for relevance scoring
  • score_threshold – Minimum relevance score to retain (e.g., 0.25)
  • verbose – Enable progress logging

Basic Usage Example

To enable URL metadata extraction, instantiate LinkPreviewConfig and pass it to the AdaptiveCrawler.crawl() method:

import asyncio
from crawl4ai import AdaptiveCrawler, LinkPreviewConfig

async def main():
    # Configure LinkPreview for internal links only

    preview_cfg = LinkPreviewConfig(
        include_internal=True,
        include_external=False,
        concurrency=8,
        timeout=5,
        verbose=False,
    )

    async with AdaptiveCrawler() as crawler:
        result = await crawler.crawl(
            url="https://example.com",
            link_preview_config=preview_cfg,
        )

    # Access metadata from extracted heads

    for link in result.links.internal:
        print(f"URL: {link.href}")
        print(f"Title: {link.head_data.get('title')}")
        print(f"Description: {link.head_data.get('description')}")

asyncio.run(main())

The head_data dictionary contains extracted metadata tags including title, description, keywords, and other meta properties found in the <head> section.

Advanced Usage with Pattern Filtering and BM25 Scoring

For targeted URL metadata extraction, combine glob patterns with BM25 relevance scoring to filter and rank links by semantic relevance:

import asyncio
from crawl4ai import AdaptiveCrawler, LinkPreviewConfig

async def main():
    preview_cfg = LinkPreviewConfig(
        include_internal=True,
        include_external=True,
        include_patterns=["*/blog/*", "*/docs/*", "*/research/*"],
        exclude_patterns=["*/admin/*", "*/login*", "*/logout*"],
        concurrency=12,
        timeout=4,
        max_links=200,
        query="machine learning",
        score_threshold=0.25,
        verbose=True,
    )

    async with AdaptiveCrawler() as crawler:
        result = await crawler.crawl(
            url="https://openai.com",
            link_preview_config=preview_cfg,
        )

    # Links are sorted by total_score (intrinsic + BM25 relevance)

    for link in result.links.internal + result.links.external:
        score = link.head_data.get('relevance_score', 0)
        print(f"Score: {score:.2f} | {link.href}")
        print(f"Title: {link.head_data.get('title')}")

asyncio.run(main())

In crawl4ai/link_preview.py, the _filter_links method applies fnmatch for pattern filtering, while calculate_total_score in crawl4ai/utils.py combines intrinsic link scores with BM25 contextual scores when a query is provided.

Reusing LinkPreview Instances Across Crawls

For batch processing workflows, instantiate a single LinkPreview object and reuse it across multiple crawl operations to avoid recreating the underlying AsyncUrlSeeder connection pool:

import asyncio
from crawl4ai import AdaptiveCrawler, LinkPreview, LinkPreviewConfig

async def main():
    # Single instance for connection pooling

    preview = LinkPreview()
    cfg = LinkPreviewConfig(concurrency=5, verbose=False)

    async with AdaptiveCrawler() as crawler:
        urls = ["https://site1.com", "https://site2.com", "https://site3.com"]
        
        for url in urls:
            result = await crawler.crawl(
                url=url,
                link_preview_config=cfg,
                link_preview=preview,  # Reuse instance

            )
            print(f"{url}: {len(result.links.internal)} links processed")

    # Clean up resources

    await preview.close()

asyncio.run(main())

This pattern improves throughput for large-scale URL metadata extraction by maintaining persistent HTTP connections through the AsyncUrlSeeder utilized by LinkPreview.

Architecture and Source Code Reference

The LinkPreview implementation spans several key files in the unclecode/crawl4ai repository:

  • crawl4ai/link_preview.py – Core LinkPreview class implementing the extraction pipeline, including _filter_links, _extract_heads_parallel, and _merge_head_data methods.
  • crawl4ai/async_configs.py – LinkPreviewConfig dataclass defining all user-configurable parameters.
  • crawl4ai/adaptive_crawler.py – Integration point where AdaptiveCrawler._crawl_with_preview invokes LinkPreview.extract_link_heads when a config is provided.
  • crawl4ai/utils.py – Contains calculate_total_score for combining intrinsic and BM25 relevance scores.
  • crawl4ai/async_url_seeder.py – Underlying AsyncUrlSeeder class performing parallel HTTP requests for head extraction.

Summary

  • LinkPreview enables lightweight URL metadata extraction by fetching only <head> sections via parallel HTTP requests.
  • Configure behavior through LinkPreviewConfig to filter links by patterns, set concurrency limits, and enable BM25 relevance scoring.
  • Access extracted metadata through the head_data attribute on Link objects, which includes title, description, and relevance scores.
  • Reuse LinkPreview instances across multiple crawls to maintain efficient HTTP connection pooling via AsyncUrlSeeder.
  • The feature integrates seamlessly with AdaptiveCrawler through the link_preview_config parameter.

Frequently Asked Questions

How does LinkPreview differ from full page crawling?

LinkPreview extracts only the HTML <head> section using lightweight HTTP requests, while full crawling downloads and processes the entire page content. This makes LinkPreview significantly faster for metadata extraction and link ranking tasks, consuming less bandwidth and processing time when you only need titles, descriptions, and meta tags.

Can I use BM25 scoring without specifying include patterns?

Yes, you can enable BM25 relevance scoring by setting the query parameter in LinkPreviewConfig without defining any include_patterns. When query is provided, the system scores all discovered links against the query text and filters results based on the score_threshold. Include patterns simply provide an additional pre-filtering layer before scoring occurs.

Links that fail to respond or return HTTP errors during head extraction are automatically excluded from the results. The AsyncUrlSeeder handles connection timeouts based on the timeout parameter in your configuration, and failed requests do not raise exceptions but rather result in the link being omitted from the final links collection returned by the crawler.

Is it possible to extract custom meta tags beyond title and description?

Yes, the head_data dictionary captured by LinkPreview contains all meta tags found in the <head> section, not just standard title and description tags. You can access custom meta properties using link.head_data.get('meta') or similar accessors depending on how the specific meta tags are parsed and stored by the AsyncUrlSeeder implementation in crawl4ai/async_url_seeder.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →