# How to Use LinkPreview for URL Metadata Extraction in crawl4ai

> Learn how to use LinkPreview for URL metadata extraction in crawl4ai. Fetch link head sections in parallel with configurable patterns and BM25 scoring.

- Repository: [UncleCode/crawl4ai](https://github.com/unclecode/crawl4ai)
- Tags: how-to-guide
- Published: 2026-03-05

---

**LinkPreview is a built-in crawl4ai feature that fetches only the `<head>` section of discovered links in parallel, applying configurable include/exclude patterns and optional BM25 relevance scoring against a query.**

LinkPreview enables high-performance URL metadata extraction without the overhead of full page loads. As implemented in the `unclecode/crawl4ai` repository, this feature integrates directly with the `AdaptiveCrawler` to enrich link discovery with title tags, meta descriptions, and relevance scores.

## What Is LinkPreview?

The `LinkPreview` class in [`crawl4ai/link_preview.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/link_preview.py) orchestrates lightweight metadata extraction from hyperlink targets. Unlike full page crawling, LinkPreview uses the `AsyncUrlSeeder` to perform parallel HTTP HEAD requests (or lightweight GETs limited to the head section), extracting only metadata tags.

Key capabilities include:

- **Parallel extraction** using configurable concurrency limits
- **Pattern-based filtering** with include/exclude glob patterns
- **BM25 relevance scoring** when a query string is provided
- **Score thresholding** to filter low-relevance links
- **Integration** with `AdaptiveCrawler` via `link_preview_config` parameter

## Configuration Options

All LinkPreview behavior is controlled through `LinkPreviewConfig`, defined in [`crawl4ai/async_configs.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/async_configs.py). This dataclass exposes the following parameters:

- `include_internal` – Fetch metadata for same-domain links
- `include_external` – Fetch metadata for external links
- `include_patterns` – List of glob patterns to match (e.g., `["*/blog/*", "*/docs/*"]`)
- `exclude_patterns` – List of glob patterns to reject (e.g., `["*/admin/*", "*/login*"]`)
- `concurrency` – Number of parallel requests (default: 8)
- `timeout` – Request timeout in seconds
- `max_links` – Maximum links to process per crawl
- `query` – Optional BM25 query string for relevance scoring
- `score_threshold` – Minimum relevance score to retain (e.g., 0.25)
- `verbose` – Enable progress logging

## Basic Usage Example

To enable URL metadata extraction, instantiate `LinkPreviewConfig` and pass it to the `AdaptiveCrawler.crawl()` method:

```python
import asyncio
from crawl4ai import AdaptiveCrawler, LinkPreviewConfig

async def main():
    # Configure LinkPreview for internal links only

    preview_cfg = LinkPreviewConfig(
        include_internal=True,
        include_external=False,
        concurrency=8,
        timeout=5,
        verbose=False,
    )

    async with AdaptiveCrawler() as crawler:
        result = await crawler.crawl(
            url="https://example.com",
            link_preview_config=preview_cfg,
        )

    # Access metadata from extracted heads

    for link in result.links.internal:
        print(f"URL: {link.href}")
        print(f"Title: {link.head_data.get('title')}")
        print(f"Description: {link.head_data.get('description')}")

asyncio.run(main())

```

The `head_data` dictionary contains extracted metadata tags including `title`, `description`, `keywords`, and other meta properties found in the `<head>` section.

## Advanced Usage with Pattern Filtering and BM25 Scoring

For targeted URL metadata extraction, combine glob patterns with BM25 relevance scoring to filter and rank links by semantic relevance:

```python
import asyncio
from crawl4ai import AdaptiveCrawler, LinkPreviewConfig

async def main():
    preview_cfg = LinkPreviewConfig(
        include_internal=True,
        include_external=True,
        include_patterns=["*/blog/*", "*/docs/*", "*/research/*"],
        exclude_patterns=["*/admin/*", "*/login*", "*/logout*"],
        concurrency=12,
        timeout=4,
        max_links=200,
        query="machine learning",
        score_threshold=0.25,
        verbose=True,
    )

    async with AdaptiveCrawler() as crawler:
        result = await crawler.crawl(
            url="https://openai.com",
            link_preview_config=preview_cfg,
        )

    # Links are sorted by total_score (intrinsic + BM25 relevance)

    for link in result.links.internal + result.links.external:
        score = link.head_data.get('relevance_score', 0)
        print(f"Score: {score:.2f} | {link.href}")
        print(f"Title: {link.head_data.get('title')}")

asyncio.run(main())

```

In [`crawl4ai/link_preview.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/link_preview.py), the `_filter_links` method applies `fnmatch` for pattern filtering, while `calculate_total_score` in [`crawl4ai/utils.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/utils.py) combines intrinsic link scores with BM25 contextual scores when a query is provided.

## Reusing LinkPreview Instances Across Crawls

For batch processing workflows, instantiate a single `LinkPreview` object and reuse it across multiple crawl operations to avoid recreating the underlying `AsyncUrlSeeder` connection pool:

```python
import asyncio
from crawl4ai import AdaptiveCrawler, LinkPreview, LinkPreviewConfig

async def main():
    # Single instance for connection pooling

    preview = LinkPreview()
    cfg = LinkPreviewConfig(concurrency=5, verbose=False)

    async with AdaptiveCrawler() as crawler:
        urls = ["https://site1.com", "https://site2.com", "https://site3.com"]
        
        for url in urls:
            result = await crawler.crawl(
                url=url,
                link_preview_config=cfg,
                link_preview=preview,  # Reuse instance

            )
            print(f"{url}: {len(result.links.internal)} links processed")

    # Clean up resources

    await preview.close()

asyncio.run(main())

```

This pattern improves throughput for large-scale URL metadata extraction by maintaining persistent HTTP connections through the `AsyncUrlSeeder` utilized by `LinkPreview`.

## Architecture and Source Code Reference

The LinkPreview implementation spans several key files in the `unclecode/crawl4ai` repository:

- **[`crawl4ai/link_preview.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/link_preview.py)** – Core `LinkPreview` class implementing the extraction pipeline, including `_filter_links`, `_extract_heads_parallel`, and `_merge_head_data` methods.
- **[`crawl4ai/async_configs.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/async_configs.py)** – `LinkPreviewConfig` dataclass defining all user-configurable parameters.
- **[`crawl4ai/adaptive_crawler.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/adaptive_crawler.py)** – Integration point where `AdaptiveCrawler._crawl_with_preview` invokes `LinkPreview.extract_link_heads` when a config is provided.
- **[`crawl4ai/utils.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/utils.py)** – Contains `calculate_total_score` for combining intrinsic and BM25 relevance scores.
- **[`crawl4ai/async_url_seeder.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/async_url_seeder.py)** – Underlying `AsyncUrlSeeder` class performing parallel HTTP requests for head extraction.

## Summary

- **LinkPreview** enables lightweight URL metadata extraction by fetching only `<head>` sections via parallel HTTP requests.
- Configure behavior through **`LinkPreviewConfig`** to filter links by patterns, set concurrency limits, and enable BM25 relevance scoring.
- Access extracted metadata through the **`head_data`** attribute on `Link` objects, which includes `title`, `description`, and relevance scores.
- Reuse **`LinkPreview`** instances across multiple crawls to maintain efficient HTTP connection pooling via `AsyncUrlSeeder`.
- The feature integrates seamlessly with **`AdaptiveCrawler`** through the `link_preview_config` parameter.

## Frequently Asked Questions

### How does LinkPreview differ from full page crawling?

LinkPreview extracts only the HTML `<head>` section using lightweight HTTP requests, while full crawling downloads and processes the entire page content. This makes LinkPreview significantly faster for metadata extraction and link ranking tasks, consuming less bandwidth and processing time when you only need titles, descriptions, and meta tags.

### Can I use BM25 scoring without specifying include patterns?

Yes, you can enable BM25 relevance scoring by setting the `query` parameter in `LinkPreviewConfig` without defining any `include_patterns`. When `query` is provided, the system scores all discovered links against the query text and filters results based on the `score_threshold`. Include patterns simply provide an additional pre-filtering layer before scoring occurs.

### What happens if a link returns a 404 or timeout during head extraction?

Links that fail to respond or return HTTP errors during head extraction are automatically excluded from the results. The `AsyncUrlSeeder` handles connection timeouts based on the `timeout` parameter in your configuration, and failed requests do not raise exceptions but rather result in the link being omitted from the final `links` collection returned by the crawler.

### Is it possible to extract custom meta tags beyond title and description?

Yes, the `head_data` dictionary captured by LinkPreview contains all meta tags found in the `<head>` section, not just standard title and description tags. You can access custom meta properties using `link.head_data.get('meta')` or similar accessors depending on how the specific meta tags are parsed and stored by the `AsyncUrlSeeder` implementation in [`crawl4ai/async_url_seeder.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/async_url_seeder.py).