# How to Use BM25 Content Filtering to Remove Noise in crawl4ai

> Learn how to use BM25 content filtering in crawl4ai to automatically remove ads and navigation. Preserve relevant content with this powerful tool.

- Repository: [UncleCode/crawl4ai](https://github.com/unclecode/crawl4ai)
- Tags: tutorial
- Published: 2026-03-05

---

**BM25ContentFilter** is the built-in relevance-based content filter that crawl4ai ships with, using the BM25Okapi algorithm and HTML tag weighting to automatically strip navigation, advertisements, and boilerplate while preserving semantically relevant sections.

The crawl4ai library provides a sophisticated **BM25 content filtering** mechanism that eliminates noise from crawled web pages before markdown generation. By leveraging probabilistic ranking and semantic tag prioritization, the `BM25ContentFilter` class ensures only high-relevance HTML sections proceed to the output pipeline.

## How BM25ContentFilter Works

The filter operates through an eight-stage pipeline defined in [`crawl4ai/content_filter_strategy.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/content_filter_strategy.py):

1. **Query Extraction**: When no explicit `user_query` is provided, the filter automatically derives a search query from the page title, `<h1>` element, meta tags, or the first substantial paragraph (lines 25-59).

2. **HTML Chunking**: The `extract_text_chunks` method splits the document body into ordered text segments while preserving tag type metadata (header versus regular content) at lines 61-71.

3. **Tokenization**: Each chunk undergoes tokenization with optional Snowball stemming based on the configured language (lines 85-95).

4. **Noise Removal**: The `clean_tokens` helper imported from [`crawl4ai/utils.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/utils.py) strips stop-words and punctuation from both query and content tokens (lines 104-105).

5. **BM25 Scoring**: The system applies `rank_bm25.BM25Okapi` to calculate relevance scores for each chunk against the query (lines 107-108).

6. **Tag Priority Boosting**: A weighting map elevates scores for semantic HTML elements (`h1`, `h2`, `strong`, etc.), ensuring headings outrank boilerplate text (lines 126-136).

7. **Threshold Filtering**: Chunks scoring below the configurable `bm25_threshold` (default 1.0) are discarded (lines 177-182).

8. **HTML Reconstruction**: The surviving chunks are returned as cleaned HTML fragments in document order, ready for markdown conversion (lines 227-230).

## Source Code Architecture

| File | Purpose |
|------|---------|
| [`crawl4ai/content_filter_strategy.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/content_filter_strategy.py) | Implements `BM25ContentFilter` with query extraction, chunking, BM25 scoring, and tag weighting |
| [`crawl4ai/utils.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/utils.py) | Provides `clean_tokens` for stop-word removal and text normalization |
| [`crawl4ai/markdown_generation_strategy.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/markdown_generation_strategy.py) | Contains `DefaultMarkdownGenerator` which consumes filtered HTML output |
| [`crawl4ai/types.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/types.py) | Defines `CrawlerRunConfig` for injecting filters into the crawling pipeline |

## Direct Usage of BM25ContentFilter

You can instantiate and run the filter independently on raw HTML strings:

```python
from crawl4ai.content_filter_strategy import BM25ContentFilter

# Raw HTML obtained from any source

raw_html = """
<html><body>
  <nav>Navigation bar …</nav>
  <h1>Main article title</h1>
  <p>This paragraph contains the core information you need.</p>
  <div class="ads">Sponsored content</div>
</body></html>
"""

# Initialize with explicit query and threshold

filter = BM25ContentFilter(user_query="core information", bm25_threshold=0.8)

# Returns list of cleaned HTML fragments

clean_chunks = filter.filter_content(raw_html)

for i, chunk in enumerate(clean_chunks, 1):
    print(f"Chunk {i} → {chunk}\n")

```

The `<nav>` and `.ads` elements receive low BM25 scores and are filtered out, while the `<h1>` and relevant `<p>` tags survive as clean HTML strings.

## Integrating BM25 Filtering into the Crawler Pipeline

For production workflows, attach the filter to the markdown generator within the crawler configuration:

```python
import asyncio
from crawl4ai import AsyncWebCrawler, CrawlerRunConfig
from crawl4ai.content_filter_strategy import BM25ContentFilter
from crawl4ai.markdown_generation_strategy import DefaultMarkdownGenerator

async def main():
    # 1. Configure the BM25 filter

    bm25_filter = BM25ContentFilter(
        user_query="Python web scraping tutorial",
        bm25_threshold=1.0,
        language="english",
        use_stemming=True
    )

    # 2. Attach to markdown generator

    md_generator = DefaultMarkdownGenerator(content_filter=bm25_filter)

    # 3. Build run configuration

    run_cfg = CrawlerRunConfig(
        cache_mode="bypass",
        markdown_generator=md_generator
    )

    # 4. Execute crawl

    async with AsyncWebCrawler() as crawler:
        result = await crawler.arun(
            "https://example.com/tutorial.html",
            config=run_cfg
        )
        print(result.markdown)

asyncio.run(main())

```

This pipeline executes the full **crawl → extract → BM25-score → prune → markdown** sequence, delivering noise-free output.

## Tuning Noise Removal Aggressiveness

Control the strictness of content removal by adjusting the `bm25_threshold` parameter:

```python

# Strict filtering: keeps only high-relevance chunks

strict_filter = BM25ContentFilter(bm25_threshold=2.0)

# Permissive filtering: retains most content while dropping obvious boilerplate

lenient_filter = BM25ContentFilter(bm25_threshold=0.5)

```

Higher values increase the influence of the tag priority map, effectively discarding navigation menus, footers, and comment sections that lack semantic relevance to the query.

## Summary

- **BM25ContentFilter** applies the BM25Okapi algorithm to rank HTML chunks by relevance to an extracted or provided query.
- The filter automatically boosts scores for semantic HTML tags (`h1`, `h2`, `strong`) to prioritize content over boilerplate.
- Default `bm25_threshold` is 1.0; increasing this value removes more noise while decreasing it preserves additional content.
- The class integrates directly with `DefaultMarkdownGenerator` and `CrawlerRunConfig` for seamless pipeline insertion.
- Source implementation resides primarily in [`crawl4ai/content_filter_strategy.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/content_filter_strategy.py) with token cleaning utilities in [`crawl4ai/utils.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/utils.py).

## Frequently Asked Questions

### What is BM25ContentFilter in crawl4ai?

**BM25ContentFilter** is a relevance-based content filtering class that uses the BM25 ranking algorithm to identify and preserve the most pertinent HTML sections while removing navigation, advertisements, and other boilerplate content from crawled pages.

### How does BM25ContentFilter determine which content to keep?

The filter tokenizes the query and HTML chunks, calculates BM25 scores using `rank_bm25.BM25Okapi`, applies multipliers based on HTML tag importance (headers and emphasized text receive higher weights), and retains only chunks scoring above the `bm25_threshold` parameter.

### Can I use BM25ContentFilter without the AsyncWebCrawler?

Yes, you can instantiate `BM25ContentFilter` directly and call `filter_content(raw_html)` on any HTML string, making it suitable for processing static HTML files or content retrieved through other HTTP clients.

### How do I adjust the strictness of the BM25 content filtering?

Modify the `bm25_threshold` parameter when initializing the filter. Values above 1.0 produce stricter filtering that removes more potential noise, while values below 1.0 yield more permissive extraction that retains marginal content alongside high-relevance sections.