How to Use BM25 Content Filtering to Remove Noise in crawl4ai

BM25ContentFilter is the built-in relevance-based content filter that crawl4ai ships with, using the BM25Okapi algorithm and HTML tag weighting to automatically strip navigation, advertisements, and boilerplate while preserving semantically relevant sections.

The crawl4ai library provides a sophisticated BM25 content filtering mechanism that eliminates noise from crawled web pages before markdown generation. By leveraging probabilistic ranking and semantic tag prioritization, the BM25ContentFilter class ensures only high-relevance HTML sections proceed to the output pipeline.

How BM25ContentFilter Works

The filter operates through an eight-stage pipeline defined in crawl4ai/content_filter_strategy.py:

  1. Query Extraction: When no explicit user_query is provided, the filter automatically derives a search query from the page title, <h1> element, meta tags, or the first substantial paragraph (lines 25-59).

  2. HTML Chunking: The extract_text_chunks method splits the document body into ordered text segments while preserving tag type metadata (header versus regular content) at lines 61-71.

  3. Tokenization: Each chunk undergoes tokenization with optional Snowball stemming based on the configured language (lines 85-95).

  4. Noise Removal: The clean_tokens helper imported from crawl4ai/utils.py strips stop-words and punctuation from both query and content tokens (lines 104-105).

  5. BM25 Scoring: The system applies rank_bm25.BM25Okapi to calculate relevance scores for each chunk against the query (lines 107-108).

  6. Tag Priority Boosting: A weighting map elevates scores for semantic HTML elements (h1, h2, strong, etc.), ensuring headings outrank boilerplate text (lines 126-136).

  7. Threshold Filtering: Chunks scoring below the configurable bm25_threshold (default 1.0) are discarded (lines 177-182).

  8. HTML Reconstruction: The surviving chunks are returned as cleaned HTML fragments in document order, ready for markdown conversion (lines 227-230).

Source Code Architecture

File Purpose
crawl4ai/content_filter_strategy.py Implements BM25ContentFilter with query extraction, chunking, BM25 scoring, and tag weighting
crawl4ai/utils.py Provides clean_tokens for stop-word removal and text normalization
crawl4ai/markdown_generation_strategy.py Contains DefaultMarkdownGenerator which consumes filtered HTML output
crawl4ai/types.py Defines CrawlerRunConfig for injecting filters into the crawling pipeline

Direct Usage of BM25ContentFilter

You can instantiate and run the filter independently on raw HTML strings:

from crawl4ai.content_filter_strategy import BM25ContentFilter

# Raw HTML obtained from any source

raw_html = """
<html><body>
  <nav>Navigation bar …</nav>
  <h1>Main article title</h1>
  <p>This paragraph contains the core information you need.</p>
  <div class="ads">Sponsored content</div>
</body></html>
"""

# Initialize with explicit query and threshold

filter = BM25ContentFilter(user_query="core information", bm25_threshold=0.8)

# Returns list of cleaned HTML fragments

clean_chunks = filter.filter_content(raw_html)

for i, chunk in enumerate(clean_chunks, 1):
    print(f"Chunk {i} → {chunk}\n")

The <nav> and .ads elements receive low BM25 scores and are filtered out, while the <h1> and relevant <p> tags survive as clean HTML strings.

Integrating BM25 Filtering into the Crawler Pipeline

For production workflows, attach the filter to the markdown generator within the crawler configuration:

import asyncio
from crawl4ai import AsyncWebCrawler, CrawlerRunConfig
from crawl4ai.content_filter_strategy import BM25ContentFilter
from crawl4ai.markdown_generation_strategy import DefaultMarkdownGenerator

async def main():
    # 1. Configure the BM25 filter

    bm25_filter = BM25ContentFilter(
        user_query="Python web scraping tutorial",
        bm25_threshold=1.0,
        language="english",
        use_stemming=True
    )

    # 2. Attach to markdown generator

    md_generator = DefaultMarkdownGenerator(content_filter=bm25_filter)

    # 3. Build run configuration

    run_cfg = CrawlerRunConfig(
        cache_mode="bypass",
        markdown_generator=md_generator
    )

    # 4. Execute crawl

    async with AsyncWebCrawler() as crawler:
        result = await crawler.arun(
            "https://example.com/tutorial.html",
            config=run_cfg
        )
        print(result.markdown)

asyncio.run(main())

This pipeline executes the full crawl → extract → BM25-score → prune → markdown sequence, delivering noise-free output.

Tuning Noise Removal Aggressiveness

Control the strictness of content removal by adjusting the bm25_threshold parameter:


# Strict filtering: keeps only high-relevance chunks

strict_filter = BM25ContentFilter(bm25_threshold=2.0)

# Permissive filtering: retains most content while dropping obvious boilerplate

lenient_filter = BM25ContentFilter(bm25_threshold=0.5)

Higher values increase the influence of the tag priority map, effectively discarding navigation menus, footers, and comment sections that lack semantic relevance to the query.

Summary

  • BM25ContentFilter applies the BM25Okapi algorithm to rank HTML chunks by relevance to an extracted or provided query.
  • The filter automatically boosts scores for semantic HTML tags (h1, h2, strong) to prioritize content over boilerplate.
  • Default bm25_threshold is 1.0; increasing this value removes more noise while decreasing it preserves additional content.
  • The class integrates directly with DefaultMarkdownGenerator and CrawlerRunConfig for seamless pipeline insertion.
  • Source implementation resides primarily in crawl4ai/content_filter_strategy.py with token cleaning utilities in crawl4ai/utils.py.

Frequently Asked Questions

What is BM25ContentFilter in crawl4ai?

BM25ContentFilter is a relevance-based content filtering class that uses the BM25 ranking algorithm to identify and preserve the most pertinent HTML sections while removing navigation, advertisements, and other boilerplate content from crawled pages.

How does BM25ContentFilter determine which content to keep?

The filter tokenizes the query and HTML chunks, calculates BM25 scores using rank_bm25.BM25Okapi, applies multipliers based on HTML tag importance (headers and emphasized text receive higher weights), and retains only chunks scoring above the bm25_threshold parameter.

Can I use BM25ContentFilter without the AsyncWebCrawler?

Yes, you can instantiate BM25ContentFilter directly and call filter_content(raw_html) on any HTML string, making it suitable for processing static HTML files or content retrieved through other HTTP clients.

How do I adjust the strictness of the BM25 content filtering?

Modify the bm25_threshold parameter when initializing the filter. Values above 1.0 produce stricter filtering that removes more potential noise, while values below 1.0 yield more permissive extraction that retains marginal content alongside high-relevance sections.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →