How to Use BM25 Content Filtering to Remove Noise in crawl4ai
BM25ContentFilter is the built-in relevance-based content filter that crawl4ai ships with, using the BM25Okapi algorithm and HTML tag weighting to automatically strip navigation, advertisements, and boilerplate while preserving semantically relevant sections.
The crawl4ai library provides a sophisticated BM25 content filtering mechanism that eliminates noise from crawled web pages before markdown generation. By leveraging probabilistic ranking and semantic tag prioritization, the BM25ContentFilter class ensures only high-relevance HTML sections proceed to the output pipeline.
How BM25ContentFilter Works
The filter operates through an eight-stage pipeline defined in crawl4ai/content_filter_strategy.py:
-
Query Extraction: When no explicit
user_queryis provided, the filter automatically derives a search query from the page title,<h1>element, meta tags, or the first substantial paragraph (lines 25-59). -
HTML Chunking: The
extract_text_chunksmethod splits the document body into ordered text segments while preserving tag type metadata (header versus regular content) at lines 61-71. -
Tokenization: Each chunk undergoes tokenization with optional Snowball stemming based on the configured language (lines 85-95).
-
Noise Removal: The
clean_tokenshelper imported fromcrawl4ai/utils.pystrips stop-words and punctuation from both query and content tokens (lines 104-105). -
BM25 Scoring: The system applies
rank_bm25.BM25Okapito calculate relevance scores for each chunk against the query (lines 107-108). -
Tag Priority Boosting: A weighting map elevates scores for semantic HTML elements (
h1,h2,strong, etc.), ensuring headings outrank boilerplate text (lines 126-136). -
Threshold Filtering: Chunks scoring below the configurable
bm25_threshold(default 1.0) are discarded (lines 177-182). -
HTML Reconstruction: The surviving chunks are returned as cleaned HTML fragments in document order, ready for markdown conversion (lines 227-230).
Source Code Architecture
| File | Purpose |
|---|---|
crawl4ai/content_filter_strategy.py |
Implements BM25ContentFilter with query extraction, chunking, BM25 scoring, and tag weighting |
crawl4ai/utils.py |
Provides clean_tokens for stop-word removal and text normalization |
crawl4ai/markdown_generation_strategy.py |
Contains DefaultMarkdownGenerator which consumes filtered HTML output |
crawl4ai/types.py |
Defines CrawlerRunConfig for injecting filters into the crawling pipeline |
Direct Usage of BM25ContentFilter
You can instantiate and run the filter independently on raw HTML strings:
from crawl4ai.content_filter_strategy import BM25ContentFilter
# Raw HTML obtained from any source
raw_html = """
<html><body>
<nav>Navigation bar …</nav>
<h1>Main article title</h1>
<p>This paragraph contains the core information you need.</p>
<div class="ads">Sponsored content</div>
</body></html>
"""
# Initialize with explicit query and threshold
filter = BM25ContentFilter(user_query="core information", bm25_threshold=0.8)
# Returns list of cleaned HTML fragments
clean_chunks = filter.filter_content(raw_html)
for i, chunk in enumerate(clean_chunks, 1):
print(f"Chunk {i} → {chunk}\n")
The <nav> and .ads elements receive low BM25 scores and are filtered out, while the <h1> and relevant <p> tags survive as clean HTML strings.
Integrating BM25 Filtering into the Crawler Pipeline
For production workflows, attach the filter to the markdown generator within the crawler configuration:
import asyncio
from crawl4ai import AsyncWebCrawler, CrawlerRunConfig
from crawl4ai.content_filter_strategy import BM25ContentFilter
from crawl4ai.markdown_generation_strategy import DefaultMarkdownGenerator
async def main():
# 1. Configure the BM25 filter
bm25_filter = BM25ContentFilter(
user_query="Python web scraping tutorial",
bm25_threshold=1.0,
language="english",
use_stemming=True
)
# 2. Attach to markdown generator
md_generator = DefaultMarkdownGenerator(content_filter=bm25_filter)
# 3. Build run configuration
run_cfg = CrawlerRunConfig(
cache_mode="bypass",
markdown_generator=md_generator
)
# 4. Execute crawl
async with AsyncWebCrawler() as crawler:
result = await crawler.arun(
"https://example.com/tutorial.html",
config=run_cfg
)
print(result.markdown)
asyncio.run(main())
This pipeline executes the full crawl → extract → BM25-score → prune → markdown sequence, delivering noise-free output.
Tuning Noise Removal Aggressiveness
Control the strictness of content removal by adjusting the bm25_threshold parameter:
# Strict filtering: keeps only high-relevance chunks
strict_filter = BM25ContentFilter(bm25_threshold=2.0)
# Permissive filtering: retains most content while dropping obvious boilerplate
lenient_filter = BM25ContentFilter(bm25_threshold=0.5)
Higher values increase the influence of the tag priority map, effectively discarding navigation menus, footers, and comment sections that lack semantic relevance to the query.
Summary
- BM25ContentFilter applies the BM25Okapi algorithm to rank HTML chunks by relevance to an extracted or provided query.
- The filter automatically boosts scores for semantic HTML tags (
h1,h2,strong) to prioritize content over boilerplate. - Default
bm25_thresholdis 1.0; increasing this value removes more noise while decreasing it preserves additional content. - The class integrates directly with
DefaultMarkdownGeneratorandCrawlerRunConfigfor seamless pipeline insertion. - Source implementation resides primarily in
crawl4ai/content_filter_strategy.pywith token cleaning utilities incrawl4ai/utils.py.
Frequently Asked Questions
What is BM25ContentFilter in crawl4ai?
BM25ContentFilter is a relevance-based content filtering class that uses the BM25 ranking algorithm to identify and preserve the most pertinent HTML sections while removing navigation, advertisements, and other boilerplate content from crawled pages.
How does BM25ContentFilter determine which content to keep?
The filter tokenizes the query and HTML chunks, calculates BM25 scores using rank_bm25.BM25Okapi, applies multipliers based on HTML tag importance (headers and emphasized text receive higher weights), and retains only chunks scoring above the bm25_threshold parameter.
Can I use BM25ContentFilter without the AsyncWebCrawler?
Yes, you can instantiate BM25ContentFilter directly and call filter_content(raw_html) on any HTML string, making it suitable for processing static HTML files or content retrieved through other HTTP clients.
How do I adjust the strictness of the BM25 content filtering?
Modify the bm25_threshold parameter when initializing the filter. Values above 1.0 produce stricter filtering that removes more potential noise, while values below 1.0 yield more permissive extraction that retains marginal content alongside high-relevance sections.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →