How to Extract Tables from HTML with TableExtractionStrategy in crawl4ai
crawl4ai extracts tables through a pluggable strategy pattern that lets you configure heuristic-based detection, disable extraction entirely, or implement custom parsing logic by subclassing TableExtractionStrategy and passing it to CrawlerRunConfig.
The crawl4ai library provides a robust, strategy-based approach to scrape tabular data from web pages using Python. By leveraging the TableExtractionStrategy abstraction defined in the source code, you can extract structured data from HTML tables while maintaining full control over detection sensitivity and parsing behavior.
Understanding the TableExtractionStrategy Architecture
In crawl4ai/table_extraction.py, the base TableExtractionStrategy class (lines 21-28) defines the contract for all table extractors. This design decouples the crawling engine from specific parsing implementations, allowing you to swap extraction logic without changing your core crawl configuration.
When you invoke AsyncWebCrawler.arun() or WebCrawler.run(), the crawler passes the parsed DOM element to your configured strategy's extract_tables method. The strategy returns a list of normalized dictionaries, which crawl4ai stores in result.tables.
Built-in Extraction Strategies
crawl4ai ships with two production-ready implementations in crawl4ai/table_extraction.py.
DefaultTableExtraction
The DefaultTableExtraction class provides intelligent, heuristic-based table detection. Found at lines 75-130, this strategy:
- Scores tables based on structural quality and content density
- Handles complex attributes like
colspanandrowspanduring normalization - Filters by size using configurable minimum row and column thresholds
- Returns standardized dictionaries containing
headers,rows,caption,summary, andmetadata(includingrow_count,column_count, andhas_headers)
The constructor (lines 75-88) accepts parameters like table_score_threshold, min_rows, and min_cols, while the extraction logic resides at lines 90-130.
NoTableExtraction
When you know a page contains no tables or want to maximize crawl speed, use NoTableExtraction (defined at lines 99-107). This no-op implementation of TableExtractionStrategy returns an empty list immediately, bypassing all table parsing overhead.
Configuring Table Extraction
You attach any strategy to the crawler via the table_extraction parameter in CrawlerRunConfig, defined in crawl4ai/types.py. During execution, the crawler invokes the strategy's extract_tables method on the page content and populates result.tables with the extracted data.
Practical Code Examples
Basic Extraction with Automatic Defaults
If you omit the table_extraction parameter, crawl4ai automatically uses DefaultTableExtraction. Tune sensitivity with the table_score_threshold parameter:
import asyncio
import pandas as pd
from crawl4ai import AsyncWebCrawler, CrawlerRunConfig, CacheMode
async def extract():
async with AsyncWebCrawler() as crawler:
cfg = CrawlerRunConfig(
cache_mode=CacheMode.BYPASS,
table_score_threshold=7
)
result = await crawler.arun(
"https://en.wikipedia.org/wiki/List_of_countries_by_GDP_(nominal)",
config=cfg
)
if result.success and result.tables:
print(f"Found {len(result.tables)} tables")
first = result.tables[0]
df = pd.DataFrame(first["rows"], columns=first["headers"] or None)
print(df.head())
print("Shape:", df.shape)
asyncio.run(extract())
Source: docs/examples/table_extraction_example.py (lines 21-50)
Custom Threshold Configuration
Instantiate DefaultTableExtraction directly to enforce stricter criteria:
from crawl4ai import DefaultTableExtraction, CrawlerRunConfig, CacheMode, AsyncWebCrawler
async def custom():
async with AsyncWebCrawler() as crawler:
table_strat = DefaultTableExtraction(
table_score_threshold=5,
min_rows=3,
min_cols=2,
verbose=True
)
cfg = CrawlerRunConfig(
cache_mode=CacheMode.BYPASS,
table_extraction=table_strat,
css_selector="div.main-content"
)
result = await crawler.arun("https://example.com/data", config=cfg)
print(f"Extracted {len(result.tables)} tables")
asyncio.run(custom())
Source: docs/examples/table_extraction_example.py (lines 60-74)
Disabling Table Extraction
To skip table processing entirely and improve performance:
from crawl4ai import NoTableExtraction, CrawlerRunConfig, CacheMode, AsyncWebCrawler
async def no_tables():
async with AsyncWebCrawler() as crawler:
cfg = CrawlerRunConfig(
cache_mode=CacheMode.BYPASS,
table_extraction=NoTableExtraction()
)
result = await crawler.arun("https://example.com", config=cfg)
print("Tables extracted:", len(result.tables))
asyncio.run(no_tables())
Source: docs/examples/table_extraction_example.py (lines 92-102)
Implementing Custom Extraction Strategies
Subclass TableExtractionStrategy to implement domain-specific detection rules. The following pattern from docs/examples/table_extraction_example.py (lines 114-170) demonstrates a financial table extractor that filters tables containing currency symbols:
from crawl4ai import TableExtractionStrategy
class FinancialTableExtraction(TableExtractionStrategy):
"""Extract only tables containing currency indicators."""
def __init__(self, currency_symbols=None, **kwargs):
super().__init__(**kwargs)
self.currency_symbols = currency_symbols or ["$", "€", "£", "¥"]
def extract_tables(self, element, **kwargs):
tables = []
for tbl in element.xpath(".//table"):
text = "".join(tbl.itertext())
if any(sym in text for sym in self.currency_symbols):
# Implement parsing logic or delegate to DefaultTableExtraction
parsed = self._parse_table(tbl)
tables.append(parsed)
return tables
Register your custom class with CrawlerRunConfig(table_extraction=YourStrategy()) exactly like the built-in options.
Summary
TableExtractionStrategyincrawl4ai/table_extraction.py(lines 21-28) defines the abstract contract all extractors must implementDefaultTableExtraction(lines 75-130) provides production-ready heuristic parsing with configurable scoring, minimum row/column thresholds, andcolspan/rowspanhandlingNoTableExtraction(lines 99-107) offers zero-overhead extraction when tables are irrelevant- Results are standardized dictionaries containing
headers,rows,caption,summary, and richmetadataavailable inresult.tables - Configuration occurs through
CrawlerRunConfig.table_extraction, processed duringAsyncWebCrawler.arun()orWebCrawler.run() - Extensibility is native: subclass the base strategy and override
extract_tablesto handle custom DOM structures or domain-specific filtering
Frequently Asked Questions
What is the default table extraction behavior in crawl4ai?
If you do not specify a table_extraction strategy in CrawlerRunConfig, crawl4ai automatically instantiates DefaultTableExtraction with sensible defaults. This strategy evaluates all <table> elements in the page, assigns quality scores based on structural heuristics, and returns normalized dictionaries for tables exceeding the internal threshold.
How do I convert extracted tables to a pandas DataFrame?
Each item in result.tables is a dictionary containing headers (list of strings or None) and rows (list of lists). Pass these directly to pandas: pd.DataFrame(table["rows"], columns=table["headers"]). The metadata field includes row_count and column_count for validation.
Can I extract tables from only specific parts of a webpage?
Yes. Pass a css_selector string to your CrawlerRunConfig to restrict processing to a specific DOM subtree. The table extraction strategy operates only within the selected region, improving performance and eliminating irrelevant tables outside your target container.
Where are the table extraction classes defined?
The strategy hierarchy lives in crawl4ai/table_extraction.py. The abstract TableExtractionStrategy appears at lines 21-28, DefaultTableExtraction implementation spans lines 75-130, and NoTableExtraction is defined at lines 99-107. These classes are re-exported through crawl4ai/__init__.py for convenient imports.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →