# How to Extract Tables from HTML with TableExtractionStrategy in crawl4ai

> Learn to extract tables from HTML using crawl4ai's TableExtractionStrategy. Customize parsing logic or use heuristic-based detection for efficient data extraction.

- Repository: [UncleCode/crawl4ai](https://github.com/unclecode/crawl4ai)
- Tags: how-to-guide
- Published: 2026-03-05

---

**crawl4ai extracts tables through a pluggable strategy pattern that lets you configure heuristic-based detection, disable extraction entirely, or implement custom parsing logic by subclassing `TableExtractionStrategy` and passing it to `CrawlerRunConfig`.**

The crawl4ai library provides a robust, strategy-based approach to scrape tabular data from web pages using Python. By leveraging the `TableExtractionStrategy` abstraction defined in the source code, you can extract structured data from HTML tables while maintaining full control over detection sensitivity and parsing behavior.

## Understanding the TableExtractionStrategy Architecture

In [`crawl4ai/table_extraction.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/table_extraction.py), the base `TableExtractionStrategy` class (lines 21-28) defines the contract for all table extractors. This design decouples the crawling engine from specific parsing implementations, allowing you to swap extraction logic without changing your core crawl configuration.

When you invoke `AsyncWebCrawler.arun()` or `WebCrawler.run()`, the crawler passes the parsed DOM element to your configured strategy's `extract_tables` method. The strategy returns a list of normalized dictionaries, which crawl4ai stores in `result.tables`.

## Built-in Extraction Strategies

crawl4ai ships with two production-ready implementations in [`crawl4ai/table_extraction.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/table_extraction.py).

### DefaultTableExtraction

The `DefaultTableExtraction` class provides intelligent, heuristic-based table detection. Found at lines 75-130, this strategy:

- **Scores tables** based on structural quality and content density
- **Handles complex attributes** like `colspan` and `rowspan` during normalization
- **Filters by size** using configurable minimum row and column thresholds
- **Returns standardized dictionaries** containing `headers`, `rows`, `caption`, `summary`, and `metadata` (including `row_count`, `column_count`, and `has_headers`)

The constructor (lines 75-88) accepts parameters like `table_score_threshold`, `min_rows`, and `min_cols`, while the extraction logic resides at lines 90-130.

### NoTableExtraction

When you know a page contains no tables or want to maximize crawl speed, use `NoTableExtraction` (defined at lines 99-107). This no-op implementation of `TableExtractionStrategy` returns an empty list immediately, bypassing all table parsing overhead.

## Configuring Table Extraction

You attach any strategy to the crawler via the `table_extraction` parameter in `CrawlerRunConfig`, defined in [`crawl4ai/types.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/types.py). During execution, the crawler invokes the strategy's `extract_tables` method on the page content and populates `result.tables` with the extracted data.

## Practical Code Examples

### Basic Extraction with Automatic Defaults

If you omit the `table_extraction` parameter, crawl4ai automatically uses `DefaultTableExtraction`. Tune sensitivity with the `table_score_threshold` parameter:

```python
import asyncio
import pandas as pd
from crawl4ai import AsyncWebCrawler, CrawlerRunConfig, CacheMode

async def extract():
    async with AsyncWebCrawler() as crawler:
        cfg = CrawlerRunConfig(
            cache_mode=CacheMode.BYPASS,
            table_score_threshold=7
        )
        result = await crawler.arun(
            "https://en.wikipedia.org/wiki/List_of_countries_by_GDP_(nominal)",
            config=cfg
        )
        if result.success and result.tables:
            print(f"Found {len(result.tables)} tables")
            first = result.tables[0]
            df = pd.DataFrame(first["rows"], columns=first["headers"] or None)
            print(df.head())
            print("Shape:", df.shape)

asyncio.run(extract())

```

*Source: [`docs/examples/table_extraction_example.py`](https://github.com/unclecode/crawl4ai/blob/main/docs/examples/table_extraction_example.py) (lines 21-50)*

### Custom Threshold Configuration

Instantiate `DefaultTableExtraction` directly to enforce stricter criteria:

```python
from crawl4ai import DefaultTableExtraction, CrawlerRunConfig, CacheMode, AsyncWebCrawler

async def custom():
    async with AsyncWebCrawler() as crawler:
        table_strat = DefaultTableExtraction(
            table_score_threshold=5,
            min_rows=3,
            min_cols=2,
            verbose=True
        )
        cfg = CrawlerRunConfig(
            cache_mode=CacheMode.BYPASS,
            table_extraction=table_strat,
            css_selector="div.main-content"
        )
        result = await crawler.arun("https://example.com/data", config=cfg)
        print(f"Extracted {len(result.tables)} tables")

asyncio.run(custom())

```

*Source: [`docs/examples/table_extraction_example.py`](https://github.com/unclecode/crawl4ai/blob/main/docs/examples/table_extraction_example.py) (lines 60-74)*

### Disabling Table Extraction

To skip table processing entirely and improve performance:

```python
from crawl4ai import NoTableExtraction, CrawlerRunConfig, CacheMode, AsyncWebCrawler

async def no_tables():
    async with AsyncWebCrawler() as crawler:
        cfg = CrawlerRunConfig(
            cache_mode=CacheMode.BYPASS,
            table_extraction=NoTableExtraction()
        )
        result = await crawler.arun("https://example.com", config=cfg)
        print("Tables extracted:", len(result.tables))

asyncio.run(no_tables())

```

*Source: [`docs/examples/table_extraction_example.py`](https://github.com/unclecode/crawl4ai/blob/main/docs/examples/table_extraction_example.py) (lines 92-102)*

## Implementing Custom Extraction Strategies

Subclass `TableExtractionStrategy` to implement domain-specific detection rules. The following pattern from [`docs/examples/table_extraction_example.py`](https://github.com/unclecode/crawl4ai/blob/main/docs/examples/table_extraction_example.py) (lines 114-170) demonstrates a financial table extractor that filters tables containing currency symbols:

```python
from crawl4ai import TableExtractionStrategy

class FinancialTableExtraction(TableExtractionStrategy):
    """Extract only tables containing currency indicators."""
    def __init__(self, currency_symbols=None, **kwargs):
        super().__init__(**kwargs)
        self.currency_symbols = currency_symbols or ["$", "€", "£", "¥"]

    def extract_tables(self, element, **kwargs):
        tables = []
        for tbl in element.xpath(".//table"):
            text = "".join(tbl.itertext())
            if any(sym in text for sym in self.currency_symbols):
                # Implement parsing logic or delegate to DefaultTableExtraction

                parsed = self._parse_table(tbl)
                tables.append(parsed)
        return tables

```

Register your custom class with `CrawlerRunConfig(table_extraction=YourStrategy())` exactly like the built-in options.

## Summary

- **`TableExtractionStrategy`** in [`crawl4ai/table_extraction.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/table_extraction.py) (lines 21-28) defines the abstract contract all extractors must implement
- **`DefaultTableExtraction`** (lines 75-130) provides production-ready heuristic parsing with configurable scoring, minimum row/column thresholds, and `colspan`/`rowspan` handling
- **`NoTableExtraction`** (lines 99-107) offers zero-overhead extraction when tables are irrelevant
- **Results** are standardized dictionaries containing `headers`, `rows`, `caption`, `summary`, and rich `metadata` available in `result.tables`
- **Configuration** occurs through `CrawlerRunConfig.table_extraction`, processed during `AsyncWebCrawler.arun()` or `WebCrawler.run()`
- **Extensibility** is native: subclass the base strategy and override `extract_tables` to handle custom DOM structures or domain-specific filtering

## Frequently Asked Questions

### What is the default table extraction behavior in crawl4ai?

If you do not specify a `table_extraction` strategy in `CrawlerRunConfig`, crawl4ai automatically instantiates `DefaultTableExtraction` with sensible defaults. This strategy evaluates all `<table>` elements in the page, assigns quality scores based on structural heuristics, and returns normalized dictionaries for tables exceeding the internal threshold.

### How do I convert extracted tables to a pandas DataFrame?

Each item in `result.tables` is a dictionary containing `headers` (list of strings or `None`) and `rows` (list of lists). Pass these directly to pandas: `pd.DataFrame(table["rows"], columns=table["headers"])`. The `metadata` field includes `row_count` and `column_count` for validation.

### Can I extract tables from only specific parts of a webpage?

Yes. Pass a `css_selector` string to your `CrawlerRunConfig` to restrict processing to a specific DOM subtree. The table extraction strategy operates only within the selected region, improving performance and eliminating irrelevant tables outside your target container.

### Where are the table extraction classes defined?

The strategy hierarchy lives in [`crawl4ai/table_extraction.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/table_extraction.py). The abstract `TableExtractionStrategy` appears at lines 21-28, `DefaultTableExtraction` implementation spans lines 75-130, and `NoTableExtraction` is defined at lines 99-107. These classes are re-exported through [`crawl4ai/__init__.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/__init__.py) for convenient imports.