# How MinerU Extracts Tables from PDFs: A Deep Dive into the Pipeline

> Discover how MinerU extracts tables from PDFs. Explore its three-stage pipeline: rasterization, wired/wireless table classification with PaddlePaddle, and extraction via UNet or RapidTable models.

- Repository: [OpenDataLab/MinerU](https://github.com/opendatalab/mineru)
- Tags: deep-dive
- Published: 2026-02-23

---

**MinerU uses a three-stage pipeline that rasterizes PDF pages, classifies tables as "wired" or "wireless" using a Paddle-based classifier, and routes them to specialized UNet or RapidTable models for HTML extraction and cross-page merging.**

MinerU is an open-source PDF parsing toolkit developed by OpenDataLab that converts complex documents into structured markdown and JSON. Understanding how MinerU extracts tables from PDFs reveals why it outperforms generic OCR solutions on scientific papers, financial reports, and scanned documents with complex layouts.

## The Three-Stage Table Extraction Pipeline

MinerU processes every PDF through a rigid three-stage architecture designed to handle both digital-born and scanned documents.

### Stage 1: PDF to Image Rasterization

The pipeline begins by converting each PDF page into a high-resolution raster image. In [`minerU/utils/pdf_reader.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/utils/pdf_reader.py), the `pdf_to_images` function (lines 61-81) uses **pypdfium2** to render pages at a configurable DPI (default 200).

The function includes memory safeguards by capping image dimensions via `max_width_or_height`, ensuring that high-resolution scans do not exhaust GPU memory during downstream processing.

### Stage 2: Table Type Classification and Detection

Once rasterized, each page image passes through a **table-type classifier** to determine the extraction strategy. The `PaddleTableClsModel` in [`minerU/model/table/cls/paddle_table_cls.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/model/table/cls/paddle_table_cls.py) (lines 15-27) predicts whether a table is **"wired"** (grid-based with visible borders) or **"wireless"** (borderless, layout-based).

Based on this classification, MinerU routes the image to one of two specialized models:

- **Wired tables**: Processed by `UnetTableModel` calling `WiredTableRecognition` in [`minerU/model/table/rec/unet_table/main.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/model/table/rec/unet_table/main.py) (lines 257-284). This UNet-based model segments the visible grid structure.
- **Wireless tables**: Processed by `RapidTableModel` calling `RapidTable` in [`minerU/model/table/rec/slanet_plus/main.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/model/table/rec/slanet_plus/main.py) (lines 36-58). This uses a vision-transformer based Slanet-plus architecture to predict cell positions without visible borders.

Both models return HTML markup, cell bounding boxes, and logical point coordinates.

### Stage 3: Post-Processing and Cross-Page Merging

Raw table outputs undergo sophisticated post-processing to handle real-world document complexities. The `perform_table_merge` function in [`minerU/utils/table_merge.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/utils/table_merge.py) executes several critical operations:

- **Column alignment**: `calculate_table_total_columns` standardizes column counts across page breaks.
- **Header detection**: `detect_table_headers` identifies repeated header rows when tables span multiple pages.
- **Structure repair**: Adjusts `colspan` and `rowspan` attributes for merged cells that cross page boundaries.
- **Caption merging**: Combines footnotes and captions when continuation markers like "(续)" appear.

The `cross_page_table_merge` function in [`minerU/backend/utils.py`](https://github.com/opendatalab/MinerU/blob/main/minerU/backend/utils.py) (lines 17-23) orchestrates this process by walking the page-wise middle-JSON structure and invoking the merge utilities.

## Wired vs. Wireless Table Detection

MinerU's dual-model approach distinguishes it from single-strategy table extractors.

**Wired Table Recognition** relies on semantic segmentation. The UNet architecture in `UnetTableModel` identifies horizontal and vertical ruling lines, reconstructs the grid topology, and uses `TableRecover` to generate HTML. This excels on scanned financial statements and academic tables with clear borders.

**Wireless Table Recognition** uses the `RapidTable` (Slanet-plus) model to predict cell positions through pure visual layout analysis. This handles modern PDFs with CSS-styled tables, academic papers using whitespace alignment, and documents where borders are implicit.

## Code Example: Running the Extraction Pipeline

The following example demonstrates how to invoke MinerU's table extraction components directly:

```python
from mineru.utils.pdf_reader import pdf_to_images
from mineru.model.table.cls.paddle_table_cls import PaddleTableClsModel
from mineru.model.table.rec.unet_table.main import UnetTableModel
from mineru.model.table.rec.slanet_plus.main import RapidTableModel
from mineru.backend.utils import cross_page_table_merge

# 1. Rasterize PDF

images = pdf_to_images("document.pdf", dpi=200)

# 2. Initialize models

classifier = PaddleTableClsModel()
wired_model = UnetTableModel(ocr_engine=None)
wireless_model = RapidTableModel(ocr_engine=None)

# 3. Process each page

tables_per_page = []
for img in images:
    table_type = classifier.predict(img)
    
    if table_type == "wired":
        result = wired_model.predict(img)
    else:
        result = wireless_model.predict(img)
        
    tables_per_page.append(result)

# 4. Merge cross-page tables

final_tables = cross_page_table_merge(tables_per_page)

```

The resulting `final_tables` contains dictionaries with the following structure:

```json
{
  "type": "table",
  "bbox": [100, 200, 500, 600],
  "html": "<table><tr><td>Cell 1</td></tr></table>",
  "cell_bboxes": [[100, 200, 150, 250], ...],
  "logic_points": [...]
}

```

## Summary

- MinerU extracts tables through a **three-stage pipeline**: PDF rasterization, type-specific detection (wired vs. wireless), and cross-page merging.
- The **PaddleTableClsModel** in [`paddle_table_cls.py`](https://github.com/opendatalab/MinerU/blob/main/paddle_table_cls.py) routes tables to either **UNet-based** wired detection or **Slanet-plus-based** wireless detection.
- **Post-processing** in [`table_merge.py`](https://github.com/opendatalab/MinerU/blob/main/table_merge.py) handles complex real-world scenarios including header repetition, colspan/rowspan repair, and multi-page table continuity.
- The final output provides **structured JSON** containing HTML markup, precise cell bounding boxes, and logical coordinates for downstream applications.

## Frequently Asked Questions

### What is the difference between wired and wireless table detection in MinerU?

**Wired table detection** processes tables with visible grid lines and borders using a UNet segmentation model (`UnetTableModel`), while **wireless table detection** handles borderless tables that rely on whitespace alignment using a vision-transformer based RapidTable model (`RapidTableModel`). The classifier in [`paddle_table_cls.py`](https://github.com/opendatalab/MinerU/blob/main/paddle_table_cls.py) automatically selects the appropriate method based on visual features of each page.

### How does MinerU handle tables that span multiple pages?

MinerU uses the `cross_page_table_merge` function in [`backend/utils.py`](https://github.com/opendatalab/MinerU/blob/main/backend/utils.py) to identify table continuations across page breaks. The system detects repeated headers using `detect_table_headers`, aligns column counts with `calculate_table_total_columns`, and repairs HTML structure by adjusting `colspan` and `rowspan` attributes. This ensures that a single logical table split across physical pages emerges as one coherent data structure in the final JSON output.

### Can I adjust the image resolution for table detection?

Yes. The `pdf_to_images` function in [`utils/pdf_reader.py`](https://github.com/opendatalab/MinerU/blob/main/utils/pdf_reader.py) accepts a `dpi` parameter (default 200) that controls the resolution of rasterized pages. Higher DPI values improve detection accuracy for small fonts or complex tables but increase memory consumption and processing time. The function also implements safeguards via `max_width_or_height` to prevent GPU memory exhaustion when processing high-resolution scans.

### What output format does MinerU produce for extracted tables?

MinerU generates a structured JSON object for each table containing: `html` (the full HTML table markup), `bbox` (the bounding box coordinates of the entire table), `cell_bboxes` (an array of bounding boxes for individual cells), and `logic_points` (logical coordinates for cell relationships). This format preserves both the visual geometry and semantic structure, enabling conversion to Markdown, CSV, or direct database insertion.