How MinerU Extracts Tables from PDFs: A Deep Dive into the Pipeline

MinerU uses a three-stage pipeline that rasterizes PDF pages, classifies tables as "wired" or "wireless" using a Paddle-based classifier, and routes them to specialized UNet or RapidTable models for HTML extraction and cross-page merging.

MinerU is an open-source PDF parsing toolkit developed by OpenDataLab that converts complex documents into structured markdown and JSON. Understanding how MinerU extracts tables from PDFs reveals why it outperforms generic OCR solutions on scientific papers, financial reports, and scanned documents with complex layouts.

The Three-Stage Table Extraction Pipeline

MinerU processes every PDF through a rigid three-stage architecture designed to handle both digital-born and scanned documents.

Stage 1: PDF to Image Rasterization

The pipeline begins by converting each PDF page into a high-resolution raster image. In minerU/utils/pdf_reader.py, the pdf_to_images function (lines 61-81) uses pypdfium2 to render pages at a configurable DPI (default 200).

The function includes memory safeguards by capping image dimensions via max_width_or_height, ensuring that high-resolution scans do not exhaust GPU memory during downstream processing.

Stage 2: Table Type Classification and Detection

Once rasterized, each page image passes through a table-type classifier to determine the extraction strategy. The PaddleTableClsModel in minerU/model/table/cls/paddle_table_cls.py (lines 15-27) predicts whether a table is "wired" (grid-based with visible borders) or "wireless" (borderless, layout-based).

Based on this classification, MinerU routes the image to one of two specialized models:

  • Wired tables: Processed by UnetTableModel calling WiredTableRecognition in minerU/model/table/rec/unet_table/main.py (lines 257-284). This UNet-based model segments the visible grid structure.
  • Wireless tables: Processed by RapidTableModel calling RapidTable in minerU/model/table/rec/slanet_plus/main.py (lines 36-58). This uses a vision-transformer based Slanet-plus architecture to predict cell positions without visible borders.

Both models return HTML markup, cell bounding boxes, and logical point coordinates.

Stage 3: Post-Processing and Cross-Page Merging

Raw table outputs undergo sophisticated post-processing to handle real-world document complexities. The perform_table_merge function in minerU/utils/table_merge.py executes several critical operations:

  • Column alignment: calculate_table_total_columns standardizes column counts across page breaks.
  • Header detection: detect_table_headers identifies repeated header rows when tables span multiple pages.
  • Structure repair: Adjusts colspan and rowspan attributes for merged cells that cross page boundaries.
  • Caption merging: Combines footnotes and captions when continuation markers like "(续)" appear.

The cross_page_table_merge function in minerU/backend/utils.py (lines 17-23) orchestrates this process by walking the page-wise middle-JSON structure and invoking the merge utilities.

Wired vs. Wireless Table Detection

MinerU's dual-model approach distinguishes it from single-strategy table extractors.

Wired Table Recognition relies on semantic segmentation. The UNet architecture in UnetTableModel identifies horizontal and vertical ruling lines, reconstructs the grid topology, and uses TableRecover to generate HTML. This excels on scanned financial statements and academic tables with clear borders.

Wireless Table Recognition uses the RapidTable (Slanet-plus) model to predict cell positions through pure visual layout analysis. This handles modern PDFs with CSS-styled tables, academic papers using whitespace alignment, and documents where borders are implicit.

Code Example: Running the Extraction Pipeline

The following example demonstrates how to invoke MinerU's table extraction components directly:

from mineru.utils.pdf_reader import pdf_to_images
from mineru.model.table.cls.paddle_table_cls import PaddleTableClsModel
from mineru.model.table.rec.unet_table.main import UnetTableModel
from mineru.model.table.rec.slanet_plus.main import RapidTableModel
from mineru.backend.utils import cross_page_table_merge

# 1. Rasterize PDF

images = pdf_to_images("document.pdf", dpi=200)

# 2. Initialize models

classifier = PaddleTableClsModel()
wired_model = UnetTableModel(ocr_engine=None)
wireless_model = RapidTableModel(ocr_engine=None)

# 3. Process each page

tables_per_page = []
for img in images:
    table_type = classifier.predict(img)
    
    if table_type == "wired":
        result = wired_model.predict(img)
    else:
        result = wireless_model.predict(img)
        
    tables_per_page.append(result)

# 4. Merge cross-page tables

final_tables = cross_page_table_merge(tables_per_page)

The resulting final_tables contains dictionaries with the following structure:

{
  "type": "table",
  "bbox": [100, 200, 500, 600],
  "html": "<table><tr><td>Cell 1</td></tr></table>",
  "cell_bboxes": [[100, 200, 150, 250], ...],
  "logic_points": [...]
}

Summary

  • MinerU extracts tables through a three-stage pipeline: PDF rasterization, type-specific detection (wired vs. wireless), and cross-page merging.
  • The PaddleTableClsModel in paddle_table_cls.py routes tables to either UNet-based wired detection or Slanet-plus-based wireless detection.
  • Post-processing in table_merge.py handles complex real-world scenarios including header repetition, colspan/rowspan repair, and multi-page table continuity.
  • The final output provides structured JSON containing HTML markup, precise cell bounding boxes, and logical coordinates for downstream applications.

Frequently Asked Questions

What is the difference between wired and wireless table detection in MinerU?

Wired table detection processes tables with visible grid lines and borders using a UNet segmentation model (UnetTableModel), while wireless table detection handles borderless tables that rely on whitespace alignment using a vision-transformer based RapidTable model (RapidTableModel). The classifier in paddle_table_cls.py automatically selects the appropriate method based on visual features of each page.

How does MinerU handle tables that span multiple pages?

MinerU uses the cross_page_table_merge function in backend/utils.py to identify table continuations across page breaks. The system detects repeated headers using detect_table_headers, aligns column counts with calculate_table_total_columns, and repairs HTML structure by adjusting colspan and rowspan attributes. This ensures that a single logical table split across physical pages emerges as one coherent data structure in the final JSON output.

Can I adjust the image resolution for table detection?

Yes. The pdf_to_images function in utils/pdf_reader.py accepts a dpi parameter (default 200) that controls the resolution of rasterized pages. Higher DPI values improve detection accuracy for small fonts or complex tables but increase memory consumption and processing time. The function also implements safeguards via max_width_or_height to prevent GPU memory exhaustion when processing high-resolution scans.

What output format does MinerU produce for extracted tables?

MinerU generates a structured JSON object for each table containing: html (the full HTML table markup), bbox (the bounding box coordinates of the entire table), cell_bboxes (an array of bounding boxes for individual cells), and logic_points (logical coordinates for cell relationships). This format preserves both the visual geometry and semantic structure, enabling conversion to Markdown, CSV, or direct database insertion.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →