Table Detection Methods in pdf-inspector: Rectangle, Line, and Heuristic Approaches

pdf-inspector extracts tables from PDFs using a three-stage pipeline that applies rectangle-based, line-based, and heuristic detection methods in sequence, stopping at the first successful match.

The pdf-inspector library by Firecrawl provides robust table extraction capabilities for PDF documents through a tiered detection strategy. Understanding the different methods for table detection in pdf-inspector helps developers optimize extraction accuracy for various PDF formats, from structured reports to scanned documents.

The Three-Stage Detection Pipeline

pdf-inspector implements three distinct detection algorithms in src/tables/mod.rs, each targeting specific PDF encoding patterns. The system prioritizes graphics-based approaches over text heuristics to maximize accuracy.

Rectangle-Based Detection

The rectangle-based method, implemented in src/tables/detect_rects.rs, analyzes PDF drawing operations to identify table structures. This method collects filled boxes, cell backgrounds, and rule lines, then applies a union-find algorithm to cluster rectangles into coherent table cells. It excels at processing PDFs that encode table structure as explicit geometric shapes.

Line-Based Detection

When rectangle detection fails, the pipeline falls back to src/tables/detect_lines.rs. This scanner identifies horizontal and vertical line operators within the page content, constructing a grid from intersecting lines to infer cell boundaries. The line-based approach works effectively for PDFs that draw tables with straight lines but lack filled rectangular backgrounds.

Heuristic Detection

The final fallback resides in src/tables/detect_heuristic.rs, which analyzes textual patterns when no explicit graphics exist. This method examines text flow, font size variations, and inter-item gaps using gap-histogram analysis. If necessary, it employs body-font clustering to group text into logical table structures.

How the Detection Pipeline Works

The orchestration logic in src/tables/mod.rs exports three primary functions: detect_tables_from_rects, detect_tables_from_lines, and detect_tables_heuristically. During processing, process_pdf_with_options in src/lib.rs invokes these detectors in strict priority order.

The pipeline executes sequentially:

  1. detect_tables_from_rects runs first with configurable tolerance parameters
  2. If empty results return, detect_tables_from_lines executes
  3. Finally, detect_tables_heuristically processes the page if previous methods fail

This tiered approach ensures that the most reliable graphics-based methods are preferred while maintaining coverage for text-heavy documents.

Using Table Detection

Developers can access table detection through the command-line interface or direct library integration.

CLI Usage

The pdf2md command automatically applies all three detection methods during conversion:


# Extract PDF to Markdown with automatic table detection

pdf2md my-report.pdf > my-report.md

# Output structured JSON including detected table data

pdf2md --json my-report.pdf > my-report.json

Programmatic Usage

For direct library access, import the detection modules from pdf_inspector::tables:

use pdf_inspector::lib::process_pdf_with_options;
use pdf_inspector::tables::{detect_rects, detect_lines, detect_heuristic};

fn main() -> anyhow::Result<()> {
    // Load PDF document
    let pdf = pdf_inspector::pdf::PdfDocument::load("my-report.pdf")?;
    
    // Run full extraction pipeline
    let markdown = process_pdf_with_options(&pdf, Default::default())?;
    
    // Direct detector access for custom logic
    let page_rects = pdf.page_rects(0)?;
    let rect_tables = detect_rects::detect_tables_from_rects(&page_rects, 3.0, 6)?;
    
    if rect_tables.is_empty() {
        let line_tables = detect_lines::detect_tables_from_lines(&page_rects)?;
        if line_tables.is_empty() {
            let heuristic_tables = detect_heuristic::detect_tables_heuristically(&page_rects)?;
        }
    }
    Ok(())
}

Summary

  • pdf-inspector employs three distinct table detection methods: rectangle-based, line-based, and heuristic
  • Detection occurs in fixed priority order, preferring explicit graphics over text analysis
  • Rectangle detection uses union-find clustering in src/tables/detect_rects.rs
  • Line detection builds grids from intersecting operators in src/tables/detect_lines.rs
  • Heuristic detection analyzes text gaps and fonts in src/tables/detect_heuristic.rs
  • The pipeline automatically selects the most reliable method available for each PDF page

Frequently Asked Questions

What order does pdf-inspector use for table detection?

The pipeline follows a strict sequence: rectangle-based detection runs first, followed by line-based detection if no tables are found, and finally heuristic detection as a fallback. This ordering prioritizes methods that rely on explicit PDF graphics operators over text inference.

When should I use the heuristic detector directly?

Access detect_tables_heuristically directly when processing PDFs known to contain implicit tables without grid lines or background rectangles, such as whitespace-aligned text columns or monospace font layouts. This method analyzes gap histograms and font clustering to identify tabular structures invisible to graphic-based detectors.

How does rectangle-based detection work internally?

The rectangle detector in src/tables/detect_rects.rs extracts drawing rectangles from PDF content streams, including filled boxes and cell backgrounds. It implements a union-find algorithm to cluster adjacent rectangles into coherent grids, effectively reconstructing table cells from explicit geometric primitives embedded in the document.

Can I disable specific detection methods?

While the default process_pdf_with_options function in src/lib.rs runs the full pipeline automatically, you can bypass unwanted methods by calling specific detector functions directly from src/tables/mod.rs. Import individual detection modules and implement custom logic that skips rectangle or line detection based on your document characteristics.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →