# Table Detection Methods in pdf-inspector: Rectangle, Line, and Heuristic Approaches

> Explore pdf-inspector's table detection methods: rectangle, line, and heuristic. Discover how this tool efficiently extracts tables from PDFs.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: deep-dive
- Published: 2026-08-08

---

**pdf-inspector extracts tables from PDFs using a three-stage pipeline that applies rectangle-based, line-based, and heuristic detection methods in sequence, stopping at the first successful match.**

The pdf-inspector library by Firecrawl provides robust table extraction capabilities for PDF documents through a tiered detection strategy. Understanding the different methods for table detection in pdf-inspector helps developers optimize extraction accuracy for various PDF formats, from structured reports to scanned documents.

## The Three-Stage Detection Pipeline

pdf-inspector implements three distinct detection algorithms in [`src/tables/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/mod.rs), each targeting specific PDF encoding patterns. The system prioritizes graphics-based approaches over text heuristics to maximize accuracy.

### Rectangle-Based Detection

The rectangle-based method, implemented in [`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs), analyzes PDF drawing operations to identify table structures. This method collects filled boxes, cell backgrounds, and rule lines, then applies a **union-find algorithm** to cluster rectangles into coherent table cells. It excels at processing PDFs that encode table structure as explicit geometric shapes.

### Line-Based Detection

When rectangle detection fails, the pipeline falls back to [`src/tables/detect_lines.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_lines.rs). This scanner identifies horizontal and vertical line operators within the page content, constructing a grid from intersecting lines to infer cell boundaries. The line-based approach works effectively for PDFs that draw tables with straight lines but lack filled rectangular backgrounds.

### Heuristic Detection

The final fallback resides in [`src/tables/detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs), which analyzes textual patterns when no explicit graphics exist. This method examines text flow, font size variations, and inter-item gaps using gap-histogram analysis. If necessary, it employs body-font clustering to group text into logical table structures.

## How the Detection Pipeline Works

The orchestration logic in [`src/tables/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/mod.rs) exports three primary functions: `detect_tables_from_rects`, `detect_tables_from_lines`, and `detect_tables_heuristically`. During processing, `process_pdf_with_options` in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) invokes these detectors in strict priority order.

The pipeline executes sequentially:

1. `detect_tables_from_rects` runs first with configurable tolerance parameters
2. If empty results return, `detect_tables_from_lines` executes
3. Finally, `detect_tables_heuristically` processes the page if previous methods fail

This tiered approach ensures that the most reliable graphics-based methods are preferred while maintaining coverage for text-heavy documents.

## Using Table Detection

Developers can access table detection through the command-line interface or direct library integration.

### CLI Usage

The `pdf2md` command automatically applies all three detection methods during conversion:

```bash

# Extract PDF to Markdown with automatic table detection

pdf2md my-report.pdf > my-report.md

# Output structured JSON including detected table data

pdf2md --json my-report.pdf > my-report.json

```

### Programmatic Usage

For direct library access, import the detection modules from `pdf_inspector::tables`:

```rust
use pdf_inspector::lib::process_pdf_with_options;
use pdf_inspector::tables::{detect_rects, detect_lines, detect_heuristic};

fn main() -> anyhow::Result<()> {
    // Load PDF document
    let pdf = pdf_inspector::pdf::PdfDocument::load("my-report.pdf")?;
    
    // Run full extraction pipeline
    let markdown = process_pdf_with_options(&pdf, Default::default())?;
    
    // Direct detector access for custom logic
    let page_rects = pdf.page_rects(0)?;
    let rect_tables = detect_rects::detect_tables_from_rects(&page_rects, 3.0, 6)?;
    
    if rect_tables.is_empty() {
        let line_tables = detect_lines::detect_tables_from_lines(&page_rects)?;
        if line_tables.is_empty() {
            let heuristic_tables = detect_heuristic::detect_tables_heuristically(&page_rects)?;
        }
    }
    Ok(())
}

```

## Summary

- pdf-inspector employs three distinct table detection methods: rectangle-based, line-based, and heuristic
- Detection occurs in fixed priority order, preferring explicit graphics over text analysis
- Rectangle detection uses union-find clustering in [`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs)
- Line detection builds grids from intersecting operators in [`src/tables/detect_lines.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_lines.rs)
- Heuristic detection analyzes text gaps and fonts in [`src/tables/detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs)
- The pipeline automatically selects the most reliable method available for each PDF page

## Frequently Asked Questions

### What order does pdf-inspector use for table detection?

The pipeline follows a strict sequence: rectangle-based detection runs first, followed by line-based detection if no tables are found, and finally heuristic detection as a fallback. This ordering prioritizes methods that rely on explicit PDF graphics operators over text inference.

### When should I use the heuristic detector directly?

Access `detect_tables_heuristically` directly when processing PDFs known to contain implicit tables without grid lines or background rectangles, such as whitespace-aligned text columns or monospace font layouts. This method analyzes gap histograms and font clustering to identify tabular structures invisible to graphic-based detectors.

### How does rectangle-based detection work internally?

The rectangle detector in [`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs) extracts drawing rectangles from PDF content streams, including filled boxes and cell backgrounds. It implements a union-find algorithm to cluster adjacent rectangles into coherent grids, effectively reconstructing table cells from explicit geometric primitives embedded in the document.

### Can I disable specific detection methods?

While the default `process_pdf_with_options` function in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) runs the full pipeline automatically, you can bypass unwanted methods by calling specific detector functions directly from [`src/tables/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/mod.rs). Import individual detection modules and implement custom logic that skips rectangle or line detection based on your document characteristics.