# How the pdf-inspector Table Detection Pipeline Prioritizes Rect, Line, and Heuristic Methods

> Discover how the pdf-inspector table detection pipeline prioritizes rect, line, and heuristic methods. Learn its smart strategy for accurate table extraction.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: internals
- Published: 2026-08-06

---

**The pdf-inspector table detection pipeline follows a strict rect-based → line-based → heuristic priority order, trying the most reliable geometric signals first before falling back to text-based inference.**

The `firecrawl/pdf-inspector` repository implements a deterministic table extraction system that processes each PDF page in three successive passes. Understanding how the pdf-inspector table detection pipeline prioritizes its methods helps developers debug extraction results and integrate the library effectively across diverse document styles.

## Priority Order: Rect → Line → Heuristic

At the highest level, the extraction orchestrator in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) enforces a rigid sequence inside the `process_pdf_with_options` function. The engine attempts the most precise geometric cues before resorting to noisier signals.

The pipeline progresses through these stages:

1. **Rect-based detection** – searches for explicit rectangle drawing operators (`re`) that form cell borders.
2. **Line-based detection** – falls back to line drawing operators (`l`) when rects are absent or insufficient.
3. **Heuristic detection** – infers structure purely from text layout and spacing as a last resort.

This ordering ensures that **low-noise, explicit PDF graphics are preferred** over inferred text alignment, minimizing false positives.

## Stage 1: Rect-Based Detection

The first pass runs `detect_tables_from_rects` inside [`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs). It scans the page content stream for `re` (rectangle) operators, then clusters spatially overlapping rects. From these clusters it builds a grid using the rectangle edges and validates that grid against the page's text items.

If this pass produces tables and at least one table has more than three columns, the pipeline accepts the rect result immediately. The specific acceptance logic, as implemented in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), checks:

```rust
if !rect_tables.is_empty() && rect_tables.iter().any(|t| t.columns.len() > 3) {
    // use rect result
}

```

This guard prevents the engine from discarding a strong rect signal just because a narrow table was detected.

## Stage 2: Line-Based Detection

When the rect pass yields no tables—or only tables with three or fewer columns—the orchestrator falls back to `detect_tables_from_lines` in [`src/tables/detect_lines.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_lines.rs). This stage treats PDF drawing lines (`l` operators) as potential column and row edges. It clusters those lines and then runs the same grid-building logic used in the rect pass.

This method catches tables that are visually drawn with lines but lack explicit closed cell borders. Because line geometry is still an explicit drawing instruction, it remains more reliable than pure text inference.

## Stage 3: Heuristic Detection

If both geometric passes fail, the pipeline finally calls `detect_tables` from [`src/tables/detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs). This stage ignores all vector graphics and analyzes only the text layout. It infers column boundaries from text baselines, splits merged number tokens, and applies structural guards such as minimum fill-rate, prose-filtering, and column-count checks.

Because this heuristic method operates on alignment and spacing alone, it can discover borderless tables that rely purely on whitespace typography. It serves as the catch-all fallback for PDFs that contain tables without any underlying line art.

## Fallback Merges Within the Priority Order

The pipeline preserves its priority order even when individual strategies need internal recovery. Inside `detect_tables_from_rects`, two additional fallback mechanisms run before the stage is considered complete:

- **Merged-cluster fallback** – if no tables emerge from initial rect clusters, the engine merges all clusters and re-runs the line-stripe strategy.
- **Cell-rect fallback** – when rect clusters fail to produce a valid grid, the engine uses rect Y-edges for rows and text X-positions for columns.

These fallbacks are implemented in [`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs) and ensure the rect strategy exhausts its possibilities before the orchestrator moves to line-based detection.

## How to Call the Pipeline in Rust

The public API in [`src/tables/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/mod.rs) re-exports the three detection functions, making the priority visible to callers. To use the full orchestrated pipeline with automatic rect → line → heuristic ordering:

```rust
use pdf_inspector::lib::process_pdf_with_options;
use pdf_inspector::process_mode::ProcessMode;

let opts = ProcessMode::default();
let result = process_pdf_with_options("my_document.pdf", &opts).unwrap();

for (i, table) in result.tables.iter().enumerate() {
    println!("Table {}: {} cols × {} rows", i + 1, table.columns.len(), table.rows.len());
}

```

To run a specific stage manually, import the individual detectors:

```rust
use pdf_inspector::tables::{detect_tables_from_rects, detect_tables_from_lines, detect_tables};

let rects = /* extract PdfRect objects from the PDF */;
let items = /* extract TextItem objects from the same page */;

let (rect_tables, _) = detect_tables_from_rects(&items, &rects, page_number);
let line_tables = detect_tables_from_lines(&items, page_number);
let heuristic_tables = detect_tables(&items, page_width, false);

```

Direct invocation bypasses the orchestrator, so you must implement your own priority logic if you want to replicate the default behavior.

## Summary

- The pdf-inspector table detection pipeline processes every page in a fixed **rect-based → line-based → heuristic** order.
- [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) (`process_pdf_with_options`) enforces this priority and only advances to the next stage when the current one fails or returns overly narrow tables.
- `detect_tables_from_rects` in [`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs) handles explicit rectangle operators and includes internal merged-cluster and cell-rect fallbacks.
- `detect_tables_from_lines` in [`src/tables/detect_lines.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_lines.rs) captures line-drawn tables that lack closed cell rectangles.
- `detect_tables` in [`src/tables/detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs) provides a pure-text fallback for borderless, alignment-based tables.

## Frequently Asked Questions

### Why does pdf-inspector prioritize rect-based detection over line-based detection?

Rectangles provide closed cell boundaries, which yield unambiguous row and column edges with minimal noise. Lines can represent borders, dividers, or decorative underlines, so they carry slightly more ambiguity until clustered into a full grid.

### Can a page return results from more than one detection method at the same time?

No. The orchestrator in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) stops at the first successful stage that meets the acceptance criteria. Rect results are returned immediately if they exist and contain a table wider than three columns; otherwise the engine tries lines, and finally heuristics.

### What happens if the rect pass finds only small tables?

If `detect_tables_from_rects` returns tables but none have more than three columns, the orchestrator treats the rect pass as insufficient and proceeds to `detect_tables_from_lines`. This prevents narrow decorative boxes from blocking detection of larger line-based or heuristic tables.

### Is it possible to skip the geometric detectors and use only heuristic detection?

Yes. You can call `detect_tables` directly from [`src/tables/detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs), passing the page text items and width. However, this bypasses the robust geometric validation in the default pipeline, so it should be reserved for documents that are known to contain only borderless, alignment-based tables.