How the pdf-inspector Table Detection Pipeline Prioritizes Rect, Line, and Heuristic Methods
The pdf-inspector table detection pipeline follows a strict rect-based → line-based → heuristic priority order, trying the most reliable geometric signals first before falling back to text-based inference.
The firecrawl/pdf-inspector repository implements a deterministic table extraction system that processes each PDF page in three successive passes. Understanding how the pdf-inspector table detection pipeline prioritizes its methods helps developers debug extraction results and integrate the library effectively across diverse document styles.
Priority Order: Rect → Line → Heuristic
At the highest level, the extraction orchestrator in src/lib.rs enforces a rigid sequence inside the process_pdf_with_options function. The engine attempts the most precise geometric cues before resorting to noisier signals.
The pipeline progresses through these stages:
- Rect-based detection – searches for explicit rectangle drawing operators (
re) that form cell borders. - Line-based detection – falls back to line drawing operators (
l) when rects are absent or insufficient. - Heuristic detection – infers structure purely from text layout and spacing as a last resort.
This ordering ensures that low-noise, explicit PDF graphics are preferred over inferred text alignment, minimizing false positives.
Stage 1: Rect-Based Detection
The first pass runs detect_tables_from_rects inside src/tables/detect_rects.rs. It scans the page content stream for re (rectangle) operators, then clusters spatially overlapping rects. From these clusters it builds a grid using the rectangle edges and validates that grid against the page's text items.
If this pass produces tables and at least one table has more than three columns, the pipeline accepts the rect result immediately. The specific acceptance logic, as implemented in src/lib.rs, checks:
if !rect_tables.is_empty() && rect_tables.iter().any(|t| t.columns.len() > 3) {
// use rect result
}
This guard prevents the engine from discarding a strong rect signal just because a narrow table was detected.
Stage 2: Line-Based Detection
When the rect pass yields no tables—or only tables with three or fewer columns—the orchestrator falls back to detect_tables_from_lines in src/tables/detect_lines.rs. This stage treats PDF drawing lines (l operators) as potential column and row edges. It clusters those lines and then runs the same grid-building logic used in the rect pass.
This method catches tables that are visually drawn with lines but lack explicit closed cell borders. Because line geometry is still an explicit drawing instruction, it remains more reliable than pure text inference.
Stage 3: Heuristic Detection
If both geometric passes fail, the pipeline finally calls detect_tables from src/tables/detect_heuristic.rs. This stage ignores all vector graphics and analyzes only the text layout. It infers column boundaries from text baselines, splits merged number tokens, and applies structural guards such as minimum fill-rate, prose-filtering, and column-count checks.
Because this heuristic method operates on alignment and spacing alone, it can discover borderless tables that rely purely on whitespace typography. It serves as the catch-all fallback for PDFs that contain tables without any underlying line art.
Fallback Merges Within the Priority Order
The pipeline preserves its priority order even when individual strategies need internal recovery. Inside detect_tables_from_rects, two additional fallback mechanisms run before the stage is considered complete:
- Merged-cluster fallback – if no tables emerge from initial rect clusters, the engine merges all clusters and re-runs the line-stripe strategy.
- Cell-rect fallback – when rect clusters fail to produce a valid grid, the engine uses rect Y-edges for rows and text X-positions for columns.
These fallbacks are implemented in src/tables/detect_rects.rs and ensure the rect strategy exhausts its possibilities before the orchestrator moves to line-based detection.
How to Call the Pipeline in Rust
The public API in src/tables/mod.rs re-exports the three detection functions, making the priority visible to callers. To use the full orchestrated pipeline with automatic rect → line → heuristic ordering:
use pdf_inspector::lib::process_pdf_with_options;
use pdf_inspector::process_mode::ProcessMode;
let opts = ProcessMode::default();
let result = process_pdf_with_options("my_document.pdf", &opts).unwrap();
for (i, table) in result.tables.iter().enumerate() {
println!("Table {}: {} cols × {} rows", i + 1, table.columns.len(), table.rows.len());
}
To run a specific stage manually, import the individual detectors:
use pdf_inspector::tables::{detect_tables_from_rects, detect_tables_from_lines, detect_tables};
let rects = /* extract PdfRect objects from the PDF */;
let items = /* extract TextItem objects from the same page */;
let (rect_tables, _) = detect_tables_from_rects(&items, &rects, page_number);
let line_tables = detect_tables_from_lines(&items, page_number);
let heuristic_tables = detect_tables(&items, page_width, false);
Direct invocation bypasses the orchestrator, so you must implement your own priority logic if you want to replicate the default behavior.
Summary
- The pdf-inspector table detection pipeline processes every page in a fixed rect-based → line-based → heuristic order.
src/lib.rs(process_pdf_with_options) enforces this priority and only advances to the next stage when the current one fails or returns overly narrow tables.detect_tables_from_rectsinsrc/tables/detect_rects.rshandles explicit rectangle operators and includes internal merged-cluster and cell-rect fallbacks.detect_tables_from_linesinsrc/tables/detect_lines.rscaptures line-drawn tables that lack closed cell rectangles.detect_tablesinsrc/tables/detect_heuristic.rsprovides a pure-text fallback for borderless, alignment-based tables.
Frequently Asked Questions
Why does pdf-inspector prioritize rect-based detection over line-based detection?
Rectangles provide closed cell boundaries, which yield unambiguous row and column edges with minimal noise. Lines can represent borders, dividers, or decorative underlines, so they carry slightly more ambiguity until clustered into a full grid.
Can a page return results from more than one detection method at the same time?
No. The orchestrator in src/lib.rs stops at the first successful stage that meets the acceptance criteria. Rect results are returned immediately if they exist and contain a table wider than three columns; otherwise the engine tries lines, and finally heuristics.
What happens if the rect pass finds only small tables?
If detect_tables_from_rects returns tables but none have more than three columns, the orchestrator treats the rect pass as insufficient and proceeds to detect_tables_from_lines. This prevents narrow decorative boxes from blocking detection of larger line-based or heuristic tables.
Is it possible to skip the geometric detectors and use only heuristic detection?
Yes. You can call detect_tables directly from src/tables/detect_heuristic.rs, passing the page text items and width. However, this bypasses the robust geometric validation in the default pipeline, so it should be reserved for documents that are known to contain only borderless, alignment-based tables.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →