Can pdf-inspector Extract Tables from PDFs? A Complete Technical Guide

Yes, pdf-inspector can extract tables from PDFs using a multi-layered detection pipeline that combines rectangle-based, line-based, and heuristic text analysis to handle diverse document layouts.

The firecrawl/pdf-inspector Rust library is engineered specifically to locate and extract tabular data from PDF documents. Unlike simple text extractors, it implements three independent detection strategies that merge geometric analysis with intelligent text parsing, making it possible to extract tables from everything from structured financial reports to scanned documents with minimal formatting.

Three-Stage Table Detection Architecture

The table extraction system in pdf-inspector operates through a prioritized cascade of three distinct detection methods. The public API in src/tables/mod.rs exposes these through focused re-exports:

pub use detect_heuristic::detect_tables;                     // high‑level heuristic detector
pub use detect_lines::detect_tables_from_lines;             // line‑based detector
pub use detect_rects::{detect_tables_from_rects, RectHintRegion}; // rectangle‑based detector

Stage 1: Rectangle-Based Detection

The primary detection strategy scans PDF content streams for drawing operators (re) that define cell-border rectangles. Located in src/tables/detect_rects.rs, this implementation uses a union-find clustering algorithm to group spatially overlapping rectangles, then constructs a grid from the rectangle edges to generate a Table object.

This approach excels with PDFs that use explicit rectangular borders to define cells, common in generated reports and formatted spreadsheets.

Stage 2: Line-Based Detection

When rectangle operators are insufficient, pdf-inspector falls back to line-based detection via src/tables/detect_lines.rs. This module identifies explicit horizontal and vertical line operators in the PDF content stream, groups them into intersecting sets, and derives row and column boundaries from the line intersections.

The detect_vector_grid_tables_from_lines function assembles these geometric elements into structured tables, handling documents that use line art rather than filled rectangles for table borders.

Stage 3: Heuristic Text Analysis

For documents where geometric cues are absent—such as scanned PDFs or minimally formatted text dumps—the library employs src/tables/detect_heuristic.rs. This pure-text analysis pipeline merges adjacent glyph items, normalizes financial-item fragments (like currency amounts split across lines), and identifies column boundaries through text-anchor pattern recognition.

The detect_tables function validates the resulting grid structure before returning finalized table objects, ensuring robust extraction even when PDF metadata is stripped or corrupted.

How to Extract Tables Using pdf-inspector

The crate exposes table extraction capabilities through both a command-line interface and direct Rust API calls. The top-level src/lib.rs re-exports these functions through the pdf_inspector::tables module.

CLI Table Extraction with pdf2md

The pdf2md binary located in src/bin/pdf2md.rs provides immediate access to table extraction without writing code. It internally invokes extractor::extract_text followed by markdown::to_markdown, which triggers the full detection pipeline:


# Extract text and tables to Markdown

pdf2md my_report.pdf > my_report.md

# Extract structured data including tables as JSON

pdf2md --json my_report.pdf > my_report.json

The CLI attempts detection in priority order (rectangles, then lines, then heuristics) and formats output using the tables::format::table_to_markdown utilities.

Programmatic Table Detection in Rust

For applications requiring granular control, the library exposes low-level detection functions that operate on parsed PDF structures. The process_pdf function in src/lib.rs populates the items and rects vectors required by the geometric detectors:

use pdf_inspector::{tables::detect_tables_from_rects, ProcessMode};

fn main() -> Result<(), Box<dyn std::error::Error>> {
    // Load and parse the PDF document
    let result = pdf_inspector::process_pdf("financial_report.pdf")?;
    
    // Extract tables using rectangle detection
    let (tables, _hints) = detect_tables_from_rects(
        &result.items,          // Vec<TextItem> from extractor
        &result.rects,          // Vec<PdfRect> from geometry parser
        1,                      // page number (1-indexed)
    );
    
    // Convert to Markdown using the Table struct's helper
    for (i, table) in tables.iter().enumerate() {
        println!("Table {}:\n{}", i + 1, table.to_markdown());
    }
    Ok(())
}

The Table struct provides formatting helpers implemented in src/tables/format.rs, including conversion to Markdown and JSON serialization.

Accessing the Heuristic Fallback

When geometric data is unavailable, call the heuristic detector directly via detect_tables from src/tables/mod.rs:

use pdf_inspector::tables::detect_tables;

// text_items obtained from extractor::fonts or extractor::content_stream
let tables = detect_tables(&text_items, page_width);

This function serves as a thin wrapper around detect_heuristic::detect_tables, performing glyph merging and column inference without requiring PdfRect or PdfLine inputs.

Inside the Detection Pipeline

Understanding the internal data flow clarifies how pdf-inspector achieves high accuracy across diverse PDF implementations. The pipeline executes five distinct phases:

  1. Geometry Collection: extractor::fonts and extractor::content_stream parse the PDF content stream to expose PdfRect rectangles and PdfLine line segments.
  2. Spatial Clustering: detect_rects::cluster_rects groups overlapping rectangles using union-find; detect_lines groups intersecting line segments.
  3. Grid Construction: detect_table_from_rect_group and detect_vector_grid_tables_from_lines snap edges to compute precise column and row boundaries, then assign TextItem objects to cells.
  4. Heuristic Processing: When geometric detection returns empty results, detect_heuristic::detect_tables merges glyph items and expands consolidated financial numbers to infer grid structure.
  5. Validation and Formatting: The tables::format module validates grid integrity and converts Table structs to Markdown or JSON representations.

This architecture ensures that if one strategy fails, the system automatically proceeds to the next, maximizing extraction success rates across the heterogeneous PDF ecosystem.

Summary

  • pdf-inspector extracts tables through three complementary strategies: rectangle-based detection, line-based detection, and heuristic text analysis.
  • The detection pipeline is implemented across src/tables/detect_rects.rs, src/tables/detect_lines.rs, and src/tables/detect_heuristic.rs, with public APIs exposed via src/tables/mod.rs.
  • Users can extract tables via the pdf2md CLI tool or programmatically using Rust functions like detect_tables_from_rects and detect_tables.
  • The system automatically falls back from geometric to heuristic methods when PDF structure varies, ensuring robust extraction from both formatted reports and plain text documents.
  • Output formatting is handled by src/tables/format.rs, supporting both Markdown and JSON serialization.

Frequently Asked Questions

How does pdf-inspector handle tables without visible borders?

When geometric cues are absent, pdf-inspector activates the heuristic detector in src/tables/detect_heuristic.rs. This module analyzes text anchor patterns, merges adjacent glyph items, and identifies column boundaries through spatial alignment rather than drawn lines, enabling extraction from borderless tables and scanned documents.

Can I extract tables from specific pages only?

Yes. The detection functions accept a page number parameter. For example, detect_tables_from_rects takes a 1-indexed page number as its third argument, allowing you to target specific pages. The process_pdf function also provides mechanisms to process individual pages before passing vectors to the detection algorithms.

What output formats does pdf-inspector support for extracted tables?

The library supports Markdown and JSON output. The Table struct implements to_markdown() for human-readable tables, while the pdf2md CLI offers a --json flag that serializes table structures including cell coordinates and text content. The formatting logic resides in src/tables/format.rs.

Is pdf-inspector suitable for extracting financial tables with merged cells?

Yes. The rectangle-based detector handles complex cell geometries through union-find clustering, while the heuristic detector specifically normalizes financial-item fragments split across lines. The detect_heuristic.rs implementation expands consolidated financial numbers and validates grid structures, making it effective for financial reports with irregular cell merging.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →