# Can pdf-inspector Extract Tables from PDFs? A Complete Technical Guide

> Learn how pdf-inspector extracts tables from PDFs. Discover its multi-layered detection pipeline combining rectangle, line, and heuristic analysis for diverse layouts.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-04

---

**Yes, pdf-inspector can extract tables from PDFs using a multi-layered detection pipeline that combines rectangle-based, line-based, and heuristic text analysis to handle diverse document layouts.**

The `firecrawl/pdf-inspector` Rust library is engineered specifically to locate and extract tabular data from PDF documents. Unlike simple text extractors, it implements three independent detection strategies that merge geometric analysis with intelligent text parsing, making it possible to extract tables from everything from structured financial reports to scanned documents with minimal formatting.

## Three-Stage Table Detection Architecture

The table extraction system in `pdf-inspector` operates through a prioritized cascade of three distinct detection methods. The public API in [`src/tables/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/mod.rs) exposes these through focused re-exports:

```rust
pub use detect_heuristic::detect_tables;                     // high‑level heuristic detector
pub use detect_lines::detect_tables_from_lines;             // line‑based detector
pub use detect_rects::{detect_tables_from_rects, RectHintRegion}; // rectangle‑based detector

```

### Stage 1: Rectangle-Based Detection

The primary detection strategy scans PDF content streams for drawing operators (`re`) that define cell-border rectangles. Located in [`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs), this implementation uses a union-find clustering algorithm to group spatially overlapping rectangles, then constructs a grid from the rectangle edges to generate a `Table` object.

This approach excels with PDFs that use explicit rectangular borders to define cells, common in generated reports and formatted spreadsheets.

### Stage 2: Line-Based Detection

When rectangle operators are insufficient, `pdf-inspector` falls back to line-based detection via [`src/tables/detect_lines.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_lines.rs). This module identifies explicit horizontal and vertical line operators in the PDF content stream, groups them into intersecting sets, and derives row and column boundaries from the line intersections.

The `detect_vector_grid_tables_from_lines` function assembles these geometric elements into structured tables, handling documents that use line art rather than filled rectangles for table borders.

### Stage 3: Heuristic Text Analysis

For documents where geometric cues are absent—such as scanned PDFs or minimally formatted text dumps—the library employs [`src/tables/detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs). This pure-text analysis pipeline merges adjacent glyph items, normalizes financial-item fragments (like currency amounts split across lines), and identifies column boundaries through text-anchor pattern recognition.

The `detect_tables` function validates the resulting grid structure before returning finalized table objects, ensuring robust extraction even when PDF metadata is stripped or corrupted.

## How to Extract Tables Using pdf-inspector

The crate exposes table extraction capabilities through both a command-line interface and direct Rust API calls. The top-level [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) re-exports these functions through the `pdf_inspector::tables` module.

### CLI Table Extraction with pdf2md

The `pdf2md` binary located in [`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) provides immediate access to table extraction without writing code. It internally invokes `extractor::extract_text` followed by `markdown::to_markdown`, which triggers the full detection pipeline:

```bash

# Extract text and tables to Markdown

pdf2md my_report.pdf > my_report.md

# Extract structured data including tables as JSON

pdf2md --json my_report.pdf > my_report.json

```

The CLI attempts detection in priority order (rectangles, then lines, then heuristics) and formats output using the `tables::format::table_to_markdown` utilities.

### Programmatic Table Detection in Rust

For applications requiring granular control, the library exposes low-level detection functions that operate on parsed PDF structures. The `process_pdf` function in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) populates the `items` and `rects` vectors required by the geometric detectors:

```rust
use pdf_inspector::{tables::detect_tables_from_rects, ProcessMode};

fn main() -> Result<(), Box<dyn std::error::Error>> {
    // Load and parse the PDF document
    let result = pdf_inspector::process_pdf("financial_report.pdf")?;
    
    // Extract tables using rectangle detection
    let (tables, _hints) = detect_tables_from_rects(
        &result.items,          // Vec<TextItem> from extractor
        &result.rects,          // Vec<PdfRect> from geometry parser
        1,                      // page number (1-indexed)
    );
    
    // Convert to Markdown using the Table struct's helper
    for (i, table) in tables.iter().enumerate() {
        println!("Table {}:\n{}", i + 1, table.to_markdown());
    }
    Ok(())
}

```

The `Table` struct provides formatting helpers implemented in [`src/tables/format.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/format.rs), including conversion to Markdown and JSON serialization.

### Accessing the Heuristic Fallback

When geometric data is unavailable, call the heuristic detector directly via `detect_tables` from [`src/tables/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/mod.rs):

```rust
use pdf_inspector::tables::detect_tables;

// text_items obtained from extractor::fonts or extractor::content_stream
let tables = detect_tables(&text_items, page_width);

```

This function serves as a thin wrapper around `detect_heuristic::detect_tables`, performing glyph merging and column inference without requiring `PdfRect` or `PdfLine` inputs.

## Inside the Detection Pipeline

Understanding the internal data flow clarifies how `pdf-inspector` achieves high accuracy across diverse PDF implementations. The pipeline executes five distinct phases:

1. **Geometry Collection**: `extractor::fonts` and `extractor::content_stream` parse the PDF content stream to expose `PdfRect` rectangles and `PdfLine` line segments.
2. **Spatial Clustering**: `detect_rects::cluster_rects` groups overlapping rectangles using union-find; `detect_lines` groups intersecting line segments.
3. **Grid Construction**: `detect_table_from_rect_group` and `detect_vector_grid_tables_from_lines` snap edges to compute precise column and row boundaries, then assign `TextItem` objects to cells.
4. **Heuristic Processing**: When geometric detection returns empty results, `detect_heuristic::detect_tables` merges glyph items and expands consolidated financial numbers to infer grid structure.
5. **Validation and Formatting**: The `tables::format` module validates grid integrity and converts `Table` structs to Markdown or JSON representations.

This architecture ensures that if one strategy fails, the system automatically proceeds to the next, maximizing extraction success rates across the heterogeneous PDF ecosystem.

## Summary

- **pdf-inspector** extracts tables through three complementary strategies: rectangle-based detection, line-based detection, and heuristic text analysis.
- The detection pipeline is implemented across [`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs), [`src/tables/detect_lines.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_lines.rs), and [`src/tables/detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs), with public APIs exposed via [`src/tables/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/mod.rs).
- Users can extract tables via the `pdf2md` CLI tool or programmatically using Rust functions like `detect_tables_from_rects` and `detect_tables`.
- The system automatically falls back from geometric to heuristic methods when PDF structure varies, ensuring robust extraction from both formatted reports and plain text documents.
- Output formatting is handled by [`src/tables/format.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/format.rs), supporting both Markdown and JSON serialization.

## Frequently Asked Questions

### How does pdf-inspector handle tables without visible borders?

When geometric cues are absent, `pdf-inspector` activates the heuristic detector in [`src/tables/detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs). This module analyzes text anchor patterns, merges adjacent glyph items, and identifies column boundaries through spatial alignment rather than drawn lines, enabling extraction from borderless tables and scanned documents.

### Can I extract tables from specific pages only?

Yes. The detection functions accept a page number parameter. For example, `detect_tables_from_rects` takes a 1-indexed page number as its third argument, allowing you to target specific pages. The `process_pdf` function also provides mechanisms to process individual pages before passing vectors to the detection algorithms.

### What output formats does pdf-inspector support for extracted tables?

The library supports Markdown and JSON output. The `Table` struct implements `to_markdown()` for human-readable tables, while the `pdf2md` CLI offers a `--json` flag that serializes table structures including cell coordinates and text content. The formatting logic resides in [`src/tables/format.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/format.rs).

### Is pdf-inspector suitable for extracting financial tables with merged cells?

Yes. The rectangle-based detector handles complex cell geometries through union-find clustering, while the heuristic detector specifically normalizes financial-item fragments split across lines. The [`detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/detect_heuristic.rs) implementation expands consolidated financial numbers and validates grid structures, making it effective for financial reports with irregular cell merging.