Can pdf-inspector Extract Tables from PDFs? A Complete Technical Guide
Yes, pdf-inspector can extract tables from PDFs using a multi-layered detection pipeline that combines rectangle-based, line-based, and heuristic text analysis to handle diverse document layouts.
The firecrawl/pdf-inspector Rust library is engineered specifically to locate and extract tabular data from PDF documents. Unlike simple text extractors, it implements three independent detection strategies that merge geometric analysis with intelligent text parsing, making it possible to extract tables from everything from structured financial reports to scanned documents with minimal formatting.
Three-Stage Table Detection Architecture
The table extraction system in pdf-inspector operates through a prioritized cascade of three distinct detection methods. The public API in src/tables/mod.rs exposes these through focused re-exports:
pub use detect_heuristic::detect_tables; // high‑level heuristic detector
pub use detect_lines::detect_tables_from_lines; // line‑based detector
pub use detect_rects::{detect_tables_from_rects, RectHintRegion}; // rectangle‑based detector
Stage 1: Rectangle-Based Detection
The primary detection strategy scans PDF content streams for drawing operators (re) that define cell-border rectangles. Located in src/tables/detect_rects.rs, this implementation uses a union-find clustering algorithm to group spatially overlapping rectangles, then constructs a grid from the rectangle edges to generate a Table object.
This approach excels with PDFs that use explicit rectangular borders to define cells, common in generated reports and formatted spreadsheets.
Stage 2: Line-Based Detection
When rectangle operators are insufficient, pdf-inspector falls back to line-based detection via src/tables/detect_lines.rs. This module identifies explicit horizontal and vertical line operators in the PDF content stream, groups them into intersecting sets, and derives row and column boundaries from the line intersections.
The detect_vector_grid_tables_from_lines function assembles these geometric elements into structured tables, handling documents that use line art rather than filled rectangles for table borders.
Stage 3: Heuristic Text Analysis
For documents where geometric cues are absent—such as scanned PDFs or minimally formatted text dumps—the library employs src/tables/detect_heuristic.rs. This pure-text analysis pipeline merges adjacent glyph items, normalizes financial-item fragments (like currency amounts split across lines), and identifies column boundaries through text-anchor pattern recognition.
The detect_tables function validates the resulting grid structure before returning finalized table objects, ensuring robust extraction even when PDF metadata is stripped or corrupted.
How to Extract Tables Using pdf-inspector
The crate exposes table extraction capabilities through both a command-line interface and direct Rust API calls. The top-level src/lib.rs re-exports these functions through the pdf_inspector::tables module.
CLI Table Extraction with pdf2md
The pdf2md binary located in src/bin/pdf2md.rs provides immediate access to table extraction without writing code. It internally invokes extractor::extract_text followed by markdown::to_markdown, which triggers the full detection pipeline:
# Extract text and tables to Markdown
pdf2md my_report.pdf > my_report.md
# Extract structured data including tables as JSON
pdf2md --json my_report.pdf > my_report.json
The CLI attempts detection in priority order (rectangles, then lines, then heuristics) and formats output using the tables::format::table_to_markdown utilities.
Programmatic Table Detection in Rust
For applications requiring granular control, the library exposes low-level detection functions that operate on parsed PDF structures. The process_pdf function in src/lib.rs populates the items and rects vectors required by the geometric detectors:
use pdf_inspector::{tables::detect_tables_from_rects, ProcessMode};
fn main() -> Result<(), Box<dyn std::error::Error>> {
// Load and parse the PDF document
let result = pdf_inspector::process_pdf("financial_report.pdf")?;
// Extract tables using rectangle detection
let (tables, _hints) = detect_tables_from_rects(
&result.items, // Vec<TextItem> from extractor
&result.rects, // Vec<PdfRect> from geometry parser
1, // page number (1-indexed)
);
// Convert to Markdown using the Table struct's helper
for (i, table) in tables.iter().enumerate() {
println!("Table {}:\n{}", i + 1, table.to_markdown());
}
Ok(())
}
The Table struct provides formatting helpers implemented in src/tables/format.rs, including conversion to Markdown and JSON serialization.
Accessing the Heuristic Fallback
When geometric data is unavailable, call the heuristic detector directly via detect_tables from src/tables/mod.rs:
use pdf_inspector::tables::detect_tables;
// text_items obtained from extractor::fonts or extractor::content_stream
let tables = detect_tables(&text_items, page_width);
This function serves as a thin wrapper around detect_heuristic::detect_tables, performing glyph merging and column inference without requiring PdfRect or PdfLine inputs.
Inside the Detection Pipeline
Understanding the internal data flow clarifies how pdf-inspector achieves high accuracy across diverse PDF implementations. The pipeline executes five distinct phases:
- Geometry Collection:
extractor::fontsandextractor::content_streamparse the PDF content stream to exposePdfRectrectangles andPdfLineline segments. - Spatial Clustering:
detect_rects::cluster_rectsgroups overlapping rectangles using union-find;detect_linesgroups intersecting line segments. - Grid Construction:
detect_table_from_rect_groupanddetect_vector_grid_tables_from_linessnap edges to compute precise column and row boundaries, then assignTextItemobjects to cells. - Heuristic Processing: When geometric detection returns empty results,
detect_heuristic::detect_tablesmerges glyph items and expands consolidated financial numbers to infer grid structure. - Validation and Formatting: The
tables::formatmodule validates grid integrity and convertsTablestructs to Markdown or JSON representations.
This architecture ensures that if one strategy fails, the system automatically proceeds to the next, maximizing extraction success rates across the heterogeneous PDF ecosystem.
Summary
- pdf-inspector extracts tables through three complementary strategies: rectangle-based detection, line-based detection, and heuristic text analysis.
- The detection pipeline is implemented across
src/tables/detect_rects.rs,src/tables/detect_lines.rs, andsrc/tables/detect_heuristic.rs, with public APIs exposed viasrc/tables/mod.rs. - Users can extract tables via the
pdf2mdCLI tool or programmatically using Rust functions likedetect_tables_from_rectsanddetect_tables. - The system automatically falls back from geometric to heuristic methods when PDF structure varies, ensuring robust extraction from both formatted reports and plain text documents.
- Output formatting is handled by
src/tables/format.rs, supporting both Markdown and JSON serialization.
Frequently Asked Questions
How does pdf-inspector handle tables without visible borders?
When geometric cues are absent, pdf-inspector activates the heuristic detector in src/tables/detect_heuristic.rs. This module analyzes text anchor patterns, merges adjacent glyph items, and identifies column boundaries through spatial alignment rather than drawn lines, enabling extraction from borderless tables and scanned documents.
Can I extract tables from specific pages only?
Yes. The detection functions accept a page number parameter. For example, detect_tables_from_rects takes a 1-indexed page number as its third argument, allowing you to target specific pages. The process_pdf function also provides mechanisms to process individual pages before passing vectors to the detection algorithms.
What output formats does pdf-inspector support for extracted tables?
The library supports Markdown and JSON output. The Table struct implements to_markdown() for human-readable tables, while the pdf2md CLI offers a --json flag that serializes table structures including cell coordinates and text content. The formatting logic resides in src/tables/format.rs.
Is pdf-inspector suitable for extracting financial tables with merged cells?
Yes. The rectangle-based detector handles complex cell geometries through union-find clustering, while the heuristic detector specifically normalizes financial-item fragments split across lines. The detect_heuristic.rs implementation expands consolidated financial numbers and validates grid structures, making it effective for financial reports with irregular cell merging.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →