How firecrawl pdf-inspector Extracts Tables from PDFs Using Three Detection Strategies
Yes, firecrawl pdf-inspector can extract tables from PDFs and render them as Markdown tables using a multi-layered detection pipeline that applies rect-based, line-based, and heuristic strategies in priority order.
The firecrawl/pdf-inspector repository is a Rust-based PDF extraction tool built for high-quality text and table recovery. Unlike simple text scrapers, its table extraction engine runs three complementary detection methods to handle everything from explicitly bordered tables to pure text-based layouts. This guide walks through how the table pipeline works, the specific modules involved, and how to use the API.
The Three-Layer Table Detection Pipeline
The core table extraction logic lives in src/tables/. When processing a page, the pipeline tries detection strategies in order of reliability, falling back to more permissive heuristics when structured clues are absent.
Rect-Based Detection (Most Reliable)
The detect_rects.rs module finds explicit PDF drawing rectangles—filled cells, border rectangles, and background shapes. By clustering these geometric primitives into grids, it reconstructs table structures with high accuracy.
This is the preferred strategy because it works directly with the PDF's vector graphics operators. When a table uses filled rectangles for zebra striping or borders, detect_tables_from_rects captures the precise cell boundaries.
Line-Based Detection (Stroke Paths)
When borders are drawn as stroke paths rather than filled rectangles, detect_lines.rs takes over. It analyzes horizontal and vertical line primitives to build a grid. This catches tables where designers used thin lines for visual separation without closed rectangular cells.
The detect_tables_from_lines function in this module assembles candidate grids from line intersections, filtering out decorative rules that don't form coherent tabular structures.
Heuristic Detection (Text-Only Fallback)
For tables without any visual borders, detect_heuristic.rs performs pure text analysis. It infers column boundaries from:
- X-position alignment of text items across rows
- Font-size patterns that distinguish headers from data cells
- Vertical spacing heuristics that detect row groupings
The detect_tables function exported from this module serves as the default entry point—it attempts all strategies internally and returns the best result.
Public API and Usage
Module Exports in src/tables/mod.rs
The public interface exposes each detector explicitly:
pub use detect_heuristic::detect_tables; // heuristic (default)
pub use detect_lines::detect_tables_from_lines; // line-based
pub use detect_rects::{detect_tables_from_rects, ...}; // rect-based
pub use detect_struct::detect_tables_from_struct_tree; // PDF structure tree
During extraction, src/extractor/mod.rs calls detect_tables on each page's TextItem collection. Detected tables populate Table structs that preserve row count, column count, and merged cell information.
Markdown Output via src/tables/format.rs
The formatter converts Table structs into standard Markdown syntax. Tables appear inline with other content—no post-processing required. The output is optimized for downstream LLM consumption and data pipelines.
Command-Line Usage
Default Markdown Extraction
# Extract full Markdown with tables rendered as Markdown tables
pdf2md my_report.pdf > my_report.md
Structured JSON Output
# Tables appear as a dedicated field for programmatic access
pdf2md --json my_report.pdf > my_report.json
The pdf2md binary is defined in src/bin/pdf2md.rs and uses ExtractorOptions::default(), which enables table detection automatically.
Programmatic Examples
Rust API
use pdf_inspector::process_pdf_with_options;
use pdf_inspector::extractor::ExtractorOptions;
let opts = ExtractorOptions::default(); // table detection enabled by default
let result = process_pdf_with_options("my_report.pdf", opts).unwrap();
for page in result.pages {
for tbl in page.tables {
println!("Page {} – Table with {} rows × {} cols",
page.number, tbl.rows.len(), tbl.columns.len());
}
}
Python FFI Wrapper
from pdf_inspector import pdf2md
markdown = pdf2md("my_report.pdf")
print(markdown) # tables appear as Markdown tables inline
Key Source Files
| File | Role |
|---|---|
src/tables/mod.rs |
Public API exports and detector orchestration |
src/tables/detect_rects.rs |
Rectangle-based detection (highest precision) |
src/tables/detect_lines.rs |
Line-stroke grid detection |
src/tables/detect_heuristic.rs |
Text-alignment heuristics (fallback) |
src/tables/format.rs |
Markdown serialization of Table structs |
src/extractor/mod.rs |
Per-page extraction that invokes table detectors |
src/bin/pdf2md.rs |
CLI entry point |
Summary
- firecrawl pdf-inspector extracts tables using three prioritized strategies: rects, lines, then text heuristics.
- Table detection runs before post-processing, so output preserves structure without additional steps.
detect_tablesinsrc/tables/detect_heuristic.rsis the default public API, while specialized functions allow direct access to specific strategies.- Markdown and JSON outputs both include tables—choose based on your downstream pipeline needs.
- The Rust-first design with Python FFI support makes it suitable for both systems integration and scripting workflows.
Frequently Asked Questions
Does pdf-inspector require tables to have visible borders?
No. While detect_rects.rs and detect_lines.rs handle bordered tables, detect_heuristic.rs detects borderless tables by analyzing text alignment, font patterns, and spacing. The default detect_tables function tries all three strategies automatically.
What output formats support table extraction?
Markdown (default) renders tables as standard | column | column | syntax. JSON mode (--json flag) includes tables as structured objects with rows, columns, and cells arrays. Both formats are produced by the same detection pipeline in src/extractor/mod.rs.
Can I disable table detection for faster processing?
The ExtractorOptions struct controls extraction behavior. While default() enables tables, you can construct custom options. Check src/extractor/mod.rs for the ExtractorOptions definition—table detection typically adds minimal overhead relative to full text extraction.
How does pdf-inspector handle merged cells?
The Table struct in src/tables/ preserves merged cell information detected during grid construction. When formatting to Markdown, merged cells are represented by empty cells in subsequent rows/columns, maintaining visual alignment. More complex spanning is preserved in JSON output's cell metadata.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →