# How firecrawl pdf-inspector Extracts Tables from PDFs Using Three Detection Strategies

> Learn how firecrawl pdf-inspector extracts tables from PDFs using rect-based, line-based, and heuristic strategies. Get your data in Markdown format.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-07

---

**Yes, firecrawl pdf-inspector can extract tables from PDFs and render them as Markdown tables using a multi-layered detection pipeline that applies rect-based, line-based, and heuristic strategies in priority order.**

The `firecrawl/pdf-inspector` repository is a Rust-based PDF extraction tool built for high-quality text and table recovery. Unlike simple text scrapers, its table extraction engine runs three complementary detection methods to handle everything from explicitly bordered tables to pure text-based layouts. This guide walks through how the table pipeline works, the specific modules involved, and how to use the API.

## The Three-Layer Table Detection Pipeline

The core table extraction logic lives in `src/tables/`. When processing a page, the pipeline tries detection strategies in order of reliability, falling back to more permissive heuristics when structured clues are absent.

### Rect-Based Detection (Most Reliable)

The **[`detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/detect_rects.rs)** module finds explicit PDF drawing rectangles—filled cells, border rectangles, and background shapes. By clustering these geometric primitives into grids, it reconstructs table structures with high accuracy.

This is the preferred strategy because it works directly with the PDF's vector graphics operators. When a table uses filled rectangles for zebra striping or borders, `detect_tables_from_rects` captures the precise cell boundaries.

### Line-Based Detection (Stroke Paths)

When borders are drawn as **stroke paths** rather than filled rectangles, **[`detect_lines.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/detect_lines.rs)** takes over. It analyzes horizontal and vertical line primitives to build a grid. This catches tables where designers used thin lines for visual separation without closed rectangular cells.

The `detect_tables_from_lines` function in this module assembles candidate grids from line intersections, filtering out decorative rules that don't form coherent tabular structures.

### Heuristic Detection (Text-Only Fallback)

For tables without any visual borders, **[`detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/detect_heuristic.rs)** performs pure text analysis. It infers column boundaries from:

- **X-position alignment** of text items across rows
- **Font-size patterns** that distinguish headers from data cells
- **Vertical spacing heuristics** that detect row groupings

The `detect_tables` function exported from this module serves as the **default entry point**—it attempts all strategies internally and returns the best result.

## Public API and Usage

### Module Exports in [`src/tables/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/mod.rs)

The public interface exposes each detector explicitly:

```rust
pub use detect_heuristic::detect_tables;               // heuristic (default)
pub use detect_lines::detect_tables_from_lines;        // line-based
pub use detect_rects::{detect_tables_from_rects, ...}; // rect-based
pub use detect_struct::detect_tables_from_struct_tree; // PDF structure tree

```

During extraction, [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) calls `detect_tables` on each page's `TextItem` collection. Detected tables populate `Table` structs that preserve **row count**, **column count**, and **merged cell information**.

### Markdown Output via [`src/tables/format.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/format.rs)

The formatter converts `Table` structs into standard Markdown syntax. Tables appear inline with other content—no post-processing required. The output is optimized for downstream LLM consumption and data pipelines.

## Command-Line Usage

### Default Markdown Extraction

```bash

# Extract full Markdown with tables rendered as Markdown tables

pdf2md my_report.pdf > my_report.md

```

### Structured JSON Output

```bash

# Tables appear as a dedicated field for programmatic access

pdf2md --json my_report.pdf > my_report.json

```

The `pdf2md` binary is defined in [`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) and uses `ExtractorOptions::default()`, which enables table detection automatically.

## Programmatic Examples

### Rust API

```rust
use pdf_inspector::process_pdf_with_options;
use pdf_inspector::extractor::ExtractorOptions;

let opts = ExtractorOptions::default();           // table detection enabled by default
let result = process_pdf_with_options("my_report.pdf", opts).unwrap();

for page in result.pages {
    for tbl in page.tables {
        println!("Page {} – Table with {} rows × {} cols",
                 page.number, tbl.rows.len(), tbl.columns.len());
    }
}

```

### Python FFI Wrapper

```python
from pdf_inspector import pdf2md

markdown = pdf2md("my_report.pdf")
print(markdown)                     # tables appear as Markdown tables inline

```

## Key Source Files

| File | Role |
|------|------|
| [`src/tables/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/mod.rs) | Public API exports and detector orchestration |
| [`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs) | Rectangle-based detection (highest precision) |
| [`src/tables/detect_lines.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_lines.rs) | Line-stroke grid detection |
| [`src/tables/detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs) | Text-alignment heuristics (fallback) |
| [`src/tables/format.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/format.rs) | Markdown serialization of `Table` structs |
| [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) | Per-page extraction that invokes table detectors |
| [`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) | CLI entry point |

## Summary

- **firecrawl pdf-inspector extracts tables** using three prioritized strategies: rects, lines, then text heuristics.
- **Table detection runs before post-processing**, so output preserves structure without additional steps.
- **`detect_tables`** in [`src/tables/detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs) is the default public API, while specialized functions allow direct access to specific strategies.
- **Markdown and JSON outputs** both include tables—choose based on your downstream pipeline needs.
- The Rust-first design with Python FFI support makes it suitable for both systems integration and scripting workflows.

## Frequently Asked Questions

### Does pdf-inspector require tables to have visible borders?

No. While **[`detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/detect_rects.rs)** and **[`detect_lines.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/detect_lines.rs)** handle bordered tables, **[`detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/detect_heuristic.rs)** detects borderless tables by analyzing text alignment, font patterns, and spacing. The default `detect_tables` function tries all three strategies automatically.

### What output formats support table extraction?

**Markdown** (default) renders tables as standard `| column | column |` syntax. **JSON** mode (`--json` flag) includes tables as structured objects with `rows`, `columns`, and `cells` arrays. Both formats are produced by the same detection pipeline in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs).

### Can I disable table detection for faster processing?

The `ExtractorOptions` struct controls extraction behavior. While `default()` enables tables, you can construct custom options. Check [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) for the `ExtractorOptions` definition—table detection typically adds minimal overhead relative to full text extraction.

### How does pdf-inspector handle merged cells?

The `Table` struct in `src/tables/` preserves merged cell information detected during grid construction. When formatting to Markdown, merged cells are represented by empty cells in subsequent rows/columns, maintaining visual alignment. More complex spanning is preserved in JSON output's cell metadata.