# How to Handle Complex PDF Layouts with Firecrawl pdf‑inspector: A Complete Technical Guide

> Master complex PDF layouts with Firecrawl pdf-inspector. Learn its multi-stage pipeline for classification, parsing, column detection, table extraction, and Markdown generation. Get precise PDF data.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-07

---

**Firecrawl pdf‑inspector handles complex PDF layouts through a multi‑stage pipeline that combines PDF type classification, content stream parsing, column detection, three‑tiered table extraction, and intelligent Markdown generation.**

The `firecrawl/pdf‑inspector` repository implements a Rust‑based extraction engine designed specifically for the chaotic reality of real‑world PDFs. Whether you're processing multi‑column academic papers, scanned documents with mixed content, or tables without clear grid lines, the library provides granular control over every stage of the conversion pipeline. This guide walks through the architecture and shows how to leverage each component for your use case.

## PDF Type Classification and Pre‑Processing

Before any text extraction begins, [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) classifies the document into one of four categories: **TextBased**, **Scanned**, **Mixed**, or **ImageBased**. This classification determines which downstream strategies to apply.

The detector runs **tiled‑scan heuristics** to identify large scanned sections that may be split across multiple tiles in the PDF structure. This prevents the extractor from treating partial scan fragments as independent content regions.

Run classification standalone to understand your document:

```bash
detect-pdf my-document.pdf --json

```

## Content Stream Parsing and Text Positioning

The core extraction engine lives in [`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs). This module walks the PDF operator stream—processing operators like `Tj`, `TJ`, `Td`, and `Tm`—while maintaining a complete **graphics state** including current font, transformation matrix, and clipping path.

The output is a flat list of **TextItem** structures, each carrying precise positioning data. This positional information becomes the foundation for all subsequent layout analysis.

```rust
use pdf_inspector::process_pdf_with_options;

let opts = pdf_inspector::ProcessOptions::default();
let md = process_pdf_with_options("my-document.pdf", opts).unwrap();
println!("{}", md);

```

## Column Detection and Reading Order Resolution

Multi‑column layouts represent one of the most common failure modes in naive PDF extractors. The [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs) module solves this through **horizontal projection histograms**.

The algorithm:

- Projects text items onto the X‑axis to find density valleys
- Extracts column boundaries from these valleys
- Classifies the layout mode as **newspaper** (sequential column reading) or **tabular** (Y‑interleaved column reading)

This ensures that text flows in logical reading order rather than raw PDF content stream order.

## Three‑Tiered Table Detection Strategy

Tables without explicit borders require sophisticated detection. pdf‑inspector implements a **cascading fallback strategy** in `src/tables/`, where the first successful method wins:

| Method | Module | When It Applies |
|--------|--------|---------------|
| **Rect‑based detection** | [`detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/detect_rects.rs) | Tables with explicit cell rectangles; uses union‑find clustering |
| **Line‑based detection** | [`detect_lines.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/detect_lines.rs) | Tables defined by horizontal and vertical line primitives |
| **Heuristic detection** | [`detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/detect_heuristic.rs) | Loosely‑structured tables; uses gap‑histogram analysis and font‑size cues |

Once detected, [`src/tables/grid.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/grid.rs) constructs the cell grid, assigns text items to cells, and handles **spanning cells** (merged cells propagate unless column count exceeds the configured limit of 10).

## Markdown Generation and Post‑Processing

The [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) module produces clean, token‑efficient Markdown optimized for LLM consumption. It works in tandem with [`src/markdown/classify.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/classify.rs) to tag each line as header, list item, code block, or caption.

Post‑processing in [`src/markdown/postprocess.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/postprocess.rs) applies final cleanups:

- Dot‑leader removal (common in tables of contents)
- Hyphenation fixes across line breaks
- URL formatting normalization

Generate Markdown output via CLI:

```bash
pdf2md my-document.pdf

```

Or request structured JSON for programmatic pipelines:

```bash
pdf2md my-document.pdf --json > output.json

```

## Tagged PDF and Accessibility Support

When source PDFs contain a **structure tree** (PDF/UA or tagged PDF), [`src/structure_tree.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/structure_tree.rs) maps standard PDF roles directly to Markdown constructs:

- `H1`–`H6` → Markdown headers
- `P` → Paragraphs
- `L` → Lists
- `Code` → Code blocks
- `BlockQuote` → Blockquotes

If the structure tree is absent or incomplete, the system falls back to font‑size heuristics for logical structure inference.

## Python Integration

For Python workflows, use the provided bindings:

```python
from pdf_inspector import pdf2md

md = pdf2md("my-document.pdf")
print(md)

```

## Summary

- **Classify first**: Use [`detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/detector.rs) to understand your PDF type before selecting extraction strategies
- **Parse positionally**: [`content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/content_stream.rs) extracts precise text positioning for layout analysis
- **Resolve reading order**: [`layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/layout.rs) handles multi‑column newspaper and tabular layouts automatically
- **Detect tables robustly**: Three cascading methods in `src/tables/` ensure tables extract regardless of border style
- **Generate clean Markdown**: [`convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/convert.rs) and [`postprocess.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/postprocess.rs) optimize output for downstream LLM consumption
- **Leverage structure trees**: [`structure_tree.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/structure_tree.rs) preserves semantic markup when available in tagged PDFs

## Frequently Asked Questions

### How does pdf‑inspector handle scanned PDFs with no embedded text?

The [`detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/detector.rs) module classifies documents as **Scanned** or **Mixed** and applies tiled‑scan heuristics to identify scanned regions. These regions can then be routed to OCR pipelines (external to the core Rust extractor) while preserving text‑based regions for direct extraction.

### Can I configure the maximum number of columns for table detection?

Yes. The spanning‑cell logic in [`src/tables/grid.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/grid.rs) applies a configurable column limit (default 10) beyond which merged cells are no longer propagated. For documents with unusually wide tables, adjust this threshold in `ProcessOptions` when using the Rust API directly.

### What happens when multiple table detection methods conflict?

The system uses **short‑circuit evaluation**: [`detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/detect_rects.rs) runs first, and if it produces a valid table grid, that result is returned immediately. Only if rect detection fails does the pipeline fall back to line‑based detection, then heuristic detection. This prioritizes precision while maintaining coverage.

### Is the Markdown output optimized for specific LLM providers?

The [`convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/convert.rs) and [`postprocess.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/postprocess.rs) modules generate **token‑efficient Markdown** by default—compact headers, normalized whitespace, and consistent list formatting. This design targets general LLM consumption rather than provider‑specific formats, ensuring broad compatibility across OpenAI, Anthropic, and open‑source model APIs.