How to Handle Complex PDF Layouts with Firecrawl pdf‑inspector: A Complete Technical Guide

Firecrawl pdf‑inspector handles complex PDF layouts through a multi‑stage pipeline that combines PDF type classification, content stream parsing, column detection, three‑tiered table extraction, and intelligent Markdown generation.

The firecrawl/pdf‑inspector repository implements a Rust‑based extraction engine designed specifically for the chaotic reality of real‑world PDFs. Whether you're processing multi‑column academic papers, scanned documents with mixed content, or tables without clear grid lines, the library provides granular control over every stage of the conversion pipeline. This guide walks through the architecture and shows how to leverage each component for your use case.

PDF Type Classification and Pre‑Processing

Before any text extraction begins, src/detector.rs classifies the document into one of four categories: TextBased, Scanned, Mixed, or ImageBased. This classification determines which downstream strategies to apply.

The detector runs tiled‑scan heuristics to identify large scanned sections that may be split across multiple tiles in the PDF structure. This prevents the extractor from treating partial scan fragments as independent content regions.

Run classification standalone to understand your document:

detect-pdf my-document.pdf --json

Content Stream Parsing and Text Positioning

The core extraction engine lives in src/extractor/content_stream.rs. This module walks the PDF operator stream—processing operators like Tj, TJ, Td, and Tm—while maintaining a complete graphics state including current font, transformation matrix, and clipping path.

The output is a flat list of TextItem structures, each carrying precise positioning data. This positional information becomes the foundation for all subsequent layout analysis.

use pdf_inspector::process_pdf_with_options;

let opts = pdf_inspector::ProcessOptions::default();
let md = process_pdf_with_options("my-document.pdf", opts).unwrap();
println!("{}", md);

Column Detection and Reading Order Resolution

Multi‑column layouts represent one of the most common failure modes in naive PDF extractors. The src/extractor/layout.rs module solves this through horizontal projection histograms.

The algorithm:

  • Projects text items onto the X‑axis to find density valleys
  • Extracts column boundaries from these valleys
  • Classifies the layout mode as newspaper (sequential column reading) or tabular (Y‑interleaved column reading)

This ensures that text flows in logical reading order rather than raw PDF content stream order.

Three‑Tiered Table Detection Strategy

Tables without explicit borders require sophisticated detection. pdf‑inspector implements a cascading fallback strategy in src/tables/, where the first successful method wins:

Method Module When It Applies
Rect‑based detection detect_rects.rs Tables with explicit cell rectangles; uses union‑find clustering
Line‑based detection detect_lines.rs Tables defined by horizontal and vertical line primitives
Heuristic detection detect_heuristic.rs Loosely‑structured tables; uses gap‑histogram analysis and font‑size cues

Once detected, src/tables/grid.rs constructs the cell grid, assigns text items to cells, and handles spanning cells (merged cells propagate unless column count exceeds the configured limit of 10).

Markdown Generation and Post‑Processing

The src/markdown/convert.rs module produces clean, token‑efficient Markdown optimized for LLM consumption. It works in tandem with src/markdown/classify.rs to tag each line as header, list item, code block, or caption.

Post‑processing in src/markdown/postprocess.rs applies final cleanups:

  • Dot‑leader removal (common in tables of contents)
  • Hyphenation fixes across line breaks
  • URL formatting normalization

Generate Markdown output via CLI:

pdf2md my-document.pdf

Or request structured JSON for programmatic pipelines:

pdf2md my-document.pdf --json > output.json

Tagged PDF and Accessibility Support

When source PDFs contain a structure tree (PDF/UA or tagged PDF), src/structure_tree.rs maps standard PDF roles directly to Markdown constructs:

  • H1–H6 → Markdown headers
  • P → Paragraphs
  • L → Lists
  • Code → Code blocks
  • BlockQuote → Blockquotes

If the structure tree is absent or incomplete, the system falls back to font‑size heuristics for logical structure inference.

Python Integration

For Python workflows, use the provided bindings:

from pdf_inspector import pdf2md

md = pdf2md("my-document.pdf")
print(md)

Summary

  • Classify first: Use detector.rs to understand your PDF type before selecting extraction strategies
  • Parse positionally: content_stream.rs extracts precise text positioning for layout analysis
  • Resolve reading order: layout.rs handles multi‑column newspaper and tabular layouts automatically
  • Detect tables robustly: Three cascading methods in src/tables/ ensure tables extract regardless of border style
  • Generate clean Markdown: convert.rs and postprocess.rs optimize output for downstream LLM consumption
  • Leverage structure trees: structure_tree.rs preserves semantic markup when available in tagged PDFs

Frequently Asked Questions

How does pdf‑inspector handle scanned PDFs with no embedded text?

The detector.rs module classifies documents as Scanned or Mixed and applies tiled‑scan heuristics to identify scanned regions. These regions can then be routed to OCR pipelines (external to the core Rust extractor) while preserving text‑based regions for direct extraction.

Can I configure the maximum number of columns for table detection?

Yes. The spanning‑cell logic in src/tables/grid.rs applies a configurable column limit (default 10) beyond which merged cells are no longer propagated. For documents with unusually wide tables, adjust this threshold in ProcessOptions when using the Rust API directly.

What happens when multiple table detection methods conflict?

The system uses short‑circuit evaluation: detect_rects.rs runs first, and if it produces a valid table grid, that result is returned immediately. Only if rect detection fails does the pipeline fall back to line‑based detection, then heuristic detection. This prioritizes precision while maintaining coverage.

Is the Markdown output optimized for specific LLM providers?

The convert.rs and postprocess.rs modules generate token‑efficient Markdown by default—compact headers, normalized whitespace, and consistent list formatting. This design targets general LLM consumption rather than provider‑specific formats, ensuring broad compatibility across OpenAI, Anthropic, and open‑source model APIs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →