How to Handle Complex PDF Layouts with Firecrawl pdf‑inspector: A Complete Technical Guide
Firecrawl pdf‑inspector handles complex PDF layouts through a multi‑stage pipeline that combines PDF type classification, content stream parsing, column detection, three‑tiered table extraction, and intelligent Markdown generation.
The firecrawl/pdf‑inspector repository implements a Rust‑based extraction engine designed specifically for the chaotic reality of real‑world PDFs. Whether you're processing multi‑column academic papers, scanned documents with mixed content, or tables without clear grid lines, the library provides granular control over every stage of the conversion pipeline. This guide walks through the architecture and shows how to leverage each component for your use case.
PDF Type Classification and Pre‑Processing
Before any text extraction begins, src/detector.rs classifies the document into one of four categories: TextBased, Scanned, Mixed, or ImageBased. This classification determines which downstream strategies to apply.
The detector runs tiled‑scan heuristics to identify large scanned sections that may be split across multiple tiles in the PDF structure. This prevents the extractor from treating partial scan fragments as independent content regions.
Run classification standalone to understand your document:
detect-pdf my-document.pdf --json
Content Stream Parsing and Text Positioning
The core extraction engine lives in src/extractor/content_stream.rs. This module walks the PDF operator stream—processing operators like Tj, TJ, Td, and Tm—while maintaining a complete graphics state including current font, transformation matrix, and clipping path.
The output is a flat list of TextItem structures, each carrying precise positioning data. This positional information becomes the foundation for all subsequent layout analysis.
use pdf_inspector::process_pdf_with_options;
let opts = pdf_inspector::ProcessOptions::default();
let md = process_pdf_with_options("my-document.pdf", opts).unwrap();
println!("{}", md);
Column Detection and Reading Order Resolution
Multi‑column layouts represent one of the most common failure modes in naive PDF extractors. The src/extractor/layout.rs module solves this through horizontal projection histograms.
The algorithm:
- Projects text items onto the X‑axis to find density valleys
- Extracts column boundaries from these valleys
- Classifies the layout mode as newspaper (sequential column reading) or tabular (Y‑interleaved column reading)
This ensures that text flows in logical reading order rather than raw PDF content stream order.
Three‑Tiered Table Detection Strategy
Tables without explicit borders require sophisticated detection. pdf‑inspector implements a cascading fallback strategy in src/tables/, where the first successful method wins:
| Method | Module | When It Applies |
|---|---|---|
| Rect‑based detection | detect_rects.rs |
Tables with explicit cell rectangles; uses union‑find clustering |
| Line‑based detection | detect_lines.rs |
Tables defined by horizontal and vertical line primitives |
| Heuristic detection | detect_heuristic.rs |
Loosely‑structured tables; uses gap‑histogram analysis and font‑size cues |
Once detected, src/tables/grid.rs constructs the cell grid, assigns text items to cells, and handles spanning cells (merged cells propagate unless column count exceeds the configured limit of 10).
Markdown Generation and Post‑Processing
The src/markdown/convert.rs module produces clean, token‑efficient Markdown optimized for LLM consumption. It works in tandem with src/markdown/classify.rs to tag each line as header, list item, code block, or caption.
Post‑processing in src/markdown/postprocess.rs applies final cleanups:
- Dot‑leader removal (common in tables of contents)
- Hyphenation fixes across line breaks
- URL formatting normalization
Generate Markdown output via CLI:
pdf2md my-document.pdf
Or request structured JSON for programmatic pipelines:
pdf2md my-document.pdf --json > output.json
Tagged PDF and Accessibility Support
When source PDFs contain a structure tree (PDF/UA or tagged PDF), src/structure_tree.rs maps standard PDF roles directly to Markdown constructs:
H1–H6→ Markdown headersP→ ParagraphsL→ ListsCode→ Code blocksBlockQuote→ Blockquotes
If the structure tree is absent or incomplete, the system falls back to font‑size heuristics for logical structure inference.
Python Integration
For Python workflows, use the provided bindings:
from pdf_inspector import pdf2md
md = pdf2md("my-document.pdf")
print(md)
Summary
- Classify first: Use
detector.rsto understand your PDF type before selecting extraction strategies - Parse positionally:
content_stream.rsextracts precise text positioning for layout analysis - Resolve reading order:
layout.rshandles multi‑column newspaper and tabular layouts automatically - Detect tables robustly: Three cascading methods in
src/tables/ensure tables extract regardless of border style - Generate clean Markdown:
convert.rsandpostprocess.rsoptimize output for downstream LLM consumption - Leverage structure trees:
structure_tree.rspreserves semantic markup when available in tagged PDFs
Frequently Asked Questions
How does pdf‑inspector handle scanned PDFs with no embedded text?
The detector.rs module classifies documents as Scanned or Mixed and applies tiled‑scan heuristics to identify scanned regions. These regions can then be routed to OCR pipelines (external to the core Rust extractor) while preserving text‑based regions for direct extraction.
Can I configure the maximum number of columns for table detection?
Yes. The spanning‑cell logic in src/tables/grid.rs applies a configurable column limit (default 10) beyond which merged cells are no longer propagated. For documents with unusually wide tables, adjust this threshold in ProcessOptions when using the Rust API directly.
What happens when multiple table detection methods conflict?
The system uses short‑circuit evaluation: detect_rects.rs runs first, and if it produces a valid table grid, that result is returned immediately. Only if rect detection fails does the pipeline fall back to line‑based detection, then heuristic detection. This prioritizes precision while maintaining coverage.
Is the Markdown output optimized for specific LLM providers?
The convert.rs and postprocess.rs modules generate token‑efficient Markdown by default—compact headers, normalized whitespace, and consistent list formatting. This design targets general LLM consumption rather than provider‑specific formats, ensuring broad compatibility across OpenAI, Anthropic, and open‑source model APIs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →