How to Use Firecrawl PDF-Inspector for Web Scraping PDFs
Firecrawl PDF-Inspector converts PDF documents into clean, structured Markdown—optimized for AI agents and LLM pipelines rather than pixel-perfect visual rendering.
This Rust-based library with Python bindings enables web scraping workflows to extract semantically meaningful text from PDFs, including robust table detection and reading-order reconstruction. The source code is available at firecrawl/pdf-inspector on GitHub.
Core Architecture: Three-Stage Pipeline
PDF-Inspector processes documents through distinct detection, extraction, and conversion phases. Understanding this flow helps you select the right integration point for your scraping pipeline.
Stage 1: PDF Type Detection
The detect-pdf CLI tool classifies documents before extraction. Located in src/bin/detect_pdf.rs, it calls into src/detector.rs to categorize PDFs as:
- TextBased – Native text with embedded fonts
- Scanned – Image-only pages requiring OCR
- Mixed – Combination of text and scanned regions
- ImageBased – Pure image content
The detector also performs tiled-scan detection for large JBIG2 or strip-based images. This classification lets your scraper route files appropriately—sending scanned documents to OCR services while processing text-based PDFs directly.
detect-pdf --json https://example.com/document.pdf
Sample output:
{"type":"Mixed","tiled_scan":false}
Stage 2: Text Extraction and Layout Reconstruction
The pdf2md binary (src/bin/pdf2md.rs) orchestrates extraction through src/extractor/mod.rs. Key subsystems include:
- Font resolution (
src/extractor/fonts.rs) – Maps PDF font objects to usable glyph data - Reading order (
src/extractor/reading_order.rs) – Reconstructs logical text flow across columns and complex layouts - Layout analysis (
src/extractor/layout.rs) – Detects columnar structures and spatial relationships
This stage parses content streams directly rather than relying on external tools, giving you fine-grained control over extraction behavior.
Stage 3: Markdown Generation and Post-Processing
Extracted content flows through src/markdown/ for final conversion:
| Module | Purpose |
|---|---|
src/markdown/classify.rs |
Tags lines as headers, lists, code blocks, or body text |
src/markdown/preprocess.rs |
Merges adjacent lines and normalizes spacing |
src/markdown/convert.rs |
Emits final Markdown with proper heading hierarchy |
Table detection runs three prioritized strategies in src/tables/:
- Rect-based detection (
detect_rects.rs) – Finds table boundaries from vector graphics - Line-based detection (
detect_lines.rs) – Identifies ruling lines and grid structures - Heuristic detection (
detect_heuristic.rs) – Falls back to pattern-based column alignment
Installation and CLI Usage
Install the Rust toolchain, then build and install PDF-Inspector directly from the repository:
cargo install --locked --git https://github.com/firecrawl/pdf-inspector.git pdf2md
Convert remote PDFs in a single command—ideal for quick scraping validation:
pdf2md https://example.com/annual-report.pdf > report.md
Pipe the output to downstream processing or inspect intermediate structure with --json flags where supported.
Python Integration for Scraping Pipelines
The Python wrapper in src/python.rs exposes process_pdf() for seamless integration with frameworks like Scrapy, BeautifulSoup, or requests-based fetchers.
import requests
import pdf_inspector
def fetch_and_extract(url: str) -> str:
"""Download PDF and return Markdown content."""
response = requests.get(url, timeout=30)
response.raise_for_status()
# pdf_bytes accepts raw bytes from any source
markdown = pdf_inspector.process_pdf(response.content)
return markdown
# Usage in your scraping workflow
content = fetch_and_extract("https://example.com/document.pdf")
This binding calls the same Rust core as the CLI, ensuring consistent output across interfaces.
Rust Library Integration
For performance-critical scrapers, call PDF-Inspector directly from Rust. The public API surface in src/lib.rs centers on process_pdf_with_options():
use pdf_inspector::{process_pdf_with_options, ProcessMode};
fn extract_from_bytes(pdf_bytes: &[u8]) -> Result<String, Box<dyn std::error::Error>> {
let markdown = process_pdf_with_options(
pdf_bytes,
ProcessMode::default(),
)?;
Ok(markdown)
}
Fetch PDFs via reqwest or tokio and stream directly into the extractor without filesystem overhead.
Handling Scanned and Mixed PDFs
When detect-pdf returns "type":"Scanned" or tiled_scan:true, route documents to OCR before Markdown conversion:
import pdf_inspector
import pytesseract # or cloud OCR service
def extract_with_fallback(pdf_bytes: bytes) -> str:
# Check document type first
doc_type = pdf_inspector.detect_pdf_type(pdf_bytes)
if doc_type.type in ("Scanned", "ImageBased") or doc_type.tiled_scan:
# Convert to images, run OCR, then optionally pass to pdf_inspector
# or use OCR output directly
return ocr_pipeline(pdf_bytes)
# Native text extraction
return pdf_inspector.process_pdf(pdf_bytes)
This pre-flight check prevents wasted processing on documents that need external handling.
Key Source Files Reference
| Path | Functionality |
|---|---|
src/lib.rs |
Public API: process_pdf_with_options(), encoding detection |
src/detector.rs |
PDF classification and tiled-scan flagging |
src/bin/pdf2md.rs |
CLI entry for Markdown conversion |
src/bin/detect_pdf.rs |
CLI for pre-extraction document analysis |
src/extractor/mod.rs |
Core extraction coordinator |
src/extractor/reading_order.rs |
Logical reading sequence reconstruction |
src/extractor/layout.rs |
Column and spatial layout detection |
src/tables/detect_rects.rs |
Primary table boundary detection |
src/tables/detect_lines.rs |
Ruling-line table identification |
src/markdown/convert.rs |
Final Markdown emission |
src/python.rs |
Python binding implementation |
Summary
- Firecrawl PDF-Inspector generates token-efficient Markdown from PDFs—designed specifically for LLM consumption rather than visual fidelity
- Detection first: Use
detect-pdfto categorize documents before extraction, handling scanned content separately - Multiple interfaces: Choose CLI for prototyping, Python bindings for existing scrapers, or Rust library for performance-critical pipelines
- Robust table handling: Three-tier detection (rect → line → heuristic) preserves tabular data as Markdown tables
- Modular source: Core logic split across
src/detector.rs,src/extractor/,src/markdown/, andsrc/tables/for targeted customization
Frequently Asked Questions
Can PDF-Inspector handle password-protected PDFs?
No. The current implementation in src/lib.rs does not include decryption routines. Remove password protection using qpdf or similar tools before processing, or decrypt in your fetching layer with pikepdf.
How does table detection compare to other PDF extractors?
PDF-Inspector's three-strategy cascade—rect-based graphics detection, ruling-line analysis, and heuristic column alignment—generally outperforms single-method approaches on academic papers and financial reports. For complex merged cells, the heuristic fallback in src/tables/detect_heuristic.rs uses whitespace patterns when vector graphics are unavailable.
Is OCR built in or required separately?
OCR is external. PDF-Inspector detects when OCR is needed (via src/detector.rs and src/text_quality.rs) but does not perform it. Integrate Tesseract, AWS Textract, or similar services when detect-pdf indicates scanned content.
What output formats are supported?
Markdown is the primary and optimized output. The conversion pipeline in src/markdown/convert.rs is purpose-built for clean, structured Markdown. For other formats, post-process the Markdown with tools like pandoc or modify src/markdown/ directly.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →