How to Use Firecrawl PDF-Inspector for Web Scraping PDFs

Firecrawl PDF-Inspector converts PDF documents into clean, structured Markdown—optimized for AI agents and LLM pipelines rather than pixel-perfect visual rendering.

This Rust-based library with Python bindings enables web scraping workflows to extract semantically meaningful text from PDFs, including robust table detection and reading-order reconstruction. The source code is available at firecrawl/pdf-inspector on GitHub.

Core Architecture: Three-Stage Pipeline

PDF-Inspector processes documents through distinct detection, extraction, and conversion phases. Understanding this flow helps you select the right integration point for your scraping pipeline.

Stage 1: PDF Type Detection

The detect-pdf CLI tool classifies documents before extraction. Located in src/bin/detect_pdf.rs, it calls into src/detector.rs to categorize PDFs as:

  • TextBased – Native text with embedded fonts
  • Scanned – Image-only pages requiring OCR
  • Mixed – Combination of text and scanned regions
  • ImageBased – Pure image content

The detector also performs tiled-scan detection for large JBIG2 or strip-based images. This classification lets your scraper route files appropriately—sending scanned documents to OCR services while processing text-based PDFs directly.

detect-pdf --json https://example.com/document.pdf

Sample output:

{"type":"Mixed","tiled_scan":false}

Stage 2: Text Extraction and Layout Reconstruction

The pdf2md binary (src/bin/pdf2md.rs) orchestrates extraction through src/extractor/mod.rs. Key subsystems include:

This stage parses content streams directly rather than relying on external tools, giving you fine-grained control over extraction behavior.

Stage 3: Markdown Generation and Post-Processing

Extracted content flows through src/markdown/ for final conversion:

Module Purpose
src/markdown/classify.rs Tags lines as headers, lists, code blocks, or body text
src/markdown/preprocess.rs Merges adjacent lines and normalizes spacing
src/markdown/convert.rs Emits final Markdown with proper heading hierarchy

Table detection runs three prioritized strategies in src/tables/:

  1. Rect-based detection (detect_rects.rs) – Finds table boundaries from vector graphics
  2. Line-based detection (detect_lines.rs) – Identifies ruling lines and grid structures
  3. Heuristic detection (detect_heuristic.rs) – Falls back to pattern-based column alignment

Installation and CLI Usage

Install the Rust toolchain, then build and install PDF-Inspector directly from the repository:

cargo install --locked --git https://github.com/firecrawl/pdf-inspector.git pdf2md

Convert remote PDFs in a single command—ideal for quick scraping validation:

pdf2md https://example.com/annual-report.pdf > report.md

Pipe the output to downstream processing or inspect intermediate structure with --json flags where supported.

Python Integration for Scraping Pipelines

The Python wrapper in src/python.rs exposes process_pdf() for seamless integration with frameworks like Scrapy, BeautifulSoup, or requests-based fetchers.

import requests
import pdf_inspector

def fetch_and_extract(url: str) -> str:
    """Download PDF and return Markdown content."""
    response = requests.get(url, timeout=30)
    response.raise_for_status()
    
    # pdf_bytes accepts raw bytes from any source

    markdown = pdf_inspector.process_pdf(response.content)
    return markdown

# Usage in your scraping workflow

content = fetch_and_extract("https://example.com/document.pdf")

This binding calls the same Rust core as the CLI, ensuring consistent output across interfaces.

Rust Library Integration

For performance-critical scrapers, call PDF-Inspector directly from Rust. The public API surface in src/lib.rs centers on process_pdf_with_options():

use pdf_inspector::{process_pdf_with_options, ProcessMode};

fn extract_from_bytes(pdf_bytes: &[u8]) -> Result<String, Box<dyn std::error::Error>> {
    let markdown = process_pdf_with_options(
        pdf_bytes,
        ProcessMode::default(),
    )?;
    Ok(markdown)
}

Fetch PDFs via reqwest or tokio and stream directly into the extractor without filesystem overhead.

Handling Scanned and Mixed PDFs

When detect-pdf returns "type":"Scanned" or tiled_scan:true, route documents to OCR before Markdown conversion:

import pdf_inspector
import pytesseract  # or cloud OCR service

def extract_with_fallback(pdf_bytes: bytes) -> str:
    # Check document type first

    doc_type = pdf_inspector.detect_pdf_type(pdf_bytes)
    
    if doc_type.type in ("Scanned", "ImageBased") or doc_type.tiled_scan:
        # Convert to images, run OCR, then optionally pass to pdf_inspector

        # or use OCR output directly

        return ocr_pipeline(pdf_bytes)
    
    # Native text extraction

    return pdf_inspector.process_pdf(pdf_bytes)

This pre-flight check prevents wasted processing on documents that need external handling.

Key Source Files Reference

Path Functionality
src/lib.rs Public API: process_pdf_with_options(), encoding detection
src/detector.rs PDF classification and tiled-scan flagging
src/bin/pdf2md.rs CLI entry for Markdown conversion
src/bin/detect_pdf.rs CLI for pre-extraction document analysis
src/extractor/mod.rs Core extraction coordinator
src/extractor/reading_order.rs Logical reading sequence reconstruction
src/extractor/layout.rs Column and spatial layout detection
src/tables/detect_rects.rs Primary table boundary detection
src/tables/detect_lines.rs Ruling-line table identification
src/markdown/convert.rs Final Markdown emission
src/python.rs Python binding implementation

Summary

  • Firecrawl PDF-Inspector generates token-efficient Markdown from PDFs—designed specifically for LLM consumption rather than visual fidelity
  • Detection first: Use detect-pdf to categorize documents before extraction, handling scanned content separately
  • Multiple interfaces: Choose CLI for prototyping, Python bindings for existing scrapers, or Rust library for performance-critical pipelines
  • Robust table handling: Three-tier detection (rect → line → heuristic) preserves tabular data as Markdown tables
  • Modular source: Core logic split across src/detector.rs, src/extractor/, src/markdown/, and src/tables/ for targeted customization

Frequently Asked Questions

Can PDF-Inspector handle password-protected PDFs?

No. The current implementation in src/lib.rs does not include decryption routines. Remove password protection using qpdf or similar tools before processing, or decrypt in your fetching layer with pikepdf.

How does table detection compare to other PDF extractors?

PDF-Inspector's three-strategy cascade—rect-based graphics detection, ruling-line analysis, and heuristic column alignment—generally outperforms single-method approaches on academic papers and financial reports. For complex merged cells, the heuristic fallback in src/tables/detect_heuristic.rs uses whitespace patterns when vector graphics are unavailable.

Is OCR built in or required separately?

OCR is external. PDF-Inspector detects when OCR is needed (via src/detector.rs and src/text_quality.rs) but does not perform it. Integrate Tesseract, AWS Textract, or similar services when detect-pdf indicates scanned content.

What output formats are supported?

Markdown is the primary and optimized output. The conversion pipeline in src/markdown/convert.rs is purpose-built for clean, structured Markdown. For other formats, post-process the Markdown with tools like pandoc or modify src/markdown/ directly.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →