# How to Use Firecrawl PDF-Inspector for Web Scraping PDFs

> Learn how to web scrape PDFs with Firecrawl PDF-Inspector. Convert PDFs to clean Markdown, perfect for AI and LLM integration. Get structured data easily.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-07

---

**Firecrawl PDF-Inspector converts PDF documents into clean, structured Markdown—optimized for AI agents and LLM pipelines rather than pixel-perfect visual rendering.**

This Rust-based library with Python bindings enables web scraping workflows to extract semantically meaningful text from PDFs, including robust table detection and reading-order reconstruction. The source code is available at `firecrawl/pdf-inspector` on GitHub.

## Core Architecture: Three-Stage Pipeline

PDF-Inspector processes documents through distinct detection, extraction, and conversion phases. Understanding this flow helps you select the right integration point for your scraping pipeline.

### Stage 1: PDF Type Detection

The `detect-pdf` CLI tool classifies documents before extraction. Located in [`src/bin/detect_pdf.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs), it calls into [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) to categorize PDFs as:

- **TextBased** – Native text with embedded fonts
- **Scanned** – Image-only pages requiring OCR
- **Mixed** – Combination of text and scanned regions
- **ImageBased** – Pure image content

The detector also performs tiled-scan detection for large JBIG2 or strip-based images. This classification lets your scraper route files appropriately—sending scanned documents to OCR services while processing text-based PDFs directly.

```bash
detect-pdf --json https://example.com/document.pdf

```

Sample output:

```json
{"type":"Mixed","tiled_scan":false}

```

### Stage 2: Text Extraction and Layout Reconstruction

The `pdf2md` binary ([`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs)) orchestrates extraction through [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs). Key subsystems include:

- **Font resolution** ([`src/extractor/fonts.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/fonts.rs)) – Maps PDF font objects to usable glyph data
- **Reading order** ([`src/extractor/reading_order.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/reading_order.rs)) – Reconstructs logical text flow across columns and complex layouts
- **Layout analysis** ([`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs)) – Detects columnar structures and spatial relationships

This stage parses content streams directly rather than relying on external tools, giving you fine-grained control over extraction behavior.

### Stage 3: Markdown Generation and Post-Processing

Extracted content flows through `src/markdown/` for final conversion:

| Module | Purpose |
|--------|---------|
| [`src/markdown/classify.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/classify.rs) | Tags lines as headers, lists, code blocks, or body text |
| [`src/markdown/preprocess.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/preprocess.rs) | Merges adjacent lines and normalizes spacing |
| [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) | Emits final Markdown with proper heading hierarchy |

Table detection runs three prioritized strategies in `src/tables/`:

1. **Rect-based detection** ([`detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/detect_rects.rs)) – Finds table boundaries from vector graphics
2. **Line-based detection** ([`detect_lines.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/detect_lines.rs)) – Identifies ruling lines and grid structures
3. **Heuristic detection** ([`detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/detect_heuristic.rs)) – Falls back to pattern-based column alignment

## Installation and CLI Usage

Install the Rust toolchain, then build and install PDF-Inspector directly from the repository:

```bash
cargo install --locked --git https://github.com/firecrawl/pdf-inspector.git pdf2md

```

Convert remote PDFs in a single command—ideal for quick scraping validation:

```bash
pdf2md https://example.com/annual-report.pdf > report.md

```

Pipe the output to downstream processing or inspect intermediate structure with `--json` flags where supported.

## Python Integration for Scraping Pipelines

The Python wrapper in [`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs) exposes `process_pdf()` for seamless integration with frameworks like Scrapy, BeautifulSoup, or `requests`-based fetchers.

```python
import requests
import pdf_inspector

def fetch_and_extract(url: str) -> str:
    """Download PDF and return Markdown content."""
    response = requests.get(url, timeout=30)
    response.raise_for_status()
    
    # pdf_bytes accepts raw bytes from any source

    markdown = pdf_inspector.process_pdf(response.content)
    return markdown

# Usage in your scraping workflow

content = fetch_and_extract("https://example.com/document.pdf")

```

This binding calls the same Rust core as the CLI, ensuring consistent output across interfaces.

## Rust Library Integration

For performance-critical scrapers, call PDF-Inspector directly from Rust. The public API surface in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) centers on `process_pdf_with_options()`:

```rust
use pdf_inspector::{process_pdf_with_options, ProcessMode};

fn extract_from_bytes(pdf_bytes: &[u8]) -> Result<String, Box<dyn std::error::Error>> {
    let markdown = process_pdf_with_options(
        pdf_bytes,
        ProcessMode::default(),
    )?;
    Ok(markdown)
}

```

Fetch PDFs via `reqwest` or `tokio` and stream directly into the extractor without filesystem overhead.

## Handling Scanned and Mixed PDFs

When `detect-pdf` returns `"type":"Scanned"` or `tiled_scan:true`, route documents to OCR before Markdown conversion:

```python
import pdf_inspector
import pytesseract  # or cloud OCR service

def extract_with_fallback(pdf_bytes: bytes) -> str:
    # Check document type first

    doc_type = pdf_inspector.detect_pdf_type(pdf_bytes)
    
    if doc_type.type in ("Scanned", "ImageBased") or doc_type.tiled_scan:
        # Convert to images, run OCR, then optionally pass to pdf_inspector

        # or use OCR output directly

        return ocr_pipeline(pdf_bytes)
    
    # Native text extraction

    return pdf_inspector.process_pdf(pdf_bytes)

```

This pre-flight check prevents wasted processing on documents that need external handling.

## Key Source Files Reference

| Path | Functionality |
|------|---------------|
| [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) | Public API: `process_pdf_with_options()`, encoding detection |
| [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) | PDF classification and tiled-scan flagging |
| [`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) | CLI entry for Markdown conversion |
| [`src/bin/detect_pdf.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs) | CLI for pre-extraction document analysis |
| [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) | Core extraction coordinator |
| [`src/extractor/reading_order.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/reading_order.rs) | Logical reading sequence reconstruction |
| [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs) | Column and spatial layout detection |
| [`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs) | Primary table boundary detection |
| [`src/tables/detect_lines.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_lines.rs) | Ruling-line table identification |
| [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) | Final Markdown emission |
| [`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs) | Python binding implementation |

## Summary

- **Firecrawl PDF-Inspector** generates token-efficient Markdown from PDFs—designed specifically for LLM consumption rather than visual fidelity
- **Detection first**: Use `detect-pdf` to categorize documents before extraction, handling scanned content separately
- **Multiple interfaces**: Choose CLI for prototyping, Python bindings for existing scrapers, or Rust library for performance-critical pipelines
- **Robust table handling**: Three-tier detection (rect → line → heuristic) preserves tabular data as Markdown tables
- **Modular source**: Core logic split across [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs), `src/extractor/`, `src/markdown/`, and `src/tables/` for targeted customization

## Frequently Asked Questions

### Can PDF-Inspector handle password-protected PDFs?

No. The current implementation in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) does not include decryption routines. Remove password protection using `qpdf` or similar tools before processing, or decrypt in your fetching layer with `pikepdf`.

### How does table detection compare to other PDF extractors?

PDF-Inspector's three-strategy cascade—rect-based graphics detection, ruling-line analysis, and heuristic column alignment—generally outperforms single-method approaches on academic papers and financial reports. For complex merged cells, the heuristic fallback in [`src/tables/detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs) uses whitespace patterns when vector graphics are unavailable.

### Is OCR built in or required separately?

OCR is external. PDF-Inspector detects when OCR is needed (via [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) and [`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs)) but does not perform it. Integrate Tesseract, AWS Textract, or similar services when `detect-pdf` indicates scanned content.

### What output formats are supported?

Markdown is the primary and optimized output. The conversion pipeline in [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) is purpose-built for clean, structured Markdown. For other formats, post-process the Markdown with tools like `pandoc` or modify `src/markdown/` directly.