# How to Get Structured Data from PDFs with Firecrawl pdf-inspector: A Complete Guide

> Effortlessly extract structured data from PDFs using Firecrawl pdf-inspector. Convert PDFs to Markdown or JSON, automatically classify document types, and process scanned pages with OCR. Get started today!

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-07

---

**Firecrawl pdf-inspector converts any PDF into clean, structured Markdown or JSON while automatically classifying document types and routing only scanned pages to OCR.**

Getting structured data from PDFs traditionally requires brittle pipelines that either lose formatting or over-rely on expensive OCR. The **firecrawl/pdf-inspector** repository solves this with a pure-Rust library that extracts position-aware text, detects tables and headings, and outputs token-efficient Markdown—all while skipping OCR for truly text-based documents. This guide covers the complete processing pipeline, API options, and practical code examples for Rust, Python, Node.js, and CLI usage.

## How the PDF Processing Pipeline Works

The library implements a three-stage architecture that loads the PDF only once, eliminating redundant I/O by sharing the same `lopdf::Document` instance across stages【/cache/repos/github.com/firecrawl/pdf-inspector/main/src/lib.rs#L90-L94】.

### Stage 1: Fast PDF Type Detection

The `detect_pdf` and `detect_pdf_type` functions in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) perform a lightweight scan of content streams to classify documents as **TextBased**, **Scanned**, **ImageBased**, or **Mixed**. Each classification includes a confidence score and a list of pages requiring OCR【/cache/repos/github.com/firecrawl/pdf-inspector/main/README.md#L14-L22】.

This detection runs without full text extraction, making it ideal for routing decisions in high-throughput pipelines.

### Stage 2: Position-Aware Text Extraction

The `extract_text_with_positions` function in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) walks the PDF content stream once, resolves fonts and `ToUnicode` CMaps, and emits structured objects:

- **`TextItem`** – Text runs with X/Y coordinates, font size, and style metadata【/cache/repos/github.com/firecrawl/pdf-inspector/main/src/lib.rs#L60-L61】
- **`PdfRect`** – Drawing objects for table detection
- **`PdfLine`** – Line geometry for grid analysis【/cache/repos/github.com/firecrawl/pdf-inspector/main/src/lib.rs#L73-L78】

### Stage 3: Markdown Conversion and Layout Analysis

The [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) module applies heuristic analysis to produce hierarchical Markdown:

- **Heading levels** inferred from font-size tiers
- **Lists and code blocks** detected from glyph patterns
- **Tables** identified via three complementary strategies: rectangle-based, line-based, and heuristic detection【/cache/repos/github.com/firecrawl/pdf-inspector/main/README.md#L16-L18】

Post-processing handles drop-caps, dot-leaders, URL linking, and page-break markers【/cache/repos/github.com/firecrawl/pdf-inspector/main/README.md#L49-L55】.

## Core Architectural Components

| Component | Responsibility | Source File |
|-----------|----------------|-------------|
| **`detect_pdf` / `detect_pdf_type`** | Fast classification without full extraction | [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) |
| **`extract_text_with_positions`** | Content-stream walker emitting positioned items | [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) |
| **`tables` module** | Three-stage table detection and formatting | [`src/tables/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/mod.rs) |
| **`markdown` module** | Font analysis, heading detection, final rendering | [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) |
| **`process_pdf_with_options`** | Public API entry point returning `PdfProcessResult` | [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)【/cache/repos/github.com/firecrawl/pdf-inspector/main/src/lib.rs#L60-L66】 |

## API Methods for Structured Data Extraction

The library exposes multiple entry points depending on your integration needs. All share the same extraction engine for consistent results.

| API | Return Value | Best For |
|-----|------------|----------|
| `process_pdf` | Full `PdfProcessResult` with PDF type, Markdown, page count, OCR routing, layout complexity | General-purpose workflows |
| `process_pdf_mem` | Same as above, operates on in-memory bytes | Web services, zero-copy pipelines |
| `extract_pages_markdown` | Per-page `PageMarkdown` with `needs_ocr` flags and `ocr_reason` | Hybrid OCR/nativetext pipelines |
| `extract_text_in_regions_mem` | Native text for arbitrary rectangular regions (PDF points) | Downstream layout model integration |
| `extract_tables_in_regions_mem` | Markdown pipe-tables for regions, OCR fallback only on failure | Table-centric OCR avoidance |
| `detect_vector_grid_in_region_mem` | Low-level vector-grid description | TSR-compatible table extraction |

## Code Examples: Extracting Structured Data from PDFs

### Rust: Full Document Processing

```rust
use pdf_inspector::process_pdf;

fn main() -> Result<(), pdf_inspector::PdfError> {
    let result = process_pdf("reports/2024-annual.pdf")?;
    
    println!("Detected type: {:?}", result.pdf_type);
    
    if let Some(md) = result.markdown {
        println!("Markdown output:\n{}", md);
    }
    
    Ok(())
}

```

### Rust: Per-Page Extraction with OCR Routing

```rust
use pdf_inspector::extract_pages_markdown;

fn main() -> Result<(), pdf_inspector::PdfError> {
    let pages = extract_pages_markdown("invoices/batch.pdf", None)?;
    
    for page in pages.pages {
        if page.needs_ocr {
            println!("Page {} needs OCR", page.page + 1);
            // Route to OCR service
        } else {
            println!("Page {} markdown:\n{}", page.page + 1, page.markdown);
        }
    }
    
    Ok(())
}

```

### Python

```python
import pdf_inspector

result = pdf_inspector.process_pdf("contract.pdf")
print("PDF type:", result.pdf_type)          # "text_based", "scanned", etc.

print("Markdown:", result.markdown)          # Full structured Markdown

print("Pages needing OCR:", result.pages_needing_ocr)

```

### Node.js

```javascript
import { readFileSync } from "fs";
import { processPdf } from "@firecrawl/pdf-inspector";

const pdf = readFileSync("proposal.pdf");
const result = processPdf(pdf);

console.log("PDF type:", result.pdfType);          // "TextBased", "Scanned", …
console.log("Markdown:", result.markdown);
console.log("OCR pages:", result.pagesNeedingOcr);

```

### CLI

```bash

# Convert PDF to clean Markdown

pdf2md research_paper.pdf

# Get JSON items with positions for downstream models

pdf2md research_paper.pdf --items-json > items.json

# Run only the classifier (fast routing)

detect-pdf scanned_document.pdf --json

```

## Key Source Files for Custom Integration

| File | Purpose |
|------|---------|
| [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) | Public API, `PdfOptions` builder, convenience functions【/cache/repos/github.com/firecrawl/pdf-inspector/main/src/lib.rs#L60-L66】 |
| [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) | Ultra-fast PDF type scanner and OCR routing【/cache/repos/github.com/firecrawl/pdf-inspector/main/src/detector.rs】 |
| [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) | Core content-stream walker producing `TextItem`, `PdfRect`, `PdfLine`【/cache/repos/github.com/firecrawl/pdf-inspector/main/src/extractor/mod.rs】 |
| [`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs) | Rectangle-based table detection (first priority stage)【/cache/repos/github.com/firecrawl/pdf-inspector/main/src/tables/detect_rects.rs】 |
| [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) | Hierarchical Markdown generation with heading/list/code detection【/cache/repos/github.com/firecrawl/pdf-inspector/main/src/markdown/convert.rs】 |

## Summary

- **Firecrawl pdf-inspector** converts PDFs to structured Markdown through a three-stage pipeline: fast type detection, position-aware extraction, and heuristic layout analysis.
- **Single-pass architecture** shares a `lopdf::Document` instance across stages, eliminating redundant I/O【/cache/repos/github.com/firecrawl/pdf-inspector/main/src/lib.rs#L90-L94】.
- **Granular OCR control** routes only genuinely scanned pages to external services, preserving native text quality and reducing costs.
- **Multiple API surfaces** support full documents, per-page processing, region-based extraction, and table-specific workflows.

## Frequently Asked Questions

### How does firecrawl pdf-inspector detect whether a PDF needs OCR?

The `detect_pdf` and `detect_pdf_type` functions in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) perform a lightweight content-stream scan without full text extraction. They classify documents as TextBased, Scanned, ImageBased, or Mixed with confidence scores, returning specific page numbers that require OCR【/cache/repos/github.com/firecrawl/pdf-inspector/main/README.md#L14-L22】.

### What structured output formats does pdf-inspector support?

The library primarily outputs **clean Markdown** with hierarchical headings, tables, lists, and code blocks. For programmatic use, it returns `PdfProcessResult` structs containing PDF type metadata, per-page Markdown objects, OCR routing information, and optional JSON item streams with precise coordinates.

### Can I extract only specific regions or tables from a PDF?

Yes. The `extract_text_in_regions_mem` and `extract_tables_in_regions_mem` APIs accept rectangular regions in PDF points and return content for those areas only. The table-specific method returns Markdown pipe-tables and falls back to OCR only when detection fails—preserving native text quality where possible.

### Is firecrawl pdf-inspector available for languages other than Rust?

The core library is pure Rust, but bindings are available for **Python** and **Node.js** as shown in the examples above. The CLI tool `pdf2md` provides language-agnostic access for shell scripting and automation workflows.