How to Get Structured Data from PDFs with Firecrawl pdf-inspector: A Complete Guide

Firecrawl pdf-inspector converts any PDF into clean, structured Markdown or JSON while automatically classifying document types and routing only scanned pages to OCR.

Getting structured data from PDFs traditionally requires brittle pipelines that either lose formatting or over-rely on expensive OCR. The firecrawl/pdf-inspector repository solves this with a pure-Rust library that extracts position-aware text, detects tables and headings, and outputs token-efficient Markdown—all while skipping OCR for truly text-based documents. This guide covers the complete processing pipeline, API options, and practical code examples for Rust, Python, Node.js, and CLI usage.

How the PDF Processing Pipeline Works

The library implements a three-stage architecture that loads the PDF only once, eliminating redundant I/O by sharing the same lopdf::Document instance across stages【/cache/repos/github.com/firecrawl/pdf-inspector/main/src/lib.rs#L90-L94】.

Stage 1: Fast PDF Type Detection

The detect_pdf and detect_pdf_type functions in src/detector.rs perform a lightweight scan of content streams to classify documents as TextBased, Scanned, ImageBased, or Mixed. Each classification includes a confidence score and a list of pages requiring OCR【/cache/repos/github.com/firecrawl/pdf-inspector/main/README.md#L14-L22】.

This detection runs without full text extraction, making it ideal for routing decisions in high-throughput pipelines.

Stage 2: Position-Aware Text Extraction

The extract_text_with_positions function in src/extractor/mod.rs walks the PDF content stream once, resolves fonts and ToUnicode CMaps, and emits structured objects:

  • TextItem – Text runs with X/Y coordinates, font size, and style metadata【/cache/repos/github.com/firecrawl/pdf-inspector/main/src/lib.rs#L60-L61】
  • PdfRect – Drawing objects for table detection
  • PdfLine – Line geometry for grid analysis【/cache/repos/github.com/firecrawl/pdf-inspector/main/src/lib.rs#L73-L78】

Stage 3: Markdown Conversion and Layout Analysis

The src/markdown/convert.rs module applies heuristic analysis to produce hierarchical Markdown:

  • Heading levels inferred from font-size tiers
  • Lists and code blocks detected from glyph patterns
  • Tables identified via three complementary strategies: rectangle-based, line-based, and heuristic detection【/cache/repos/github.com/firecrawl/pdf-inspector/main/README.md#L16-L18】

Post-processing handles drop-caps, dot-leaders, URL linking, and page-break markers【/cache/repos/github.com/firecrawl/pdf-inspector/main/README.md#L49-L55】.

Core Architectural Components

Component Responsibility Source File
detect_pdf / detect_pdf_type Fast classification without full extraction src/detector.rs
extract_text_with_positions Content-stream walker emitting positioned items src/extractor/mod.rs
tables module Three-stage table detection and formatting src/tables/mod.rs
markdown module Font analysis, heading detection, final rendering src/markdown/convert.rs
process_pdf_with_options Public API entry point returning PdfProcessResult src/lib.rs【/cache/repos/github.com/firecrawl/pdf-inspector/main/src/lib.rs#L60-L66】

API Methods for Structured Data Extraction

The library exposes multiple entry points depending on your integration needs. All share the same extraction engine for consistent results.

API Return Value Best For
process_pdf Full PdfProcessResult with PDF type, Markdown, page count, OCR routing, layout complexity General-purpose workflows
process_pdf_mem Same as above, operates on in-memory bytes Web services, zero-copy pipelines
extract_pages_markdown Per-page PageMarkdown with needs_ocr flags and ocr_reason Hybrid OCR/nativetext pipelines
extract_text_in_regions_mem Native text for arbitrary rectangular regions (PDF points) Downstream layout model integration
extract_tables_in_regions_mem Markdown pipe-tables for regions, OCR fallback only on failure Table-centric OCR avoidance
detect_vector_grid_in_region_mem Low-level vector-grid description TSR-compatible table extraction

Code Examples: Extracting Structured Data from PDFs

Rust: Full Document Processing

use pdf_inspector::process_pdf;

fn main() -> Result<(), pdf_inspector::PdfError> {
    let result = process_pdf("reports/2024-annual.pdf")?;
    
    println!("Detected type: {:?}", result.pdf_type);
    
    if let Some(md) = result.markdown {
        println!("Markdown output:\n{}", md);
    }
    
    Ok(())
}

Rust: Per-Page Extraction with OCR Routing

use pdf_inspector::extract_pages_markdown;

fn main() -> Result<(), pdf_inspector::PdfError> {
    let pages = extract_pages_markdown("invoices/batch.pdf", None)?;
    
    for page in pages.pages {
        if page.needs_ocr {
            println!("Page {} needs OCR", page.page + 1);
            // Route to OCR service
        } else {
            println!("Page {} markdown:\n{}", page.page + 1, page.markdown);
        }
    }
    
    Ok(())
}

Python

import pdf_inspector

result = pdf_inspector.process_pdf("contract.pdf")
print("PDF type:", result.pdf_type)          # "text_based", "scanned", etc.

print("Markdown:", result.markdown)          # Full structured Markdown

print("Pages needing OCR:", result.pages_needing_ocr)

Node.js

import { readFileSync } from "fs";
import { processPdf } from "@firecrawl/pdf-inspector";

const pdf = readFileSync("proposal.pdf");
const result = processPdf(pdf);

console.log("PDF type:", result.pdfType);          // "TextBased", "Scanned", …
console.log("Markdown:", result.markdown);
console.log("OCR pages:", result.pagesNeedingOcr);

CLI


# Convert PDF to clean Markdown

pdf2md research_paper.pdf

# Get JSON items with positions for downstream models

pdf2md research_paper.pdf --items-json > items.json

# Run only the classifier (fast routing)

detect-pdf scanned_document.pdf --json

Key Source Files for Custom Integration

File Purpose
src/lib.rs Public API, PdfOptions builder, convenience functions【/cache/repos/github.com/firecrawl/pdf-inspector/main/src/lib.rs#L60-L66】
src/detector.rs Ultra-fast PDF type scanner and OCR routing【/cache/repos/github.com/firecrawl/pdf-inspector/main/src/detector.rs】
src/extractor/mod.rs Core content-stream walker producing TextItem, PdfRect, PdfLine【/cache/repos/github.com/firecrawl/pdf-inspector/main/src/extractor/mod.rs】
src/tables/detect_rects.rs Rectangle-based table detection (first priority stage)【/cache/repos/github.com/firecrawl/pdf-inspector/main/src/tables/detect_rects.rs】
src/markdown/convert.rs Hierarchical Markdown generation with heading/list/code detection【/cache/repos/github.com/firecrawl/pdf-inspector/main/src/markdown/convert.rs】

Summary

  • Firecrawl pdf-inspector converts PDFs to structured Markdown through a three-stage pipeline: fast type detection, position-aware extraction, and heuristic layout analysis.
  • Single-pass architecture shares a lopdf::Document instance across stages, eliminating redundant I/O【/cache/repos/github.com/firecrawl/pdf-inspector/main/src/lib.rs#L90-L94】.
  • Granular OCR control routes only genuinely scanned pages to external services, preserving native text quality and reducing costs.
  • Multiple API surfaces support full documents, per-page processing, region-based extraction, and table-specific workflows.

Frequently Asked Questions

How does firecrawl pdf-inspector detect whether a PDF needs OCR?

The detect_pdf and detect_pdf_type functions in src/detector.rs perform a lightweight content-stream scan without full text extraction. They classify documents as TextBased, Scanned, ImageBased, or Mixed with confidence scores, returning specific page numbers that require OCR【/cache/repos/github.com/firecrawl/pdf-inspector/main/README.md#L14-L22】.

What structured output formats does pdf-inspector support?

The library primarily outputs clean Markdown with hierarchical headings, tables, lists, and code blocks. For programmatic use, it returns PdfProcessResult structs containing PDF type metadata, per-page Markdown objects, OCR routing information, and optional JSON item streams with precise coordinates.

Can I extract only specific regions or tables from a PDF?

Yes. The extract_text_in_regions_mem and extract_tables_in_regions_mem APIs accept rectangular regions in PDF points and return content for those areas only. The table-specific method returns Markdown pipe-tables and falls back to OCR only when detection fails—preserving native text quality where possible.

Is firecrawl pdf-inspector available for languages other than Rust?

The core library is pure Rust, but bindings are available for Python and Node.js as shown in the examples above. The CLI tool pdf2md provides language-agnostic access for shell scripting and automation workflows.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →