How to Use Firecrawl pdf‑inspector for Programmatic PDF Analysis

Yes—Firecrawl pdf‑inspector is purpose-built for programmatic PDF analysis, offering native Rust APIs with bindings for Python, Node.js, and WebAssembly.

This open‑source library exposes a consistent extraction and classification engine across all runtimes, making it ideal for automated workflows in scripts, backend services, or browser applications. The same core logic powers every interface, so your choice of language doesn't compromise analysis quality.

Programmatic Interfaces Overview

The firecrawl/pdf-inspector repository structures its capabilities into layered components that you can access through multiple entry points:

Interface Package Name Best For
Rust (native) pdf-inspector crate High‑performance systems, embedded logic
Python pdf-inspector (PyPI) Data pipelines, ML workflows, scripting
Node.js @firecrawl/pdf-inspector Serverless functions, API services
WebAssembly @firecrawl/pdf-inspector-wasm Client‑side parsing, privacy‑first apps

Each binding wraps the identical engine implemented in src/lib.rs, where key functions like process_pdf, detect_pdf, and extract_text orchestrate PDF loading, type classification, and content extraction.

Core API Functions

The public API in src/lib.rs provides three primary entry points for programmatic use:

  • process_pdf – Complete analysis pipeline returning PDF type, extracted text, markdown output, and layout structure.
  • detect_pdf – Lightweight classification without full extraction; returns PdfClassification (TextBased, Scanned, ImageBased, or Mixed).
  • extract_text – Plain text extraction only, skipping markdown generation for faster processing.

Classification logic lives in src/detector.rs, which applies multiple scan strategies to categorize input files. The extraction pipeline in src/extractor/mod.rs handles content stream parsing, font decoding, and positional text grouping into TextItem and TextLine structures.

Python Integration

Install via pip and use synchronously:

import pdf_inspector

# Full analysis with markdown output

result = pdf_inspector.process_pdf("annual_report.pdf")
print(result.pdf_type)      # "text_based", "scanned", etc.

print(result.markdown)      # Structured markdown string

# Detection-only for routing logic

classification = pdf_inspector.detect_pdf("unknown.pdf")
if classification.is_text_based:
    content = pdf_inspector.extract_text("unknown.pdf")

The Python package includes type stubs (pdf_inspector.pyi) for IDE autocomplete and static checking.

Node.js Integration

The N‑API binding supports both file paths and Buffer objects for serverless environments:

const { processPdf, detectPdf, extractText } = require("@firecrawl/pdf-inspector");

async function analyzeDocument(buffer) {
  const result = await processPdf(buffer);
  return {
    type: result.pdf_type,
    markdown: result.markdown,
    pageCount: result.pages.length
  };
}

See napi/README.md in the repository for async streaming patterns and error handling details.

WebAssembly Browser Usage

For client‑side parsing without server round‑trips:

<script type="module">
import init, { process_pdf } from "@firecrawl/pdf-inspector-wasm";

await init();
const response = await fetch("/document.pdf");
const result = process_pdf(new Uint8Array(await response.arrayBuffer()));

// result.markdown contains the structured output
document.getElementById("output").textContent = result.markdown;
</script>

The WASM build is distributed via CDN and documented in wasm/README.md.

Rust Direct Integration

For systems requiring maximum control or custom extraction pipelines:

use pdf_inspector::{process_pdf, PdfProcessResult};

fn analyze_upload(path: &str) -> Result<String, Box<dyn std::error::Error>> {
    let result: PdfProcessResult = process_pdf(path)?;
    
    match result.pdf_type.as_str() {
        "text_based" => Ok(result.markdown.unwrap_or_default()),
        "scanned" => Err("OCR required for scanned documents".into()),
        _ => Ok(result.plain_text.unwrap_or_default())
    }
}

Direct crate usage lets you configure extractor behavior and handle TextItem collections before markdown conversion in src/markdown/convert.rs.

Architecture Deep Dive

Understanding the internal structure helps optimize programmatic calls:

Component Source File Responsibility
Detector src/detector.rs Multi‑strategy PDF classification
Extractor src/extractor/mod.rs Content parsing, font decoding, layout analysis
Markdown generator src/markdown/convert.rs Token‑efficient markdown with headers, lists, tables
Utilities src/tounicode.rs, src/text_utils.rs Unicode, CJK, RTL text handling

The markdown generator specifically optimizes for LLM context windows by producing clean structural markers rather than visual formatting.

Performance Considerations

For programmatic workflows at scale:

  • Use detect_pdf first to route documents—it's significantly faster than full extraction.
  • Prefer extract_text over process_pdf** when you don't need markdown or layout data.
  • Buffer reuse in Node.js and WASM reduces memory allocations for batch processing.
  • Rust native offers zero‑copy access to internal structures; use it for sub‑millisecond latency requirements.

Summary

  • Firecrawl pdf‑inspector exposes identical core functionality through Rust, Python, Node.js, and WebAssembly interfaces.
  • Entry functions process_pdf, detect_pdf, and extract_text in src/lib.rs provide tiered analysis depth for different use cases.
  • Python and Node.js bindings are thin wrappers; WASM enables browser‑native parsing without server dependencies.
  • The layered architecture (detector → extractor → markdown generator) ensures consistent results across all runtimes.

Frequently Asked Questions

Does pdf‑inspector require a running Firecrawl service?

No. The library operates entirely offline with no external API calls. All processing happens locally using the Rust core engine.

Can I use pdf‑inspector in a serverless environment?

Yes. The Node.js package works in AWS Lambda, Vercel Edge Functions, and Cloudflare Workers. The WASM build runs in Deno and WinterCG‑compatible runtimes.

What PDF types can the detector identify?

The detector in src/detector.rs classifies files as TextBased, Scanned, ImageBased, or Mixed. Scanned detection relies on text density analysis rather than OCR, so it flags documents requiring external OCR services without attempting extraction.

Is the Python binding pure Python or does it require Rust compilation?

The PyPI package distributes pre‑compiled wheels for common platforms. You only need a Rust toolchain if installing from source or targeting an unsupported architecture.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →