# How to Use Firecrawl pdf‑inspector for Programmatic PDF Analysis

> Learn to use Firecrawl pdf-inspector for programmatic PDF analysis. Access native Rust APIs with Python, Node.js, and WebAssembly bindings. Automate your PDF tasks today.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-07

---

**Yes—Firecrawl pdf‑inspector is purpose-built for programmatic PDF analysis, offering native Rust APIs with bindings for Python, Node.js, and WebAssembly.**

This open‑source library exposes a consistent extraction and classification engine across all runtimes, making it ideal for automated workflows in scripts, backend services, or browser applications. The same core logic powers every interface, so your choice of language doesn't compromise analysis quality.

## Programmatic Interfaces Overview

The `firecrawl/pdf-inspector` repository structures its capabilities into layered components that you can access through multiple entry points:

| Interface | Package Name | Best For |
|-----------|--------------|----------|
| **Rust (native)** | `pdf-inspector` crate | High‑performance systems, embedded logic |
| **Python** | `pdf-inspector` (PyPI) | Data pipelines, ML workflows, scripting |
| **Node.js** | `@firecrawl/pdf-inspector` | Serverless functions, API services |
| **WebAssembly** | `@firecrawl/pdf-inspector-wasm` | Client‑side parsing, privacy‑first apps |

Each binding wraps the identical engine implemented in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), where key functions like `process_pdf`, `detect_pdf`, and `extract_text` orchestrate PDF loading, type classification, and content extraction.

## Core API Functions

The public API in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) provides three primary entry points for programmatic use:

- **`process_pdf`** – Complete analysis pipeline returning PDF type, extracted text, markdown output, and layout structure.
- **`detect_pdf`** – Lightweight classification without full extraction; returns `PdfClassification` (TextBased, Scanned, ImageBased, or Mixed).
- **`extract_text`** – Plain text extraction only, skipping markdown generation for faster processing.

Classification logic lives in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs), which applies multiple scan strategies to categorize input files. The extraction pipeline in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) handles content stream parsing, font decoding, and positional text grouping into `TextItem` and `TextLine` structures.

## Python Integration

Install via pip and use synchronously:

```python
import pdf_inspector

# Full analysis with markdown output

result = pdf_inspector.process_pdf("annual_report.pdf")
print(result.pdf_type)      # "text_based", "scanned", etc.

print(result.markdown)      # Structured markdown string

# Detection-only for routing logic

classification = pdf_inspector.detect_pdf("unknown.pdf")
if classification.is_text_based:
    content = pdf_inspector.extract_text("unknown.pdf")

```

The Python package includes type stubs (`pdf_inspector.pyi`) for IDE autocomplete and static checking.

## Node.js Integration

The N‑API binding supports both file paths and `Buffer` objects for serverless environments:

```javascript
const { processPdf, detectPdf, extractText } = require("@firecrawl/pdf-inspector");

async function analyzeDocument(buffer) {
  const result = await processPdf(buffer);
  return {
    type: result.pdf_type,
    markdown: result.markdown,
    pageCount: result.pages.length
  };
}

```

See [`napi/README.md`](https://github.com/firecrawl/pdf-inspector/blob/main/napi/README.md) in the repository for async streaming patterns and error handling details.

## WebAssembly Browser Usage

For client‑side parsing without server round‑trips:

```html
<script type="module">
import init, { process_pdf } from "@firecrawl/pdf-inspector-wasm";

await init();
const response = await fetch("/document.pdf");
const result = process_pdf(new Uint8Array(await response.arrayBuffer()));

// result.markdown contains the structured output
document.getElementById("output").textContent = result.markdown;
</script>

```

The WASM build is distributed via CDN and documented in [`wasm/README.md`](https://github.com/firecrawl/pdf-inspector/blob/main/wasm/README.md).

## Rust Direct Integration

For systems requiring maximum control or custom extraction pipelines:

```rust
use pdf_inspector::{process_pdf, PdfProcessResult};

fn analyze_upload(path: &str) -> Result<String, Box<dyn std::error::Error>> {
    let result: PdfProcessResult = process_pdf(path)?;
    
    match result.pdf_type.as_str() {
        "text_based" => Ok(result.markdown.unwrap_or_default()),
        "scanned" => Err("OCR required for scanned documents".into()),
        _ => Ok(result.plain_text.unwrap_or_default())
    }
}

```

Direct crate usage lets you configure extractor behavior and handle `TextItem` collections before markdown conversion in [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs).

## Architecture Deep Dive

Understanding the internal structure helps optimize programmatic calls:

| Component | Source File | Responsibility |
|-----------|-------------|--------------|
| Detector | [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) | Multi‑strategy PDF classification |
| Extractor | [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) | Content parsing, font decoding, layout analysis |
| Markdown generator | [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) | Token‑efficient markdown with headers, lists, tables |
| Utilities | [`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs), [`src/text_utils.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_utils.rs) | Unicode, CJK, RTL text handling |

The markdown generator specifically optimizes for LLM context windows by producing clean structural markers rather than visual formatting.

## Performance Considerations

For programmatic workflows at scale:

- **Use `detect_pdf`** first to route documents—it's significantly faster than full extraction.
- **Prefer `extract_text`** over `process_pdf`** when you don't need markdown or layout data.
- **Buffer reuse** in Node.js and WASM reduces memory allocations for batch processing.
- **Rust native** offers zero‑copy access to internal structures; use it for sub‑millisecond latency requirements.

## Summary

- Firecrawl pdf‑inspector exposes **identical core functionality** through Rust, Python, Node.js, and WebAssembly interfaces.
- Entry functions `process_pdf`, `detect_pdf`, and `extract_text` in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) provide tiered analysis depth for different use cases.
- Python and Node.js bindings are thin wrappers; WASM enables browser‑native parsing without server dependencies.
- The layered architecture (detector → extractor → markdown generator) ensures consistent results across all runtimes.

## Frequently Asked Questions

### Does pdf‑inspector require a running Firecrawl service?

No. The library operates entirely offline with no external API calls. All processing happens locally using the Rust core engine.

### Can I use pdf‑inspector in a serverless environment?

Yes. The Node.js package works in AWS Lambda, Vercel Edge Functions, and Cloudflare Workers. The WASM build runs in Deno and WinterCG‑compatible runtimes.

### What PDF types can the detector identify?

The detector in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) classifies files as **TextBased**, **Scanned**, **ImageBased**, or **Mixed**. Scanned detection relies on text density analysis rather than OCR, so it flags documents requiring external OCR services without attempting extraction.

### Is the Python binding pure Python or does it require Rust compilation?

The PyPI package distributes pre‑compiled wheels for common platforms. You only need a Rust toolchain if installing from source or targeting an unsupported architecture.