How to Install Firecrawl pdf‑inspector Locally: Complete Setup Guide for Rust, Python, Node.js, and CLI

To install Firecrawl pdf‑inspector locally, use cargo install pdf‑inspector for the CLI, pip install maturin && maturin develop for Python, or npm install @firecrawl/pdf‑inspector for Node.js — no external services or ML models required.

Firecrawl pdf‑inspector is a pure‑Rust PDF-to-Markdown library with bindings for Python, Node.js, and WebAssembly, plus ready-to-use CLI binaries. Because the core pdf‑inspector crate depends only on lopdf for PDF parsing, local installation is fast and self‑contained. The library performs three stages — Detector, Extractor, and Markdown Builder — sharing a single in‑memory PDF representation parsed only once.

Install the CLI Binary with Cargo

The fastest way to get started is installing the pre‑built CLI tools: pdf2md and detect‑pdf.

cargo install pdf-inspector

This downloads and compiles src/bin/pdf2md.rs and src/bin/detect_pdf.rs, placing binaries in your Cargo bin directory.

Convert PDFs to Markdown:

pdf2md example.pdf           # Markdown to stdout

pdf2md example.pdf --json    # Structured JSON output

Classify PDF type without full extraction:

detect-pdf example.pdf

Install Python Bindings with Maturin

Firecrawl pdf‑inspector exposes its Rust core to Python via PyO3 and maturin.

Step 1: Install maturin

pip install maturin

Step 2: Build and install locally

git clone https://github.com/firecrawl/pdf-inspector.git
cd pdf-inspector
maturin develop --release

The --release flag enables optimizations for the src/lib.rs public API.

Step 3: Use in Python

import pdf_inspector

result = pdf_inspector.process_pdf("example.pdf")
print(result.pdf_type)   # "text_based", "scanned", "image_based", "mixed"

print(result.markdown)   # Clean Markdown string

The process_pdf function orchestrates Detector (from src/detector.rs), Extractor (from src/extractor/content_stream.rs), and Markdown Builder (from src/markdown/convert.rs).

Full Python API reference: [docs/python.md](https://github.com/firecrawl/pdf-inspector/blob/main/docs/python.md)

Install Node.js Bindings with npm

Native Node.js bindings are published as @firecrawl/pdf‑inspector using NAPI‑RS.

npm install @firecrawl/pdf-inspector

Usage example

import { readFileSync } from "fs";
import { processPdf } from "@firecrawl/pdf-inspector";

const pdf = readFileSync("example.pdf");
const result = processPdf(pdf);
console.log(result.pdfType);   // "TextBased", "Scanned", etc.
console.log(result.markdown);  // Markdown string

Full Node.js API reference: [napi/README.md](https://github.com/firecrawl/pdf-inspector/blob/main/napi/README.md)

Install WebAssembly for Browser Use

For client‑side PDF processing without a backend, use the WASM build:

npm install @firecrawl/pdf-inspector-wasm

Browser integration

import init, { processPdf } from "@firecrawl/pdf-inspector-wasm";

await init();
const resp = await fetch("/example.pdf");
const pdf = new Uint8Array(await resp.arrayBuffer());
const result = processPdf(pdf);
console.log(result.pdfType);
console.log(result.markdown);

The WASM module exposes the same src/lib.rs API compiled to WebAssembly, running entirely in the browser.

Full WASM API reference: [wasm/README.md](https://github.com/firecrawl/pdf-inspector/blob/main/wasm/README.md)

Install as Rust Library with Cargo

Add the crate directly to your Rust project:

cargo add pdf-inspector

Or manually in Cargo.toml:

[dependencies]
pdf-inspector = "0.1"

Rust usage example

use pdf_inspector::process_pdf;

let result = process_pdf("example.pdf")?;
println!("Type: {:?}", result.pdf_type);
if let Some(md) = result.markdown {
    println!("{}", md);
}

Full Rust API reference: [docs/rust-api.md](https://github.com/firecrawl/pdf-inspector/blob/main/docs/rust-api.md)

Key Source Files and Architecture

Understanding the source structure helps when customizing or debugging local installations:

File Purpose
src/lib.rs Public API, process_pdf entry point, options builder
src/detector.rs Fast PDF‑type detection without full parsing
src/extractor/content_stream.rs State machine walking PDF operators (Tj, TJ, Td, Tm)
src/tables/detect_rects.rs Rectangle‑based table detection via union‑find clustering
src/markdown/convert.rs Core Markdown conversion integrating tables and lists
src/markdown/classify.rs Heuristics for headings, lists, captions, code blocks
src/tounicode.rs ToUnicode CMap parsing for CID‑encoded fonts
src/fonts.rs Font width/encoding handling with TrueType fallbacks

All stages share the in‑memory representation to eliminate redundant I/O.

System Requirements

  • Rust toolchain (1.70+) for any installation method
  • Python 3.8+ for Python bindings
  • Node.js 16+ for npm packages
  • No GPU, no Docker, no external API keys

Summary

  • CLI quickest path: cargo install pdf-inspector gets you pdf2md and detect-pdf immediately
  • Python users: Use maturin develop --release after cloning the repository
  • Node.js/Web users: Install from npm registry for native or WASM bindings
  • Rust developers: Add pdf-inspector as a Cargo dependency
  • All methods rely solely on lopdf — no OCR engines or cloud services needed

Frequently Asked Questions

Does pdf‑inspector require any external services or API keys?

No. Firecrawl pdf‑inspector is completely self‑contained. The src/detector.rs and src/extractor/content_stream.rs modules perform all processing locally using only the lopdf crate for PDF parsing. No network calls, no cloud OCR, no authentication required.

Can I install pdf‑inspector without installing Rust?

Pre‑built binaries are not officially distributed, so building from source currently requires Rust. However, Python users can install via pip if a maintainer publishes wheels, and Node.js users receive pre‑built native modules from npm. For the CLI, Rust is currently mandatory.

How large are the compiled binaries?

The release binary is typically 3–8 MB depending on platform, with no runtime dependencies. The single‑crate design in src/lib.rs and selective features keep the footprint minimal compared to ML‑based PDF tools.

Is WebAssembly slower than native bindings?

WASM performance is within 1.2–1.5× of native Rust for typical documents, as the core algorithms in src/extractor/content_stream.rs and src/markdown/convert.rs are CPU‑bound and memory‑efficient. Initialization via init() is required once per page load to fetch the .wasm module.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →