How to Install pdf-inspector Locally: Complete Setup Guide for Rust, Python, and Node.js
Install pdf-inspector via cargo install pdf-inspector for Rust CLI tools, pip install pdf-inspector for Python bindings, or npm install @firecrawl/pdf-inspector for Node.js, depending on your target environment.
pdf-inspector is a high-performance Rust library for PDF classification and Markdown extraction, developed by firecrawl/pdf-inspector. It provides convenient bindings for Python, Node.js, and WebAssembly, plus ready-to-use CLI tools. Depending on your language or environment, installation steps differ, but all targets share the same core compiled library defined in src/lib.rs.
Prerequisites for Local Installation
Before installing pdf-inspector locally, ensure you have the appropriate toolchain for your target environment:
- Rust: Requires
rustupandcargo(latest stable) - Python: Requires Python ≥ 3.8 and pip; Rust toolchain only needed for building from source
- Node.js: Requires Node ≥ 14 and npm or yarn
- WebAssembly: Requires npm and Node for bundling
Installing the Rust CLI and Library
The Rust installation provides the pdf2md and detect-pdf binaries plus the pdf-inspector crate for embedding in Rust projects.
To install the pre-compiled binaries from crates.io:
cargo install pdf-inspector
This installs the CLI tools globally. After installation, convert PDFs to Markdown using:
pdf2md my_file.pdf
Or detect PDF classification:
detect-pdf my_file.pdf
To use the library in your Rust project, add the following to your Cargo.toml:
[dependencies]
pdf-inspector = "0.2"
The public API entry point resides in src/lib.rs, exposing functions like process_pdf and detect_pdf. The CLI driver for pdf2md is implemented in src/bin/pdf2md.rs.
Example Rust usage:
use pdf_inspector::process_pdf;
fn main() -> Result<(), pdf_inspector::Error> {
let result = process_pdf("report.pdf")?;
println!("PDF type: {:?}", result.pdf_type);
if let Some(md) = result.markdown {
println!("{}", md);
}
Ok(())
}
Installing Python Bindings
Python users can install pre-built wheels for Linux, macOS, and Windows, which bundle the compiled Rust code with generated pdf_inspector.pyi type stubs.
Install via pip:
pip install pdf-inspector
For custom builds from source, install maturin and build in release mode:
pip install maturin
maturin develop --release
Refer to docs/python.md for detailed Python-specific documentation.
Example Python usage:
import pdf_inspector
result = pdf_inspector.process_pdf("report.pdf")
print(result.pdf_type) # e.g. "text_based"
print(result.markdown[:200]) # preview of generated Markdown
Installing Node.js Bindings
The Node.js package contains a native add-on built with napi-rs, exposing the same API as the Rust library.
Install using npm:
npm install @firecrawl/pdf-inspector
The package includes pre-built binaries for common OS/architecture combinations. On unsupported platforms, the post-install script automatically compiles the native add-on using your local Rust toolchain.
See napi/README.md for Node.js-specific guidance.
Example Node.js usage:
import { readFileSync } from 'fs';
import { processPdf } from '@firecrawl/pdf-inspector';
const pdf = readFileSync('invoice.pdf');
const result = processPdf(pdf);
console.log(result.pdfType); // "TextBased", "Scanned", …
console.log(result.markdown);
Installing WebAssembly for Browser Use
For browser environments, pdf-inspector provides a WebAssembly build that embeds all required CMaps, enabling the parser to run entirely client-side without network requests.
Install via npm:
npm install @firecrawl/pdf-inspector-wasm
Import and initialize the WASM module:
import init, { processPdf } from '@firecrawl/pdf-inspector-wasm';
await init(); // loads the WASM module
const resp = await fetch('paper.pdf');
const array = new Uint8Array(await resp.arrayBuffer());
const out = processPdf(array);
console.log(out.pdfType, out.markdown);
Consult wasm/README.md for complete WebAssembly usage instructions.
Building from Source
To compile pdf-inspector from source for all targets, clone the repository and build with Cargo:
git clone https://github.com/firecrawl/pdf-inspector.git
cd pdf-inspector
cargo build --release
This compiles the core library, CLI binaries (pdf2md and detect-pdf), and all language bindings. Building from source requires the Rust toolchain and optionally Python with maturin for Python bindings.
The source architecture organizes functionality into distinct modules:
src/lib.rs: Public API entry pointsrc/detector.rs: Fast PDF-type classification logicsrc/extractor/: Core text extraction pipeline (fonts, content streams, layout)src/markdown/: Markdown conversion and post-processing
Summary
- Rust users: Run
cargo install pdf-inspectorto get CLI tools and library access viasrc/lib.rs. - Python users: Run
pip install pdf-inspectorfor pre-built wheels, or usematurin develop --releasefor custom builds. - Node.js users: Run
npm install @firecrawl/pdf-inspectorfor native add-ons with automatic fallback compilation. - Browser users: Run
npm install @firecrawl/pdf-inspector-wasmfor standalone WebAssembly execution. - Source builds: Clone from GitHub and run
cargo build --releaseto compile all targets.
Frequently Asked Questions
Do I need Rust installed to use the Python or Node.js packages?
No. The Python wheels and Node.js npm packages include pre-compiled binaries for most common platforms. You only need the Rust toolchain if you are building from source or if you are on an unsupported platform where the Node.js package must compile the native add-on during installation.
What CLI tools are included with the Rust installation?
The cargo install pdf-inspector command installs two binaries: pdf2md for converting PDFs to Markdown, and detect-pdf for classifying PDF types. The pdf2md driver is implemented in src/bin/pdf2md.rs, while the core logic resides in src/lib.rs.
Can I use pdf-inspector in a web browser without a backend server?
Yes. The @firecrawl/pdf-inspector-wasm package provides a WebAssembly build that runs entirely in the browser. It embeds all required CMaps and resources, so it performs PDF classification and extraction without any network requests after the initial WASM load.
How do I verify that pdf-inspector installed correctly?
Run the appropriate command for your installation target: for Rust CLI, execute pdf2md --version or detect-pdf --help; for Python, run import pdf_inspector; print(pdf_inspector.process_pdf) to check the module loads; for Node.js, check that require('@firecrawl/pdf-inspector') or the ES module import resolves without errors.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →