How to Install Firecrawl pdf‑inspector Locally: Complete Setup Guide for Rust, Python, Node.js, and CLI
To install Firecrawl pdf‑inspector locally, use cargo install pdf‑inspector for the CLI, pip install maturin && maturin develop for Python, or npm install @firecrawl/pdf‑inspector for Node.js — no external services or ML models required.
Firecrawl pdf‑inspector is a pure‑Rust PDF-to-Markdown library with bindings for Python, Node.js, and WebAssembly, plus ready-to-use CLI binaries. Because the core pdf‑inspector crate depends only on lopdf for PDF parsing, local installation is fast and self‑contained. The library performs three stages — Detector, Extractor, and Markdown Builder — sharing a single in‑memory PDF representation parsed only once.
Install the CLI Binary with Cargo
The fastest way to get started is installing the pre‑built CLI tools: pdf2md and detect‑pdf.
cargo install pdf-inspector
This downloads and compiles src/bin/pdf2md.rs and src/bin/detect_pdf.rs, placing binaries in your Cargo bin directory.
Convert PDFs to Markdown:
pdf2md example.pdf # Markdown to stdout
pdf2md example.pdf --json # Structured JSON output
Classify PDF type without full extraction:
detect-pdf example.pdf
Install Python Bindings with Maturin
Firecrawl pdf‑inspector exposes its Rust core to Python via PyO3 and maturin.
Step 1: Install maturin
pip install maturin
Step 2: Build and install locally
git clone https://github.com/firecrawl/pdf-inspector.git
cd pdf-inspector
maturin develop --release
The --release flag enables optimizations for the src/lib.rs public API.
Step 3: Use in Python
import pdf_inspector
result = pdf_inspector.process_pdf("example.pdf")
print(result.pdf_type) # "text_based", "scanned", "image_based", "mixed"
print(result.markdown) # Clean Markdown string
The process_pdf function orchestrates Detector (from src/detector.rs), Extractor (from src/extractor/content_stream.rs), and Markdown Builder (from src/markdown/convert.rs).
Full Python API reference: [docs/python.md](https://github.com/firecrawl/pdf-inspector/blob/main/docs/python.md)
Install Node.js Bindings with npm
Native Node.js bindings are published as @firecrawl/pdf‑inspector using NAPI‑RS.
npm install @firecrawl/pdf-inspector
Usage example
import { readFileSync } from "fs";
import { processPdf } from "@firecrawl/pdf-inspector";
const pdf = readFileSync("example.pdf");
const result = processPdf(pdf);
console.log(result.pdfType); // "TextBased", "Scanned", etc.
console.log(result.markdown); // Markdown string
Full Node.js API reference: [napi/README.md](https://github.com/firecrawl/pdf-inspector/blob/main/napi/README.md)
Install WebAssembly for Browser Use
For client‑side PDF processing without a backend, use the WASM build:
npm install @firecrawl/pdf-inspector-wasm
Browser integration
import init, { processPdf } from "@firecrawl/pdf-inspector-wasm";
await init();
const resp = await fetch("/example.pdf");
const pdf = new Uint8Array(await resp.arrayBuffer());
const result = processPdf(pdf);
console.log(result.pdfType);
console.log(result.markdown);
The WASM module exposes the same src/lib.rs API compiled to WebAssembly, running entirely in the browser.
Full WASM API reference: [wasm/README.md](https://github.com/firecrawl/pdf-inspector/blob/main/wasm/README.md)
Install as Rust Library with Cargo
Add the crate directly to your Rust project:
cargo add pdf-inspector
Or manually in Cargo.toml:
[dependencies]
pdf-inspector = "0.1"
Rust usage example
use pdf_inspector::process_pdf;
let result = process_pdf("example.pdf")?;
println!("Type: {:?}", result.pdf_type);
if let Some(md) = result.markdown {
println!("{}", md);
}
Full Rust API reference: [docs/rust-api.md](https://github.com/firecrawl/pdf-inspector/blob/main/docs/rust-api.md)
Key Source Files and Architecture
Understanding the source structure helps when customizing or debugging local installations:
| File | Purpose |
|---|---|
src/lib.rs |
Public API, process_pdf entry point, options builder |
src/detector.rs |
Fast PDF‑type detection without full parsing |
src/extractor/content_stream.rs |
State machine walking PDF operators (Tj, TJ, Td, Tm) |
src/tables/detect_rects.rs |
Rectangle‑based table detection via union‑find clustering |
src/markdown/convert.rs |
Core Markdown conversion integrating tables and lists |
src/markdown/classify.rs |
Heuristics for headings, lists, captions, code blocks |
src/tounicode.rs |
ToUnicode CMap parsing for CID‑encoded fonts |
src/fonts.rs |
Font width/encoding handling with TrueType fallbacks |
All stages share the in‑memory representation to eliminate redundant I/O.
System Requirements
- Rust toolchain (1.70+) for any installation method
- Python 3.8+ for Python bindings
- Node.js 16+ for npm packages
- No GPU, no Docker, no external API keys
Summary
- CLI quickest path:
cargo install pdf-inspectorgets youpdf2mdanddetect-pdfimmediately - Python users: Use
maturin develop --releaseafter cloning the repository - Node.js/Web users: Install from npm registry for native or WASM bindings
- Rust developers: Add
pdf-inspectoras a Cargo dependency - All methods rely solely on
lopdf— no OCR engines or cloud services needed
Frequently Asked Questions
Does pdf‑inspector require any external services or API keys?
No. Firecrawl pdf‑inspector is completely self‑contained. The src/detector.rs and src/extractor/content_stream.rs modules perform all processing locally using only the lopdf crate for PDF parsing. No network calls, no cloud OCR, no authentication required.
Can I install pdf‑inspector without installing Rust?
Pre‑built binaries are not officially distributed, so building from source currently requires Rust. However, Python users can install via pip if a maintainer publishes wheels, and Node.js users receive pre‑built native modules from npm. For the CLI, Rust is currently mandatory.
How large are the compiled binaries?
The release binary is typically 3–8 MB depending on platform, with no runtime dependencies. The single‑crate design in src/lib.rs and selective features keep the footprint minimal compared to ML‑based PDF tools.
Is WebAssembly slower than native bindings?
WASM performance is within 1.2–1.5× of native Rust for typical documents, as the core algorithms in src/extractor/content_stream.rs and src/markdown/convert.rs are CPU‑bound and memory‑efficient. Initialization via init() is required once per page load to fetch the .wasm module.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →