Programming Languages Supported by Firecrawl PDF‑Inspector: A Complete Guide
Firecrawl PDF‑Inspector supports Rust, Python, Node.js/JavaScript, and WebAssembly (browser) through native bindings and packages, plus a standalone CLI tool usable from any environment.
The firecrawl/pdf‑inspector repository provides a high‑performance PDF classification and text‑extraction engine written in Rust. To maximize accessibility across different development ecosystems, the project exposes its core functionality through multiple language bindings while maintaining a single, optimized Rust codebase for all PDF processing operations.
Rust: Native Core Implementation
The Rust crate represents the canonical implementation and primary entry point for all functionality.
In src/lib.rs, the library exposes public APIs such as process_pdf() and classify_pdf() that serve as the foundation for all other language bindings. The core modules handle PDF type detection, content streaming parsing, table extraction, and Markdown conversion.
use pdf_inspector::process_pdf;
fn main() -> Result<(), Box<dyn std::error::Error>> {
let result = process_pdf("document.pdf")?;
println!("Type: {:?}", result.pdf_type);
if let Some(md) = result.markdown {
println!("{}", md);
}
Ok(())
}
Add the dependency to Cargo.toml:
[dependencies]
pdf-inspector = "0.1.0"
Python: PyO3 Bindings with Maturin
Python developers access firecrawl PDF‑inspector through PyO3 bindings compiled using maturin. The resulting package installs via pip and exposes idiomatic Python interfaces to the underlying Rust engine.
import pdf_inspector
result = pdf_inspector.process_pdf("document.pdf")
print(result.pdf_type) # e.g., "text_based"
print(result.markdown) # Markdown string or None
The Python binding source and build instructions live in docs/python.md. This approach gives Python users the same extraction speed as native Rust without requiring any Rust toolchain knowledge.
Node.js and JavaScript: N‑API Bindings
For Node.js environments, the project provides N‑API bindings generated with napi‑rs. The package @firecrawl/pdf-inspector is published to npm and supports both CommonJS and ES module workflows.
import { readFileSync } from "fs";
import { processPdf } from "@firecrawl/pdf-inspector";
const pdf = readFileSync("document.pdf");
const result = processPdf(pdf);
console.log(result.pdfType); // "TextBased", "Scanned", etc.
console.log(result.markdown);
The Node.js wrapper implementation and documentation reside in napi/README.md. N‑API ensures compatibility across Node.js versions without requiring recompilation for each release.
Browser WebAssembly: wasm‑bindgen Module
WebAssembly (WASM) support enables firecrawl PDF‑inspector to run directly in browsers and Web Workers. The @firecrawl/pdf-inspector-wasm npm package uses wasm‑bindgen to bridge JavaScript and the compiled Rust module.
<script type="module">
import init, { processPdf } from "@firecrawl/pdf-inspector-wasm";
async function run() {
await init(); // loads the WASM module
const response = await fetch("document.pdf");
const array = new Uint8Array(await response.arrayBuffer());
const result = processPdf(array);
console.log(result.pdfType);
console.log(result.markdown);
}
run();
</script>
Documentation for browser integration appears in wasm/README.md. The WASM build targets the same src/lib.rs API surface, ensuring feature parity with native deployments.
Command‑Line Interface: Language‑Agnostic Access
Beyond library bindings, firecrawl PDF‑inspector ships standalone CLI tools built directly from the Rust source. These executables—pdf2md and detect-pdf—provide shell‑accessible PDF processing without any language runtime requirements.
The CLI implementation lives in src/bin/pdf2md.rs. This makes the functionality available from any shell environment regardless of the caller's primary programming language.
Architecture: One Core, Multiple Bindings
The multi‑language support in firecrawl PDF‑inspector follows a unified architecture:
PDF bytes
├─► detector → PdfType (TextBased / Scanned / ImageBased / Mixed)
└─► extractor
├─ fonts
├─ content_stream → TextItems + PdfRects
├─ xobjects
├─ links
└─ layout → column detection → reading order
├─► tables → rectangle‑based & heuristic detection
└─► markdown → analysis → preprocess → convert → classify → postprocess
All language bindings call into the same Rust public APIs (process_pdf, classify_pdf, etc.) and marshal results into idiomatic structures. This design guarantees consistent extraction quality and identical performance characteristics across every supported platform.
Summary
- Rust – Native crate
pdf-inspectorwith direct API access viasrc/lib.rs - Python – PyO3 bindings installable through
pip, documented indocs/python.md - Node.js/JavaScript – N‑API bindings as
@firecrawl/pdf-inspector, sourced fromnapi/README.md - WebAssembly – Browser module
@firecrawl/pdf-inspector-wasm, built with wasm‑bindgen perwasm/README.md - CLI – Standalone
pdf2mdanddetect-pdfbinaries fromsrc/bin/pdf2md.rs
Every binding routes through the core Rust implementation, eliminating duplication while maximizing language ecosystem compatibility.
Frequently Asked Questions
How do I install firecrawl PDF‑inspector for Python?
Install via pip using the maturin‑built wheel: pip install pdf-inspector. The PyO3 bindings automatically handle the Rust runtime, so no separate Rust installation is required. Consult docs/python.md for build‑from‑source instructions if you need custom compilation flags.
Can I use firecrawl PDF‑inspector in a React or Vue frontend application?
Yes. Import @firecrawl/pdf-inspector-wasm and call init() to load the module before processing PDFs. The WASM build runs entirely in the browser without backend dependencies, making it suitable for client‑side PDF analysis workflows.
Is the Node.js N‑API binding compatible with all Node.js versions?
The napi‑rs toolchain targets N‑API v6+, which covers Node.js 14.x and later. Check napi/README.md for specific version matrices and prebuilt binary availability for your platform.
What PDF extraction features are available across all language bindings?
All bindings expose the complete feature set: PDF type detection (text‑based, scanned, image‑based, mixed), text extraction with layout preservation, table detection using rectangle‑based heuristics, link extraction, and Markdown conversion with reading‑order reconstruction. Core modules like src/detector.rs, src/extractor/content_stream.rs, and src/markdown/convert.rs power these capabilities universally.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →