# Firecrawl PDF-Inspector Dependencies: Complete Guide for Native, WASM, and Python Builds

> Explore Firecrawl PDF-Inspector dependencies for native, WASM, and Python builds. Get a complete guide to setting up your development environment.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: dependencies
- Published: 2026-08-07

---

**Firecrawl PDF-Inspector requires Rust ≥1.70 and declares all dependencies in [`Cargo.toml`](https://github.com/firecrawl/pdf-inspector/blob/main/Cargo.toml), with conditional compilation for native (`x86_64`), WebAssembly (`wasm32`), and optional Python binding targets.**

The firecrawl/pdf-inspector repository is a Rust-based PDF parsing engine designed for high-performance text extraction and Markdown conversion. Understanding its dependency structure is essential for successful compilation across different target platforms. This guide breaks down every crate requirement based on the actual source code in [`Cargo.toml`](https://github.com/firecrawl/pdf-inspector/blob/main/Cargo.toml).

## Core Dependencies Required for All Builds

These crates are included unconditionally, regardless of your target architecture:

- **thiserror 2.0** — Structured error handling throughout the codebase. All public error types in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) derive from this crate.

- **log 0.4** — Logging façade used by extractors, detectors, and font handlers. The actual implementation varies by target.

- **regex 1.10** — Pattern matching for text processing, table detection, and PDF structure parsing.

- **once_cell 1.19** — Lazy static initialization, primarily for caching compiled regexes in [`src/text_utils.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_utils.rs) and [`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs).

- **unicode-normalization 0.1** — NFKC normalization for cleaning extracted text and handling ligature expansion.

- **ttf-parser 0.25** — TrueType font parsing required for CID-to-Unicode cmap extraction in embedded fonts.

- **lopdf 0.42.0** — The foundational PDF parsing library. This crate uses **target-specific feature flags**:
  - `rayon` feature for native builds (parallel page processing)
  - `wasm_js` feature for WebAssembly builds (single-threaded, no filesystem dependence)

## Native Build Dependencies (`!wasm32`)

When compiling for `x86_64` or other native architectures, the following additional crates are enabled via `cfg(not(target_arch = "wasm32"))`:

- **rayon 1.10** — Parallel iterator support that `lopdf` leverages for concurrent page parsing. This significantly speeds up large document processing.

- **env_logger 0.11** — Concrete logging implementation for CLI binaries (`pdf2md` and `detect-pdf`).

Build the native binaries with:

```bash
cargo build --release

```

Run the extracted tools:

```bash

# Convert PDF to Markdown

./target/release/pdf2md my_document.pdf > output.md

# Detect PDF composition type

./target/release/detect-pdf --json my_document.pdf

```

## WebAssembly Build Dependencies (`wasm32`)

For browser or serverless environments, the WASM target excludes threading and filesystem dependencies:

- **include_dir 0.7** — Embeds bundled CMap resources directly into the compiled binary. This eliminates runtime filesystem access since WASM sandboxes cannot read arbitrary paths.

The `lopdf` crate automatically switches to its `wasm_js` feature in this configuration.

## Optional Python Bindings Dependencies

To generate the `pdf_inspector` Python package, enable the `python` feature:

- **pyo3 0.25** — Rust/Python interoperability layer that generates extension modules compatible with Python ≥3.8.

Build with Python support:

```bash
cargo build --release --features python

```

Usage in Python:

```python
import pdf_inspector

# Extract PDF to structured JSON

result = pdf_inspector.pdf2md("my_document.pdf", json=True)
print(result)

```

## Development-Only Dependencies

- **tempfile 3.3** — Used exclusively by the test suite for temporary file handling. Not included in release builds.

## Rust Library Integration

Use PDF-Inspector as a crate dependency in your own Rust projects:

```rust
use pdf_inspector::process_pdf_with_options;

fn main() -> Result<(), pdf_inspector::Error> {
    let opts = pdf_inspector::Options::default();
    let md = process_pdf_with_options("my_document.pdf", opts)?;
    println!("{}", md);
    Ok(())
}

```

## Dependency Configuration in Cargo.toml

The [`Cargo.toml`](https://github.com/firecrawl/pdf-inspector/blob/main/Cargo.toml) file gates dependencies using conditional compilation:

```toml
[dependencies]

# ... core crates ...

[target.'cfg(not(target_arch = "wasm32"))'.dependencies]
rayon = "1.10"
env_logger = "0.11"

[target.'cfg(target_arch = "wasm32")'.dependencies]
include_dir = "0.7"

[features]
python = ["pyo3"]

```

Key source files consuming these dependencies:
- [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) — Public API and error types
- [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) — Orchestrates text, layout, and table extraction
- [`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs) — Font decoding with `ttf-parser` and bundled CMaps
- [`src/text_utils.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_utils.rs) — Unicode operations with `unicode-normalization`

## Prerequisites Summary

| Requirement | Version | Purpose |
|-------------|---------|---------|
| Rust toolchain | ≥1.70 | Compilation |
| Python | ≥3.8 | Optional Python bindings |
| Cargo | Ships with Rust | Dependency resolution and builds |

## Summary

- **Firecrawl PDF-Inspector dependencies** are declared entirely in [`Cargo.toml`](https://github.com/firecrawl/pdf-inspector/blob/main/Cargo.toml) with no external system libraries required
- **Always included**: `thiserror`, `log`, `regex`, `once_cell`, `unicode-normalization`, `ttf-parser`, and `lopdf`
- **Native only**: `rayon` and `env_logger` for parallel processing and CLI logging
- **WASM only**: `include_dir` for embedded CMap resources
- **Python optional**: `pyo3` behind the `python` feature flag
- **Minimum Rust version**: 1.70

## Frequently Asked Questions

### What is the minimum Rust version for firecrawl PDF-Inspector?

Rust 1.70 or newer is required. This version ensures compatibility with the `lopdf 0.42.0` crate and modern workspace resolver features used in the build configuration.

### Do I need Python to compile firecrawl PDF-Inspector?

No. Python ≥3.8 is only required if you enable the `python` feature with `cargo build --features python`. The core library and CLI binaries compile without any Python toolchain.

### Why does the WASM build use different dependencies?

The `wasm32` target cannot access the filesystem or spawn threads. The `include_dir` crate embeds CMap data directly into the binary, while `lopdf` switches to its `wasm_js` feature to disable rayon-based parallelism and filesystem operations.

### Can I use firecrawl PDF-Inspector without installing Rust?

Not for compilation. Pre-built binaries are not officially distributed, so you must compile from source. However, once built, the Python wheel can be distributed and installed on systems without Rust.