# How to Secure Data When Using pdf‑inspector: A Complete Security Guide

> Secure your sensitive data with pdf-inspector. Learn how this local PDF processing tool protects your documents through sandboxing and resource limits. Get the complete security guide now.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: security-guide
- Published: 2026-08-04

---

**`pdf‑inspector` processes PDFs entirely locally without network calls, making it inherently secure for sensitive documents when combined with sandboxing, resource limits, and proper output handling.**

`pdf‑inspector` is a pure‑Rust library (with Python and WASM bindings) that extracts text, layout, and tables from PDF files and converts the result to structured Markdown. The tool works *only* on local PDF bytes — it performs no network calls and stores no data remotely. This architecture makes it well‑suited for handling sensitive documents, but security‑conscious deployments should follow three core pillars: memory safety, denial‑of‑service mitigation, and data isolation.

## Memory Safety: Rust's Core Guarantee

The `pdf‑inspector` core is written in **safe Rust**, as seen in the public API in [[`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs). All parsing leverages the *lopdf* crate, with any unsafe blocks confined to well‑audited helpers such as CMap parsing.

The project's CI enforces strict quality gates:

- `cargo clippy -- -D warnings` runs on every commit
- The full test suite (`cargo test`) executes automatically to catch regressions

This prevents crashes, out‑of‑bounds reads, or undefined behavior triggered by malicious PDFs.

## Denial‑of‑Service Mitigation

Malicious PDFs can attempt to trigger unbounded allocations or infinite loops. `pdf‑inspector` implements several safeguards:

| Defense | Implementation Location | Behavior |
|---------|------------------------|----------|
| **Tiled‑scan detection limits** | [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) | Stops after a configurable pixel threshold |
| **Table width limits** | [`src/tables/grid.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/grid.rs) | Aborts column detection after 25 columns |
| **Cell propagation skip** | [`src/tables/grid.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/grid.rs) | Skips expensive operations for very wide tables |

Users can further restrict resources via **environment variables** or **command‑line flags**:

```bash
pdf2md sensitive.pdf --max-pages 5 --timeout 30

```

## Data Isolation: Keeping Extracted Content Local

The binary at [[`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) reads from a caller‑supplied file path and writes Markdown to a caller‑chosen destination. Critical characteristics:

- **No temporary files** in shared directories
- **In‑memory results** returned as strings via `process_pdf_with_options`
- **No telemetry or outbound connections**

For high‑security environments, sandbox the process using Docker, Firecracker, or OS‑level isolation with read‑only mounts.

## Secure Usage Examples

### Rust Library

```rust
use pdf_inspector::process_pdf_with_options;
use pdf_inspector::ProcessOptions;

fn main() -> Result<(), Box<dyn std::error::Error>> {
    // Limit processing to 5 pages with 30-second timeout
    let opts = ProcessOptions {
        max_pages: Some(5),
        timeout_secs: Some(30),
        ..Default::default()
    };

    // Read from a trusted, read-only source
    let pdf_bytes = std::fs::read("sensitive_report.pdf")?;
    let markdown = process_pdf_with_options(&pdf_bytes, opts)?;
    println!("{}", markdown);
    Ok(())
}

```

The `process_pdf_with_options` function defined in [[`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) accepts explicit limits and returns results without filesystem side effects.

### Python Binding

```python
import pdf_inspector

# Read PDF in read-only binary mode

with open("sensitive_report.pdf", "rb") as f:
    pdf_data = f.read()

# Process with resource constraints

markdown = pdf_inspector.process_pdf(
    pdf_data,
    max_pages=5,
    timeout=30  # seconds

)

print(markdown)

```

Python bindings delegate to the Rust core via [[`napi/src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/napi/src/lib.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/napi/src/lib.rs).

### Docker Sandbox (CLI)

```bash
docker run --rm \
    -v $(pwd):/data:ro \
    --read-only \
    --tmpfs /tmp:noexec,nosuid,size=100m \
    firecrawl/pdf-inspector \
    pdf2md /data/sensitive_report.pdf --max-pages 5 --timeout 30

```

Flags explained:
- `-v $(pwd):/data:ro` — Mount input as **read‑only**
- `--read-only` — Make container filesystem immutable
- `--tmpfs /tmp` — Ephemeral, size‑limited temporary storage

## Security Best‑Practice Checklist

- **Run in a sandbox** — Docker with `--read-only` and `--tmpfs`, or Firecracker microVMs
- **Limit resources** — Use `--max-pages`, `--timeout`, and set `RUST_LOG` to appropriate levels
- **Validate input size** — Reject files exceeding your threshold (CLI rejects PDFs > 2 GB by default)
- **Audit dependencies** — Run `cargo audit` in CI; the project pins `lopdf` and related crates
- **Treat output as untrusted** — Escape Markdown before embedding in HTML or downstream formats

## Key Source Files for Security Audit

| File | Security Relevance |
|------|-------------------|
| [[`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) | Public API with resource options |
| [[`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) | PDF classification and scan‑detection limits |
| [[`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) | Content‑stream parsing orchestration |
| [[`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs) | Table detection with union‑find clustering |
| [[`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) | CLI entry point with flag handling |
| [[`napi/src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/napi/src/lib.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/napi/src/lib.rs) | Binding layer for Python/Node.js |

## Summary

- **`pdf‑inspector` operates entirely offline** — no network calls, no remote storage
- **Safe Rust core** prevents memory corruption from malicious PDFs
- **Built‑in resource limits** guard against DoS via `--max-pages`, `--timeout`, and internal thresholds
- **Sandbox deployment** with read‑only mounts and ephemeral storage provides defense in depth
- **Dependency auditing** (`cargo audit`) and upstream CI checks maintain supply‑chain integrity

## Frequently Asked Questions

### Does pdf‑inspector send PDF data to any external service?

No. According to the `firecrawl/pdf‑inspector` source code, the library processes PDF bytes entirely in‑memory without network calls. The CLI reads from a local path and writes to a caller‑specified output path. All processing happens on‑host.

### How does pdf‑inspector protect against malicious PDFs that might crash the system?

The core is written in safe Rust with unsafe blocks isolated to well‑audited helpers. CI enforces `cargo clippy -- -D warnings` and runs `cargo test` on every commit. The `lopdf` crate handles parsing defensively, and resource limits in `ProcessOptions` prevent runaway execution.

### What resource limits can I set when processing untrusted PDFs?

Set `max_pages` and `timeout_secs` in `ProcessOptions` (Rust) or `max_pages`/`timeout` in Python. The CLI accepts `--max-pages` and `--timeout` flags. Internally, table detection caps columns at 25 and skips expensive propagation for wide tables.

### Is the extracted Markdown safe to render directly in a web application?

Treat all extracted output as untrusted. The Markdown contains text extracted from the PDF without sanitization. Escape or sanitize content before embedding in HTML, JSON, or other contexts where special characters have semantic meaning.