How to Secure Data When Using pdf‑inspector: A Complete Security Guide
pdf‑inspector processes PDFs entirely locally without network calls, making it inherently secure for sensitive documents when combined with sandboxing, resource limits, and proper output handling.
pdf‑inspector is a pure‑Rust library (with Python and WASM bindings) that extracts text, layout, and tables from PDF files and converts the result to structured Markdown. The tool works only on local PDF bytes — it performs no network calls and stores no data remotely. This architecture makes it well‑suited for handling sensitive documents, but security‑conscious deployments should follow three core pillars: memory safety, denial‑of‑service mitigation, and data isolation.
Memory Safety: Rust's Core Guarantee
The pdf‑inspector core is written in safe Rust, as seen in the public API in [src/lib.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs). All parsing leverages the lopdf crate, with any unsafe blocks confined to well‑audited helpers such as CMap parsing.
The project's CI enforces strict quality gates:
cargo clippy -- -D warningsruns on every commit- The full test suite (
cargo test) executes automatically to catch regressions
This prevents crashes, out‑of‑bounds reads, or undefined behavior triggered by malicious PDFs.
Denial‑of‑Service Mitigation
Malicious PDFs can attempt to trigger unbounded allocations or infinite loops. pdf‑inspector implements several safeguards:
| Defense | Implementation Location | Behavior |
|---|---|---|
| Tiled‑scan detection limits | src/detector.rs |
Stops after a configurable pixel threshold |
| Table width limits | src/tables/grid.rs |
Aborts column detection after 25 columns |
| Cell propagation skip | src/tables/grid.rs |
Skips expensive operations for very wide tables |
Users can further restrict resources via environment variables or command‑line flags:
pdf2md sensitive.pdf --max-pages 5 --timeout 30
Data Isolation: Keeping Extracted Content Local
The binary at [src/bin/pdf2md.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) reads from a caller‑supplied file path and writes Markdown to a caller‑chosen destination. Critical characteristics:
- No temporary files in shared directories
- In‑memory results returned as strings via
process_pdf_with_options - No telemetry or outbound connections
For high‑security environments, sandbox the process using Docker, Firecracker, or OS‑level isolation with read‑only mounts.
Secure Usage Examples
Rust Library
use pdf_inspector::process_pdf_with_options;
use pdf_inspector::ProcessOptions;
fn main() -> Result<(), Box<dyn std::error::Error>> {
// Limit processing to 5 pages with 30-second timeout
let opts = ProcessOptions {
max_pages: Some(5),
timeout_secs: Some(30),
..Default::default()
};
// Read from a trusted, read-only source
let pdf_bytes = std::fs::read("sensitive_report.pdf")?;
let markdown = process_pdf_with_options(&pdf_bytes, opts)?;
println!("{}", markdown);
Ok(())
}
The process_pdf_with_options function defined in [src/lib.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) accepts explicit limits and returns results without filesystem side effects.
Python Binding
import pdf_inspector
# Read PDF in read-only binary mode
with open("sensitive_report.pdf", "rb") as f:
pdf_data = f.read()
# Process with resource constraints
markdown = pdf_inspector.process_pdf(
pdf_data,
max_pages=5,
timeout=30 # seconds
)
print(markdown)
Python bindings delegate to the Rust core via [napi/src/lib.rs](https://github.com/firecrawl/pdf-inspector/blob/main/napi/src/lib.rs).
Docker Sandbox (CLI)
docker run --rm \
-v $(pwd):/data:ro \
--read-only \
--tmpfs /tmp:noexec,nosuid,size=100m \
firecrawl/pdf-inspector \
pdf2md /data/sensitive_report.pdf --max-pages 5 --timeout 30
Flags explained:
-v $(pwd):/data:ro— Mount input as read‑only--read-only— Make container filesystem immutable--tmpfs /tmp— Ephemeral, size‑limited temporary storage
Security Best‑Practice Checklist
- Run in a sandbox — Docker with
--read-onlyand--tmpfs, or Firecracker microVMs - Limit resources — Use
--max-pages,--timeout, and setRUST_LOGto appropriate levels - Validate input size — Reject files exceeding your threshold (CLI rejects PDFs > 2 GB by default)
- Audit dependencies — Run
cargo auditin CI; the project pinslopdfand related crates - Treat output as untrusted — Escape Markdown before embedding in HTML or downstream formats
Key Source Files for Security Audit
| File | Security Relevance |
|---|---|
[src/lib.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) |
Public API with resource options |
[src/detector.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) |
PDF classification and scan‑detection limits |
[src/extractor/mod.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) |
Content‑stream parsing orchestration |
[src/tables/detect_rects.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs) |
Table detection with union‑find clustering |
[src/bin/pdf2md.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) |
CLI entry point with flag handling |
[napi/src/lib.rs](https://github.com/firecrawl/pdf-inspector/blob/main/napi/src/lib.rs) |
Binding layer for Python/Node.js |
Summary
pdf‑inspectoroperates entirely offline — no network calls, no remote storage- Safe Rust core prevents memory corruption from malicious PDFs
- Built‑in resource limits guard against DoS via
--max-pages,--timeout, and internal thresholds - Sandbox deployment with read‑only mounts and ephemeral storage provides defense in depth
- Dependency auditing (
cargo audit) and upstream CI checks maintain supply‑chain integrity
Frequently Asked Questions
Does pdf‑inspector send PDF data to any external service?
No. According to the firecrawl/pdf‑inspector source code, the library processes PDF bytes entirely in‑memory without network calls. The CLI reads from a local path and writes to a caller‑specified output path. All processing happens on‑host.
How does pdf‑inspector protect against malicious PDFs that might crash the system?
The core is written in safe Rust with unsafe blocks isolated to well‑audited helpers. CI enforces cargo clippy -- -D warnings and runs cargo test on every commit. The lopdf crate handles parsing defensively, and resource limits in ProcessOptions prevent runaway execution.
What resource limits can I set when processing untrusted PDFs?
Set max_pages and timeout_secs in ProcessOptions (Rust) or max_pages/timeout in Python. The CLI accepts --max-pages and --timeout flags. Internally, table detection caps columns at 25 and skips expensive propagation for wide tables.
Is the extracted Markdown safe to render directly in a web application?
Treat all extracted output as untrusted. The Markdown contains text extracted from the PDF without sanitization. Escape or sanitize content before embedding in HTML, JSON, or other contexts where special characters have semantic meaning.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →