# How to Process PDFs from Memory Buffer vs File Path in pdf-inspector

> Learn how to process PDFs from memory buffer or file path using pdf-inspector. Understand the differences and choose the best method for your needs. Optimize your PDF processing workflow.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-08

---

**The `pdf-inspector` crate exposes two primary entry points—`process_pdf_from_path` for filesystem sources and `process_pdf_from_bytes` for in-memory buffers—both feeding an identical extraction pipeline defined in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs).**

The `firecrawl/pdf-inspector` repository provides a Rust-based PDF extraction engine with Python bindings that converts complex documents into structured Markdown. Whether you are reading files from local storage or handling uploaded byte streams in a web service, understanding how to **process PDFs from memory buffer vs file path** ensures you choose the most efficient integration strategy for your application.

## Processing PDFs from a File Path

When your PDF resides on disk, the library offers high-level convenience functions that handle file I/O internally before invoking the unified extraction logic.

### Rust API: `process_pdf_from_path`

In [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), the public API exposes `process_pdf_from_path`, which accepts a string slice representing the filesystem path and returns a `Result<ExtractedPdf>`. This function delegates to `PdfDocument::load_from_path` before passing the document through the detection and extraction stages.

```rust
use pdf_inspector::process_pdf_from_path;

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let extracted = process_pdf_from_path("documents/report.pdf")?;
    println!("{}", extracted);  // Implements Display for Markdown output
    Ok(())
}

```

### Python API: `process_pdf`

The Python wrapper in [`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs) maps the path-based Rust function to `process_pdf`, accepting a string path and returning a dictionary containing the Markdown representation and metadata.

```python
from pdf_inspector import process_pdf

result = process_pdf("documents/report.pdf")
print(result["markdown"])

```

## Processing PDFs from a Memory Buffer

For serverless environments, network streams, or encrypted storage workflows, processing bytes directly avoids temporary file creation and disk latency.

### Rust API: `process_pdf_from_bytes`

The `process_pdf_from_bytes` function defined in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) takes a `&[u8]` slice and returns `Result<ExtractedPdf>`. According to the source code, this entry point calls `PdfDocument::load_from_memory` to initialize the document structure before processing continues through [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs).

```rust
use pdf_inspector::process_pdf_from_bytes;
use std::fs;

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let buffer = fs::read("documents/report.pdf")?;  // Or any Vec<u8> source
    let extracted = process_pdf_from_bytes(&buffer)?;
    println!("{}", extracted);
    Ok(())
}

```

### Python API: `process_pdf_from_bytes`

The Python binding exposes `process_pdf_from_bytes`, which accepts a `bytes` object. As implemented in [`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs), this function bridges the Rust byte slice API, allowing direct processing of memory buffers from HTTP requests or cloud storage without writing to disk.

```python
from pdf_inspector import process_pdf_from_bytes

with open("documents/report.pdf", "rb") as f:
    pdf_bytes = f.read()

result = process_pdf_from_bytes(pdf_bytes)
print(result["markdown"])

```

## What Happens Under the Hood

Both entry points converge on the same internal architecture. After the initial load step, the library constructs a `PdfDocument` object and executes the following pipeline:

1. **Type Detection** ([`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)): Classifies the PDF as text-based, scanned, or mixed.
2. **Content Extraction** ([`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs)): Extracts text, layout information, and tabular data.
3. **Markdown Conversion** ([`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs)): Formats the extracted structure into clean Markdown.

The only divergence occurs at the ingestion boundary:
- **File path** variants use `PdfDocument::load_from_path` (see [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)).
- **Memory buffer** variants use `PdfDocument::load_from_memory` (see [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)).

## Advanced Configuration Options

For fine-grained control over detection thresholds, page selection, or custom formatters, the crate provides `_with_options` variants for both input methods. These accept configuration structs that propagate through the pipeline without altering the core extraction logic defined in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs).

```rust
use pdf_inspector::{process_pdf_from_path_with_options, ExtractionOptions};

let opts = ExtractionOptions::default();
let result = process_pdf_from_path_with_options("scan.pdf", &opts)?;

```

## Summary

- `process_pdf_from_path` and `process_pdf` (Python) handle on-disk files via `PdfDocument::load_from_path`.
- `process_pdf_from_bytes` handles `&[u8]` in Rust and `bytes` in Python via `PdfDocument::load_from_memory`.
- Both methods feed identical extraction pipelines defined in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs), [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs), and [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs).
- Optional `_with_options` variants exist for both input types to customize extraction behavior.

## Frequently Asked Questions

### Is there a performance difference between processing from disk versus memory?

According to the source implementation in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), the performance difference is negligible for the extraction phase itself because both methods build the same internal `PdfDocument` representation. The memory buffer approach eliminates filesystem I/O overhead, which may improve latency in high-throughput serverless environments, though it requires holding the entire PDF in RAM.

### Can I process partial byte streams or do I need the complete file in memory?

The current API requires the complete PDF as a contiguous byte slice (`&[u8]`). As implemented in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), `process_pdf_from_bytes` expects the full buffer to initialize the `PdfDocument` structure. Streaming or chunked processing is not supported in the public API.

### How does error handling differ between the path and bytes methods?

Both methods return `Result<ExtractedPdf>` in Rust, mapping to Python exceptions via the bindings in [`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs). Path-based calls may raise IO errors during file opening, while byte-based calls fail only during parsing. The downstream extraction errors (detection, layout analysis) are identical regardless of input source.

### Are the optional configuration parameters identical for both input methods?

Yes. The library exposes `process_pdf_from_path_with_options` and `process_pdf_from_bytes_with_options` in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), both accepting the same `ExtractionOptions` struct. This ensures feature parity whether you are processing from a file path or a memory buffer.