# How to Process PDFs from a Memory Buffer Using `process_pdf_mem` in pdf-inspector

> Learn to process PDFs from memory buffers with pdf_inspector's process_pdf_mem. Get structured detection and extraction results instantly without file I/O. Optimize your PDF handling.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-06

---

**Use `pdf_inspector::process_pdf_mem` to process PDFs directly from a `&[u8]` byte slice without writing to disk, returning structured detection and extraction results in a single pass.**

The `pdf-inspector` Rust crate provides a high-performance, filesystem-free API for server-side PDF processing. The `process_pdf_mem` function in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) enables direct processing of PDF data loaded from HTTP responses, databases, or other in-memory sources—eliminating I/O overhead and temporary file management.

## The `process_pdf_mem` Function

The primary entry point for in-memory PDF processing is `process_pdf_mem`, defined at **lines 97–100 of [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)**:

```rust
pub fn process_pdf_mem(buffer: &[u8]) -> Result<PdfProcessResult, PdfError> {
    process_pdf_mem_with_options(buffer, PdfOptions::new())
}

```

This function accepts a **byte slice** (`&[u8]`) containing the complete PDF payload and returns a `PdfProcessResult` with:

- **PDF type detection** (scanned vs. native, structured vs. unstructured)
- **Content extraction** with layout analysis
- **Markdown generation** (unless running in analysis-only mode)
- **OCR requirements** for pages needing text recognition

### Internal Pipeline (Single-Pass Architecture)

According to the `pdf-inspector` source code, `process_pdf_mem` executes four internal steps without re-parsing the document:

1. **`validate_pdf_bytes`** — Validates the byte slice header and structure
2. **`load_document_from_mem_with_password`** — Loads the PDF into a `lopdf::Document`
3. **Shared document reference** — Reuses the parsed document for both detection and extraction
4. **`process_document`** — Drives detection, OCR routing, layout analysis, and Markdown rendering

Because the buffer is processed **only once**, this API minimizes latency for high-throughput server pipelines.

## Basic Usage: Processing PDF Bytes

The simplest approach loads PDF data into a `Vec<u8>` and passes it directly to `process_pdf_mem`:

```rust
use pdf_inspector::process_pdf_mem;

fn main() -> Result<(), pdf_inspector::PdfError> {
    // pdf_bytes sourced from HTTP response, database, or memory-mapped file
    let pdf_bytes: Vec<u8> = fetch_pdf_from_api().await?;

    let result = process_pdf_mem(&pdf_bytes)?;

    println!("Detected PDF type: {:?}", result.pdf_type);
    
    if let Some(markdown) = result.markdown {
        println!("Extracted content:\n{}", markdown);
    }

    Ok(())
}

```

The `&[u8]` parameter accepts any owned or borrowed byte container, including `Vec<u8>`, `&[u8]` slices, or `bytes::Bytes` via `as_ref()`.

## Advanced Usage: `process_pdf_mem_with_options`

For custom behavior—password decryption, page filtering, or analysis-only mode—use **`process_pdf_mem_with_options`**:

```rust
use pdf_inspector::{process_pdf_mem_with_options, PdfOptions, ProcessMode};

fn main() -> Result<(), pdf_inspector::PdfError> {
    let pdf_bytes: Vec<u8> = std::fs::read("encrypted.pdf")?; // or any source

    let options = PdfOptions::new()
        .mode(ProcessMode::Analyze)        // Detection only; skip Markdown
        .password(Some("secret".into()))   // Decrypt password-protected PDF
        .pages([1, 2, 3]);                 // 1-indexed page range filter

    let result = process_pdf_mem_with_options(&pdf_bytes, options)?;

    println!("Pages requiring OCR: {:?}", result.pages_needing_ocr);
    println!("Detected structure: {:?}", result.pdf_type);

    // No markdown field populated in ProcessMode::Analyze
    assert!(result.markdown.is_none());

    Ok(())
}

```

### `ProcessMode` Variants

| Variant | Behavior | Use Case |
|---------|----------|----------|
| `ProcessMode::Full` | Detection + extraction + Markdown | Complete content processing (default) |
| `ProcessMode::Analyze` | Detection only | Quick classification before expensive OCR |
| `ProcessMode::Extract` | Extraction without Markdown | Structured data extraction |

These variants are defined in **[`src/process_mode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/process_mode.rs)** and control pipeline depth without changing the function signature.

## Key Types and Return Values

The **`PdfProcessResult`** struct (from **[`src/types.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs)**) contains:

```rust
pub struct PdfProcessResult {
    pub pdf_type: PdfType,           // Detection classification
    pub pages_needing_ocr: Vec<u32>, // 1-indexed pages without extractable text
    pub markdown: Option<String>,    // Generated Markdown (Full mode)
    pub metadata: Option<PdfMetadata>, // Document properties
    pub error: Option<String>,       // Non-fatal processing issues
}

```

Access fields directly or pattern-match on `pdf_type` to branch processing logic:

```rust
match result.pdf_type {
    PdfType::NativeStructured => {
        // Direct text extraction available
        println!("{}", result.markdown.unwrap_or_default());
    }
    PdfType::ScannedImage => {
        // Route to OCR pipeline
        queue_for_ocr(result.pages_needing_ocr);
    }
    _ => {}
}

```

## Memory Buffer Sources

The `&[u8]` interface integrates cleanly with common Rust async patterns:

- **HTTP clients**: `reqwest::Response::bytes().await?`
- **Databases**: `sqlx::query_scalar::<Vec<u8>>()` or `tokio_postgres` binary fields
- **Message queues**: Kafka/Redis payloads as `Bytes`
- **Memory-mapped files**: `memmap2::Mmap` dereferenced to `&[u8]`

All sources avoid the filesystem entirely while leveraging `pdf-inspector`'s single-pass parsing.

## Summary

- **`process_pdf_mem`** in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) provides filesystem-free PDF processing from any `&[u8]` source
- The single-pass architecture parses the PDF once, then shares the `lopdf::Document` between detection and extraction
- Use **`process_pdf_mem_with_options`** for password decryption, page filtering, or `ProcessMode::Analyze` to skip Markdown generation
- The **`PdfProcessResult`** struct in [`src/types.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs) exposes detection results, OCR requirements, and optional Markdown output
- Ideal for server-side pipelines processing PDFs from HTTP APIs, databases, or in-memory caches

## Frequently Asked Questions

### What is the difference between `process_pdf_mem` and `detect_pdf_mem`?

`detect_pdf_mem` performs **only** PDF type detection and returns a lightweight `PdfDetectResult` without extraction or Markdown. Use it when you need rapid classification before deciding whether to run full processing. `process_pdf_mem` always runs the complete pipeline (unless restricted by `ProcessMode::Analyze`).

### Can `process_pdf_mem` handle password-protected PDFs?

Yes. Pass the password through `PdfOptions::password()` when calling `process_pdf_mem_with_options`. The underlying `load_document_from_mem_with_password` function attempts decryption during the initial document load; invalid passwords return `PdfError::InvalidPassword`.

### How do I process only specific pages from a memory buffer?

Chain `.pages([start, end])` or `.pages([1, 3, 5])` onto your `PdfOptions` before calling `process_pdf_mem_with_options`. Page numbers are **1-indexed** and inclusive. Invalid page ranges are clamped to document bounds with a warning logged to `result.error`.

### Is `process_pdf_mem` thread-safe for concurrent processing?

The function itself is **pure and stateless**—it takes `&[u8]` and returns `PdfProcessResult` without global mutable state. You can safely call it from multiple threads or async tasks, provided each call receives its own byte slice. The underlying `lopdf::Document` parsing is not shared across calls.