How to Process PDFs from a Memory Buffer Using `process_pdf_mem` in pdf-inspector

Use pdf_inspector::process_pdf_mem to process PDFs directly from a &[u8] byte slice without writing to disk, returning structured detection and extraction results in a single pass.

The pdf-inspector Rust crate provides a high-performance, filesystem-free API for server-side PDF processing. The process_pdf_mem function in src/lib.rs enables direct processing of PDF data loaded from HTTP responses, databases, or other in-memory sources—eliminating I/O overhead and temporary file management.

The process_pdf_mem Function

The primary entry point for in-memory PDF processing is process_pdf_mem, defined at lines 97–100 of src/lib.rs:

pub fn process_pdf_mem(buffer: &[u8]) -> Result<PdfProcessResult, PdfError> {
    process_pdf_mem_with_options(buffer, PdfOptions::new())
}

This function accepts a byte slice (&[u8]) containing the complete PDF payload and returns a PdfProcessResult with:

  • PDF type detection (scanned vs. native, structured vs. unstructured)
  • Content extraction with layout analysis
  • Markdown generation (unless running in analysis-only mode)
  • OCR requirements for pages needing text recognition

Internal Pipeline (Single-Pass Architecture)

According to the pdf-inspector source code, process_pdf_mem executes four internal steps without re-parsing the document:

  1. validate_pdf_bytes — Validates the byte slice header and structure
  2. load_document_from_mem_with_password — Loads the PDF into a lopdf::Document
  3. Shared document reference — Reuses the parsed document for both detection and extraction
  4. process_document — Drives detection, OCR routing, layout analysis, and Markdown rendering

Because the buffer is processed only once, this API minimizes latency for high-throughput server pipelines.

Basic Usage: Processing PDF Bytes

The simplest approach loads PDF data into a Vec<u8> and passes it directly to process_pdf_mem:

use pdf_inspector::process_pdf_mem;

fn main() -> Result<(), pdf_inspector::PdfError> {
    // pdf_bytes sourced from HTTP response, database, or memory-mapped file
    let pdf_bytes: Vec<u8> = fetch_pdf_from_api().await?;

    let result = process_pdf_mem(&pdf_bytes)?;

    println!("Detected PDF type: {:?}", result.pdf_type);
    
    if let Some(markdown) = result.markdown {
        println!("Extracted content:\n{}", markdown);
    }

    Ok(())
}

The &[u8] parameter accepts any owned or borrowed byte container, including Vec<u8>, &[u8] slices, or bytes::Bytes via as_ref().

Advanced Usage: process_pdf_mem_with_options

For custom behavior—password decryption, page filtering, or analysis-only mode—use process_pdf_mem_with_options:

use pdf_inspector::{process_pdf_mem_with_options, PdfOptions, ProcessMode};

fn main() -> Result<(), pdf_inspector::PdfError> {
    let pdf_bytes: Vec<u8> = std::fs::read("encrypted.pdf")?; // or any source

    let options = PdfOptions::new()
        .mode(ProcessMode::Analyze)        // Detection only; skip Markdown
        .password(Some("secret".into()))   // Decrypt password-protected PDF
        .pages([1, 2, 3]);                 // 1-indexed page range filter

    let result = process_pdf_mem_with_options(&pdf_bytes, options)?;

    println!("Pages requiring OCR: {:?}", result.pages_needing_ocr);
    println!("Detected structure: {:?}", result.pdf_type);

    // No markdown field populated in ProcessMode::Analyze
    assert!(result.markdown.is_none());

    Ok(())
}

ProcessMode Variants

Variant Behavior Use Case
ProcessMode::Full Detection + extraction + Markdown Complete content processing (default)
ProcessMode::Analyze Detection only Quick classification before expensive OCR
ProcessMode::Extract Extraction without Markdown Structured data extraction

These variants are defined in src/process_mode.rs and control pipeline depth without changing the function signature.

Key Types and Return Values

The PdfProcessResult struct (from src/types.rs) contains:

pub struct PdfProcessResult {
    pub pdf_type: PdfType,           // Detection classification
    pub pages_needing_ocr: Vec<u32>, // 1-indexed pages without extractable text
    pub markdown: Option<String>,    // Generated Markdown (Full mode)
    pub metadata: Option<PdfMetadata>, // Document properties
    pub error: Option<String>,       // Non-fatal processing issues
}

Access fields directly or pattern-match on pdf_type to branch processing logic:

match result.pdf_type {
    PdfType::NativeStructured => {
        // Direct text extraction available
        println!("{}", result.markdown.unwrap_or_default());
    }
    PdfType::ScannedImage => {
        // Route to OCR pipeline
        queue_for_ocr(result.pages_needing_ocr);
    }
    _ => {}
}

Memory Buffer Sources

The &[u8] interface integrates cleanly with common Rust async patterns:

  • HTTP clients: reqwest::Response::bytes().await?
  • Databases: sqlx::query_scalar::<Vec<u8>>() or tokio_postgres binary fields
  • Message queues: Kafka/Redis payloads as Bytes
  • Memory-mapped files: memmap2::Mmap dereferenced to &[u8]

All sources avoid the filesystem entirely while leveraging pdf-inspector's single-pass parsing.

Summary

  • process_pdf_mem in src/lib.rs provides filesystem-free PDF processing from any &[u8] source
  • The single-pass architecture parses the PDF once, then shares the lopdf::Document between detection and extraction
  • Use process_pdf_mem_with_options for password decryption, page filtering, or ProcessMode::Analyze to skip Markdown generation
  • The PdfProcessResult struct in src/types.rs exposes detection results, OCR requirements, and optional Markdown output
  • Ideal for server-side pipelines processing PDFs from HTTP APIs, databases, or in-memory caches

Frequently Asked Questions

What is the difference between process_pdf_mem and detect_pdf_mem?

detect_pdf_mem performs only PDF type detection and returns a lightweight PdfDetectResult without extraction or Markdown. Use it when you need rapid classification before deciding whether to run full processing. process_pdf_mem always runs the complete pipeline (unless restricted by ProcessMode::Analyze).

Can process_pdf_mem handle password-protected PDFs?

Yes. Pass the password through PdfOptions::password() when calling process_pdf_mem_with_options. The underlying load_document_from_mem_with_password function attempts decryption during the initial document load; invalid passwords return PdfError::InvalidPassword.

How do I process only specific pages from a memory buffer?

Chain .pages([start, end]) or .pages([1, 3, 5]) onto your PdfOptions before calling process_pdf_mem_with_options. Page numbers are 1-indexed and inclusive. Invalid page ranges are clamped to document bounds with a warning logged to result.error.

Is process_pdf_mem thread-safe for concurrent processing?

The function itself is pure and stateless—it takes &[u8] and returns PdfProcessResult without global mutable state. You can safely call it from multiple threads or async tasks, provided each call receives its own byte slice. The underlying lopdf::Document parsing is not shared across calls.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →