How to Process PDFs from a Memory Buffer Using `process_pdf_mem` in pdf-inspector
Use pdf_inspector::process_pdf_mem to process PDFs directly from a &[u8] byte slice without writing to disk, returning structured detection and extraction results in a single pass.
The pdf-inspector Rust crate provides a high-performance, filesystem-free API for server-side PDF processing. The process_pdf_mem function in src/lib.rs enables direct processing of PDF data loaded from HTTP responses, databases, or other in-memory sources—eliminating I/O overhead and temporary file management.
The process_pdf_mem Function
The primary entry point for in-memory PDF processing is process_pdf_mem, defined at lines 97–100 of src/lib.rs:
pub fn process_pdf_mem(buffer: &[u8]) -> Result<PdfProcessResult, PdfError> {
process_pdf_mem_with_options(buffer, PdfOptions::new())
}
This function accepts a byte slice (&[u8]) containing the complete PDF payload and returns a PdfProcessResult with:
- PDF type detection (scanned vs. native, structured vs. unstructured)
- Content extraction with layout analysis
- Markdown generation (unless running in analysis-only mode)
- OCR requirements for pages needing text recognition
Internal Pipeline (Single-Pass Architecture)
According to the pdf-inspector source code, process_pdf_mem executes four internal steps without re-parsing the document:
validate_pdf_bytes— Validates the byte slice header and structureload_document_from_mem_with_password— Loads the PDF into alopdf::Document- Shared document reference — Reuses the parsed document for both detection and extraction
process_document— Drives detection, OCR routing, layout analysis, and Markdown rendering
Because the buffer is processed only once, this API minimizes latency for high-throughput server pipelines.
Basic Usage: Processing PDF Bytes
The simplest approach loads PDF data into a Vec<u8> and passes it directly to process_pdf_mem:
use pdf_inspector::process_pdf_mem;
fn main() -> Result<(), pdf_inspector::PdfError> {
// pdf_bytes sourced from HTTP response, database, or memory-mapped file
let pdf_bytes: Vec<u8> = fetch_pdf_from_api().await?;
let result = process_pdf_mem(&pdf_bytes)?;
println!("Detected PDF type: {:?}", result.pdf_type);
if let Some(markdown) = result.markdown {
println!("Extracted content:\n{}", markdown);
}
Ok(())
}
The &[u8] parameter accepts any owned or borrowed byte container, including Vec<u8>, &[u8] slices, or bytes::Bytes via as_ref().
Advanced Usage: process_pdf_mem_with_options
For custom behavior—password decryption, page filtering, or analysis-only mode—use process_pdf_mem_with_options:
use pdf_inspector::{process_pdf_mem_with_options, PdfOptions, ProcessMode};
fn main() -> Result<(), pdf_inspector::PdfError> {
let pdf_bytes: Vec<u8> = std::fs::read("encrypted.pdf")?; // or any source
let options = PdfOptions::new()
.mode(ProcessMode::Analyze) // Detection only; skip Markdown
.password(Some("secret".into())) // Decrypt password-protected PDF
.pages([1, 2, 3]); // 1-indexed page range filter
let result = process_pdf_mem_with_options(&pdf_bytes, options)?;
println!("Pages requiring OCR: {:?}", result.pages_needing_ocr);
println!("Detected structure: {:?}", result.pdf_type);
// No markdown field populated in ProcessMode::Analyze
assert!(result.markdown.is_none());
Ok(())
}
ProcessMode Variants
| Variant | Behavior | Use Case |
|---|---|---|
ProcessMode::Full |
Detection + extraction + Markdown | Complete content processing (default) |
ProcessMode::Analyze |
Detection only | Quick classification before expensive OCR |
ProcessMode::Extract |
Extraction without Markdown | Structured data extraction |
These variants are defined in src/process_mode.rs and control pipeline depth without changing the function signature.
Key Types and Return Values
The PdfProcessResult struct (from src/types.rs) contains:
pub struct PdfProcessResult {
pub pdf_type: PdfType, // Detection classification
pub pages_needing_ocr: Vec<u32>, // 1-indexed pages without extractable text
pub markdown: Option<String>, // Generated Markdown (Full mode)
pub metadata: Option<PdfMetadata>, // Document properties
pub error: Option<String>, // Non-fatal processing issues
}
Access fields directly or pattern-match on pdf_type to branch processing logic:
match result.pdf_type {
PdfType::NativeStructured => {
// Direct text extraction available
println!("{}", result.markdown.unwrap_or_default());
}
PdfType::ScannedImage => {
// Route to OCR pipeline
queue_for_ocr(result.pages_needing_ocr);
}
_ => {}
}
Memory Buffer Sources
The &[u8] interface integrates cleanly with common Rust async patterns:
- HTTP clients:
reqwest::Response::bytes().await? - Databases:
sqlx::query_scalar::<Vec<u8>>()ortokio_postgresbinary fields - Message queues: Kafka/Redis payloads as
Bytes - Memory-mapped files:
memmap2::Mmapdereferenced to&[u8]
All sources avoid the filesystem entirely while leveraging pdf-inspector's single-pass parsing.
Summary
process_pdf_meminsrc/lib.rsprovides filesystem-free PDF processing from any&[u8]source- The single-pass architecture parses the PDF once, then shares the
lopdf::Documentbetween detection and extraction - Use
process_pdf_mem_with_optionsfor password decryption, page filtering, orProcessMode::Analyzeto skip Markdown generation - The
PdfProcessResultstruct insrc/types.rsexposes detection results, OCR requirements, and optional Markdown output - Ideal for server-side pipelines processing PDFs from HTTP APIs, databases, or in-memory caches
Frequently Asked Questions
What is the difference between process_pdf_mem and detect_pdf_mem?
detect_pdf_mem performs only PDF type detection and returns a lightweight PdfDetectResult without extraction or Markdown. Use it when you need rapid classification before deciding whether to run full processing. process_pdf_mem always runs the complete pipeline (unless restricted by ProcessMode::Analyze).
Can process_pdf_mem handle password-protected PDFs?
Yes. Pass the password through PdfOptions::password() when calling process_pdf_mem_with_options. The underlying load_document_from_mem_with_password function attempts decryption during the initial document load; invalid passwords return PdfError::InvalidPassword.
How do I process only specific pages from a memory buffer?
Chain .pages([start, end]) or .pages([1, 3, 5]) onto your PdfOptions before calling process_pdf_mem_with_options. Page numbers are 1-indexed and inclusive. Invalid page ranges are clamped to document bounds with a warning logged to result.error.
Is process_pdf_mem thread-safe for concurrent processing?
The function itself is pure and stateless—it takes &[u8] and returns PdfProcessResult without global mutable state. You can safely call it from multiple threads or async tasks, provided each call receives its own byte slice. The underlying lopdf::Document parsing is not shared across calls.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →