How to Process PDFs from Memory Buffer vs File Path in pdf-inspector
The pdf-inspector crate exposes two primary entry points—process_pdf_from_path for filesystem sources and process_pdf_from_bytes for in-memory buffers—both feeding an identical extraction pipeline defined in src/extractor/mod.rs.
The firecrawl/pdf-inspector repository provides a Rust-based PDF extraction engine with Python bindings that converts complex documents into structured Markdown. Whether you are reading files from local storage or handling uploaded byte streams in a web service, understanding how to process PDFs from memory buffer vs file path ensures you choose the most efficient integration strategy for your application.
Processing PDFs from a File Path
When your PDF resides on disk, the library offers high-level convenience functions that handle file I/O internally before invoking the unified extraction logic.
Rust API: process_pdf_from_path
In src/lib.rs, the public API exposes process_pdf_from_path, which accepts a string slice representing the filesystem path and returns a Result<ExtractedPdf>. This function delegates to PdfDocument::load_from_path before passing the document through the detection and extraction stages.
use pdf_inspector::process_pdf_from_path;
fn main() -> Result<(), Box<dyn std::error::Error>> {
let extracted = process_pdf_from_path("documents/report.pdf")?;
println!("{}", extracted); // Implements Display for Markdown output
Ok(())
}
Python API: process_pdf
The Python wrapper in src/python.rs maps the path-based Rust function to process_pdf, accepting a string path and returning a dictionary containing the Markdown representation and metadata.
from pdf_inspector import process_pdf
result = process_pdf("documents/report.pdf")
print(result["markdown"])
Processing PDFs from a Memory Buffer
For serverless environments, network streams, or encrypted storage workflows, processing bytes directly avoids temporary file creation and disk latency.
Rust API: process_pdf_from_bytes
The process_pdf_from_bytes function defined in src/lib.rs takes a &[u8] slice and returns Result<ExtractedPdf>. According to the source code, this entry point calls PdfDocument::load_from_memory to initialize the document structure before processing continues through src/extractor/mod.rs.
use pdf_inspector::process_pdf_from_bytes;
use std::fs;
fn main() -> Result<(), Box<dyn std::error::Error>> {
let buffer = fs::read("documents/report.pdf")?; // Or any Vec<u8> source
let extracted = process_pdf_from_bytes(&buffer)?;
println!("{}", extracted);
Ok(())
}
Python API: process_pdf_from_bytes
The Python binding exposes process_pdf_from_bytes, which accepts a bytes object. As implemented in src/python.rs, this function bridges the Rust byte slice API, allowing direct processing of memory buffers from HTTP requests or cloud storage without writing to disk.
from pdf_inspector import process_pdf_from_bytes
with open("documents/report.pdf", "rb") as f:
pdf_bytes = f.read()
result = process_pdf_from_bytes(pdf_bytes)
print(result["markdown"])
What Happens Under the Hood
Both entry points converge on the same internal architecture. After the initial load step, the library constructs a PdfDocument object and executes the following pipeline:
- Type Detection (
src/detector.rs): Classifies the PDF as text-based, scanned, or mixed. - Content Extraction (
src/extractor/mod.rs): Extracts text, layout information, and tabular data. - Markdown Conversion (
src/markdown/convert.rs): Formats the extracted structure into clean Markdown.
The only divergence occurs at the ingestion boundary:
- File path variants use
PdfDocument::load_from_path(seesrc/lib.rs). - Memory buffer variants use
PdfDocument::load_from_memory(seesrc/lib.rs).
Advanced Configuration Options
For fine-grained control over detection thresholds, page selection, or custom formatters, the crate provides _with_options variants for both input methods. These accept configuration structs that propagate through the pipeline without altering the core extraction logic defined in src/extractor/mod.rs.
use pdf_inspector::{process_pdf_from_path_with_options, ExtractionOptions};
let opts = ExtractionOptions::default();
let result = process_pdf_from_path_with_options("scan.pdf", &opts)?;
Summary
process_pdf_from_pathandprocess_pdf(Python) handle on-disk files viaPdfDocument::load_from_path.process_pdf_from_byteshandles&[u8]in Rust andbytesin Python viaPdfDocument::load_from_memory.- Both methods feed identical extraction pipelines defined in
src/extractor/mod.rs,src/detector.rs, andsrc/markdown/convert.rs. - Optional
_with_optionsvariants exist for both input types to customize extraction behavior.
Frequently Asked Questions
Is there a performance difference between processing from disk versus memory?
According to the source implementation in src/lib.rs, the performance difference is negligible for the extraction phase itself because both methods build the same internal PdfDocument representation. The memory buffer approach eliminates filesystem I/O overhead, which may improve latency in high-throughput serverless environments, though it requires holding the entire PDF in RAM.
Can I process partial byte streams or do I need the complete file in memory?
The current API requires the complete PDF as a contiguous byte slice (&[u8]). As implemented in src/lib.rs, process_pdf_from_bytes expects the full buffer to initialize the PdfDocument structure. Streaming or chunked processing is not supported in the public API.
How does error handling differ between the path and bytes methods?
Both methods return Result<ExtractedPdf> in Rust, mapping to Python exceptions via the bindings in src/python.rs. Path-based calls may raise IO errors during file opening, while byte-based calls fail only during parsing. The downstream extraction errors (detection, layout analysis) are identical regardless of input source.
Are the optional configuration parameters identical for both input methods?
Yes. The library exposes process_pdf_from_path_with_options and process_pdf_from_bytes_with_options in src/lib.rs, both accepting the same ExtractionOptions struct. This ensures feature parity whether you are processing from a file path or a memory buffer.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →