Performance Characteristics of pdf-inspector: How to Optimize for Large PDFs
pdf-inspector processes text-based PDFs in approximately 2.35 milliseconds per document by leveraging single-pass parsing, lightweight stream sampling for classification, and a three-tiered detection system that avoids expensive OCR and ML models unless absolutely necessary.
The firecrawl/pdf-inspector library is engineered for speed-first PDF processing in Rust, delivering sub-50-millisecond classification and sub-3-millisecond extraction times for typical documents. By avoiding heavyweight ML frameworks and redundant I/O operations, pdf-inspector maintains a minimal resource footprint while handling large-scale PDF workloads. Understanding these performance characteristics enables developers to maximize throughput when processing high volumes of documents or multi-megabyte files.
Single-Pass Document Loading and Memory Efficiency
The foundation of pdf-inspector's speed lies in its single-pass document loading architecture. In src/lib.rs, the public API functions load_document_from_path and load_document_from_mem parse the PDF once and share the same internal representation between the classifier and extractor. This eliminates duplicate parsing work and redundant I/O operations that plague traditional multi-stage pipelines.
Memory efficiency is achieved through lightweight data structures defined in src/types.rs. The library stores text items as compact TextItem structs containing only essential coordinates, font index, and optional underline flags. Unlike rendering-based PDF libraries, pdf-inspector never allocates large intermediate canvases, keeping memory usage proportional to document complexity rather than page dimensions.
The dependency footprint remains minimal with only lopdf as an external crate, as specified in Cargo.toml. This low-level parsing library avoids pulling in heavyweight frameworks, resulting in small binary sizes and fast startup times critical for serverless deployments.
Fast Classification via Stream Sampling
PDF type detection in src/detector.rs uses stream sampling to classify documents in approximately 10-50 milliseconds. Rather than analyzing every byte, the detector examines a small subset of content streams to determine whether a PDF is text-based, scanned, image-based, or mixed. The classification result is cached, allowing the extractor to skip OCR processing entirely when it isn't needed.
This approach ensures that text-based PDFs bypass expensive OCR pipelines, while scanned documents are flagged for alternative processing. The short-circuit logic prevents wasted computation on documents that don't require computer vision techniques.
Intelligent Layout and Table Detection
For complex layouts, pdf-inspector implements algorithms optimized for computational efficiency. Column detection in src/extractor/layout.rs builds horizontal projection histograms and identifies valleys to separate newspaper-style multi-column content. This algorithm runs in O(N) time over the page's text items, significantly faster than expensive layout engines that render entire pages.
Table detection follows a three-tiered strategy implemented across three source files:
src/tables/detect_rects.rs– Rectangle-based detection using drawing operations and fast union-find data structuressrc/tables/detect_heuristic.rs– Heuristic alignment detection for structured gridssrc/tables/detect_lines.rs– Line-grid detection only when the first two methods fail
This cascading approach short-circuits expensive processing for the majority of pages, applying computationally intensive line detection only when simpler methods prove insufficient.
Benchmark Results and Throughput
According to the benchmark methodology documented in docs/benchmarking.md, pdf-inspector achieves exceptional performance on modern hardware. On an Apple M4 Pro, processing 200 PDFs without OCR takes 0.470 seconds total, equating to approximately 2.35 milliseconds per document with consistent median speeds across five full-corpus runs.
The practical performance profile includes:
- Classification latency: Under 50 milliseconds per document
- Extraction latency: Under 3 milliseconds for typical text-based PDFs
- Throughput: Approximately 400 PDFs per second on a modern laptop in single-threaded execution
- Scalability: Linear scaling with CPU cores; multi-process pipelines can process thousands of PDFs per second due to the library's pure Rust implementation with no external service dependencies
Optimization Strategies for Large PDFs
To maximize performance when processing large PDFs or high-volume workloads, implement these targeted optimizations:
Avoid Unnecessary OCR
Invoke the detect-pdf functionality first and only run OCR engines on pages explicitly flagged as Scanned. Since pdf-inspector accurately identifies text-based content through stream sampling, you prevent the 100-1000x performance penalty associated with unnecessary optical character recognition.
Reuse PdfDocument Instances
When you need both classification and extraction, call process_pdf once rather than invoking separate operations. The internal document representation is reused automatically, eliminating redundant parsing overhead for large files.
Profile with Structured Logging
Enable targeted debugging through the RUST_LOG environment variable. Configure pdf_inspector::extractor::layout=debug or pdf_inspector=debug to emit fine-grained diagnostics for specific components. This allows precise profiling of runtime bottlenecks without code recompilation or instrumentation overhead, as documented in docs/debugging.md.
Measuring Performance in Practice
Monitor processing times using these implementation patterns:
Command Line Timing:
time pdf2md sample.pdf > /dev/null
Rust Explicit Timing:
use std::time::Instant;
use pdf_inspector::process_pdf;
fn main() -> anyhow::Result<()> {
let start = Instant::now();
let result = process_pdf("sample.pdf")?;
let elapsed = start.elapsed();
println!("PDF type: {:?}", result.pdf_type);
println!("Markdown length: {}", result.markdown.as_ref().map(|s| s.len()).unwrap_or(0));
println!("Processing time: {:.2?}", elapsed);
Ok(())
}
Python Batch Processing:
import time
import pdf_inspector
def process_batch(paths):
start = time.time()
for p in paths:
r = pdf_inspector.process_pdf(p)
# Access r.pdf_type and r.markdown as needed
print(f"Processed {len(paths)} PDFs in {time.time() - start:.2f}s")
process_batch(["doc1.pdf", "doc2.pdf", "doc3.pdf"])
Summary
- pdf-inspector achieves sub-3-millisecond extraction times for text-based PDFs through single-pass parsing in
src/lib.rsand lightweight data structures insrc/types.rs - Stream sampling in
src/detector.rsclassifies documents in under 50ms while caching results to skip unnecessary OCR - Three-tiered table detection applies the cheapest algorithms first in
src/tables/detect_rects.rs, falling back to expensive line-grid detection only when required - The library processes approximately 400 PDFs per second single-threaded and scales linearly with CPU cores
- Optimize large PDF workloads by detecting document types before OCR, reusing
PdfDocumentinstances, and using targetedRUST_LOGdebugging for profiling
Frequently Asked Questions
How does pdf-inspector handle large PDF files without running out of memory?
pdf-inspector processes large PDFs using memory-efficient data structures defined in src/types.rs, storing only essential text metadata like coordinates and font indices rather than rendering full page bitmaps. The single-pass architecture in src/lib.rs streams content without loading redundant intermediate representations, ensuring memory usage remains proportional to document complexity rather than file size.
Can pdf-inspector process thousands of PDFs in parallel?
Yes. Because pdf-inspector is implemented in pure Rust with no external service dependencies, it scales linearly with available CPU cores. Multi-process pipelines can achieve throughput of thousands of PDFs per second, with each single-threaded instance processing approximately 400 PDFs per second on modern hardware like the Apple M4 Pro.
When should I enable OCR for pdf-inspector processed documents?
You should only enable OCR for documents classified as Scanned by the detection routine in src/detector.rs. The library's stream sampling mechanism accurately identifies text-based PDFs in 10-50 milliseconds, allowing you to bypass OCR entirely for native digital documents and avoid the significant performance penalty associated with optical character recognition.
How can I identify which processing stage is slowing down my PDF workflow?
Use the structured logging system controlled by RUST_LOG environment variables, as documented in docs/debugging.md. Set specific component levels like pdf_inspector::extractor::layout=debug or pdf_inspector::tables=debug to isolate bottlenecks in column detection, table extraction, or classification without recompiling the library or adding instrumentation code.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →