# Performance Characteristics of pdf-inspector: How to Optimize for Large PDFs

> Discover pdf-inspector performance benchmarks and optimization tips for large PDFs. Learn how single-pass parsing achieves 2.35ms per document processing.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: performance
- Published: 2026-08-08

---

**pdf-inspector processes text-based PDFs in approximately 2.35 milliseconds per document by leveraging single-pass parsing, lightweight stream sampling for classification, and a three-tiered detection system that avoids expensive OCR and ML models unless absolutely necessary.**

The firecrawl/pdf-inspector library is engineered for speed-first PDF processing in Rust, delivering sub-50-millisecond classification and sub-3-millisecond extraction times for typical documents. By avoiding heavyweight ML frameworks and redundant I/O operations, pdf-inspector maintains a minimal resource footprint while handling large-scale PDF workloads. Understanding these performance characteristics enables developers to maximize throughput when processing high volumes of documents or multi-megabyte files.

## Single-Pass Document Loading and Memory Efficiency

The foundation of pdf-inspector's speed lies in its **single-pass document loading** architecture. In [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), the public API functions `load_document_from_path` and `load_document_from_mem` parse the PDF once and share the same internal representation between the classifier and extractor. This eliminates duplicate parsing work and redundant I/O operations that plague traditional multi-stage pipelines.

Memory efficiency is achieved through lightweight data structures defined in [`src/types.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs). The library stores text items as compact `TextItem` structs containing only essential coordinates, font index, and optional underline flags. Unlike rendering-based PDF libraries, pdf-inspector never allocates large intermediate canvases, keeping memory usage proportional to document complexity rather than page dimensions.

The dependency footprint remains minimal with only `lopdf` as an external crate, as specified in [`Cargo.toml`](https://github.com/firecrawl/pdf-inspector/blob/main/Cargo.toml). This low-level parsing library avoids pulling in heavyweight frameworks, resulting in small binary sizes and fast startup times critical for serverless deployments.

## Fast Classification via Stream Sampling

PDF type detection in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) uses **stream sampling** to classify documents in approximately 10-50 milliseconds. Rather than analyzing every byte, the detector examines a small subset of content streams to determine whether a PDF is text-based, scanned, image-based, or mixed. The classification result is cached, allowing the extractor to skip OCR processing entirely when it isn't needed.

This approach ensures that **text-based PDFs bypass expensive OCR pipelines**, while scanned documents are flagged for alternative processing. The short-circuit logic prevents wasted computation on documents that don't require computer vision techniques.

## Intelligent Layout and Table Detection

For complex layouts, pdf-inspector implements algorithms optimized for computational efficiency. **Column detection** in [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs) builds horizontal projection histograms and identifies valleys to separate newspaper-style multi-column content. This algorithm runs in O(N) time over the page's text items, significantly faster than expensive layout engines that render entire pages.

**Table detection** follows a three-tiered strategy implemented across three source files:

1. [`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs) – Rectangle-based detection using drawing operations and fast union-find data structures
2. [`src/tables/detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs) – Heuristic alignment detection for structured grids  
3. [`src/tables/detect_lines.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_lines.rs) – Line-grid detection only when the first two methods fail

This cascading approach short-circuits expensive processing for the majority of pages, applying computationally intensive line detection only when simpler methods prove insufficient.

## Benchmark Results and Throughput

According to the benchmark methodology documented in [`docs/benchmarking.md`](https://github.com/firecrawl/pdf-inspector/blob/main/docs/benchmarking.md), pdf-inspector achieves exceptional performance on modern hardware. On an Apple M4 Pro, processing 200 PDFs without OCR takes **0.470 seconds total**, equating to approximately **2.35 milliseconds per document** with consistent median speeds across five full-corpus runs.

The practical performance profile includes:

- **Classification latency**: Under 50 milliseconds per document
- **Extraction latency**: Under 3 milliseconds for typical text-based PDFs  
- **Throughput**: Approximately 400 PDFs per second on a modern laptop in single-threaded execution
- **Scalability**: Linear scaling with CPU cores; multi-process pipelines can process thousands of PDFs per second due to the library's pure Rust implementation with no external service dependencies

## Optimization Strategies for Large PDFs

To maximize performance when processing large PDFs or high-volume workloads, implement these targeted optimizations:

**Avoid Unnecessary OCR**

Invoke the `detect-pdf` functionality first and only run OCR engines on pages explicitly flagged as `Scanned`. Since pdf-inspector accurately identifies text-based content through stream sampling, you prevent the 100-1000x performance penalty associated with unnecessary optical character recognition.

**Reuse PdfDocument Instances**

When you need both classification and extraction, call `process_pdf` once rather than invoking separate operations. The internal document representation is reused automatically, eliminating redundant parsing overhead for large files.

**Profile with Structured Logging**

Enable targeted debugging through the `RUST_LOG` environment variable. Configure `pdf_inspector::extractor::layout=debug` or `pdf_inspector=debug` to emit fine-grained diagnostics for specific components. This allows precise profiling of runtime bottlenecks without code recompilation or instrumentation overhead, as documented in [`docs/debugging.md`](https://github.com/firecrawl/pdf-inspector/blob/main/docs/debugging.md).

### Measuring Performance in Practice

Monitor processing times using these implementation patterns:

**Command Line Timing:**

```bash
time pdf2md sample.pdf > /dev/null

```

**Rust Explicit Timing:**

```rust
use std::time::Instant;
use pdf_inspector::process_pdf;

fn main() -> anyhow::Result<()> {
    let start = Instant::now();
    let result = process_pdf("sample.pdf")?;
    let elapsed = start.elapsed();
    println!("PDF type: {:?}", result.pdf_type);
    println!("Markdown length: {}", result.markdown.as_ref().map(|s| s.len()).unwrap_or(0));
    println!("Processing time: {:.2?}", elapsed);
    Ok(())
}

```

**Python Batch Processing:**

```python
import time
import pdf_inspector

def process_batch(paths):
    start = time.time()
    for p in paths:
        r = pdf_inspector.process_pdf(p)
        # Access r.pdf_type and r.markdown as needed

    print(f"Processed {len(paths)} PDFs in {time.time() - start:.2f}s")

process_batch(["doc1.pdf", "doc2.pdf", "doc3.pdf"])

```

## Summary

- pdf-inspector achieves **sub-3-millisecond extraction times** for text-based PDFs through single-pass parsing in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) and lightweight data structures in [`src/types.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs)
- **Stream sampling** in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) classifies documents in under 50ms while caching results to skip unnecessary OCR
- **Three-tiered table detection** applies the cheapest algorithms first in [`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs), falling back to expensive line-grid detection only when required
- The library processes approximately **400 PDFs per second** single-threaded and scales linearly with CPU cores
- Optimize large PDF workloads by detecting document types before OCR, reusing `PdfDocument` instances, and using targeted `RUST_LOG` debugging for profiling

## Frequently Asked Questions

### How does pdf-inspector handle large PDF files without running out of memory?

pdf-inspector processes large PDFs using memory-efficient data structures defined in [`src/types.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs), storing only essential text metadata like coordinates and font indices rather than rendering full page bitmaps. The single-pass architecture in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) streams content without loading redundant intermediate representations, ensuring memory usage remains proportional to document complexity rather than file size.

### Can pdf-inspector process thousands of PDFs in parallel?

Yes. Because pdf-inspector is implemented in pure Rust with no external service dependencies, it scales linearly with available CPU cores. Multi-process pipelines can achieve throughput of thousands of PDFs per second, with each single-threaded instance processing approximately 400 PDFs per second on modern hardware like the Apple M4 Pro.

### When should I enable OCR for pdf-inspector processed documents?

You should only enable OCR for documents classified as `Scanned` by the detection routine in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs). The library's stream sampling mechanism accurately identifies text-based PDFs in 10-50 milliseconds, allowing you to bypass OCR entirely for native digital documents and avoid the significant performance penalty associated with optical character recognition.

### How can I identify which processing stage is slowing down my PDF workflow?

Use the structured logging system controlled by `RUST_LOG` environment variables, as documented in [`docs/debugging.md`](https://github.com/firecrawl/pdf-inspector/blob/main/docs/debugging.md). Set specific component levels like `pdf_inspector::extractor::layout=debug` or `pdf_inspector::tables=debug` to isolate bottlenecks in column detection, table extraction, or classification without recompiling the library or adding instrumentation code.