# Performance Considerations for pdf-inspector: Optimizing High-Speed PDF Processing

> Discover performance considerations for pdf-inspector. Achieve sub-3ms extraction latency with optimized PDF processing, single-pass loading, and tiered detection.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: performance
- Published: 2026-08-04

---

**pdf-inspector achieves sub-3ms extraction latency for text-based PDFs by avoiding heavyweight ML models, using single-pass document loading, and employing tiered detection algorithms that short-circuit expensive processing steps.**

The `firecrawl/pdf-inspector` library is engineered for speed-first PDF processing in Rust, designed to classify and extract content without the overhead of traditional OCR engines or machine learning frameworks. Understanding the specific architectural decisions behind its performance profile helps developers maximize throughput in production environments. Below is a technical breakdown of the performance considerations for pdf-inspector, mapped directly to its implementation.

## Single-Pass Document Loading

The library eliminates redundant I/O by parsing the PDF document exactly once, regardless of how many operations you perform on it. In [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), the public API functions `load_document_from_path` and `load_document_from_mem` create a single internal representation that is shared between the classifier and the extractor.

This means when you call `process_pdf`, the document is loaded into memory once and reused for both PDF type detection and Markdown extraction. Avoiding duplicate parsing work significantly reduces latency, particularly for large files or high-throughput batch jobs.

## Fast Classification via Stream Sampling

Before any expensive processing begins, pdf-inspector uses a lightweight detection phase to determine document type. The detector in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) samples a small subset of the PDF's content streams—typically completing in 10–50ms—to classify the document as text-based, scanned, image-based, or mixed.

The classification result is cached internally, allowing the extractor to skip OCR entirely when processing native text PDFs. This short-circuiting behavior is critical for performance: since OCR engines can take seconds per page, avoiding unnecessary optical character recognition provides orders-of-magnitude speed improvements for typical document corpora.

## Lightweight Dependency Footprint

Performance isn't just about runtime speed; it's also about startup time and binary size. The library maintains a minimal dependency tree, with `lopdf` as its only external crate for low-level PDF parsing. As specified in [`Cargo.toml`](https://github.com/firecrawl/pdf-inspector/blob/main/Cargo.toml), this deliberate constraint avoids pulling in heavyweight PDF libraries or ML frameworks that would bloat the binary and slow initialization.

The pure Rust implementation with zero external service dependencies means pdf-inspector scales linearly with available CPU cores without network latency or model loading overhead.

## Efficient Layout Analysis Algorithms

Multi-column document processing often requires expensive layout engines, but pdf-inspector uses an optimized histogram-based approach. In [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs), the column detection routine builds a horizontal projection histogram and identifies valleys to separate columns.

This algorithm runs in **O(N)** time over the page's text items, where N is the number of text elements, making it significantly cheaper than complex layout analysis engines. The same module handles line grouping and reading-order computation using memory-efficient geometric calculations rather than heavy intermediate representations.

## Tiered Table Detection Strategy

Table extraction implements a three-tiered short-circuiting strategy that prioritizes computational cheapness:

1. **Rectangle-based detection** ([`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs)) – Uses drawing operations and fast union-find data structures to identify tables by their border rectangles.
2. **Heuristic alignment detection** ([`src/tables/detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs)) – Applies pattern matching to detect tables based on text alignment when borders are absent.
3. **Line-grid detection** ([`src/tables/detect_lines.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_lines.rs)) – Only executes if the first two methods fail, using more expensive line intersection analysis.

This tiered approach ensures the majority of pages exit early at the rectangle or heuristic stage, avoiding the computational cost of full line-grid analysis except when absolutely necessary.

## Memory-Efficient Data Structures

The library minimizes heap allocations by using lightweight structs defined in [`src/types.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs). The `TextItem` and `PdfRect` types store only essential data: coordinates, font index, and optional underline flags. 

Unlike PDF processing libraries that allocate large intermediate canvases or raster buffers, pdf-inspector operates directly on the PDF's content streams using these compact structures. This keeps memory usage modest even when processing hundreds of documents concurrently.

## Profiling with Structured Logging

When optimizing specific workloads, developers can enable fine-grained diagnostics without recompiling. The library uses the `RUST_LOG` environment variable to emit performance data (e.g., `pdf_inspector::extractor::layout=debug`), as documented in [`docs/debugging.md`](https://github.com/firecrawl/pdf-inspector/blob/main/docs/debugging.md).

To minimize overhead during profiling, enable debug logging only for specific components rather than the entire library. This allows precise measurement of which stage—classification, layout analysis, or table detection—dominates runtime for your specific PDF corpus.

## Benchmark Results and Throughput

According to the methodology documented in [`docs/benchmarking.md`](https://github.com/firecrawl/pdf-inspector/blob/main/docs/benchmarking.md), pdf-inspector demonstrates consistent high performance:

- **Latency**: Classification completes in under 50ms; full extraction and Markdown conversion takes under 3ms for typical text-based PDFs.
- **Throughput**: Approximately 400 PDFs per second on a modern laptop (Apple M4 Pro) in single-threaded execution.
- **Scalability**: Linear scaling with CPU cores; multi-process pipelines can achieve thousands of PDFs per second.

On the benchmark hardware, processing 200 PDFs without OCR required only 0.470 seconds total—approximately **2.35ms per document** across five full-corpus runs.

## Optimization Best Practices

To achieve optimal performance in production deployments:

- **Avoid unnecessary OCR** – Always invoke detection first and only run external OCR engines on pages flagged as `Scanned`.
- **Reuse the document handle** – When you need both classification and extraction, use `process_pdf` once rather than calling separate functions; the internal `PdfDocument` is reused automatically.
- **Batch process efficiently** – Process documents in parallel using Rust's concurrency primitives or multi-process Python workers to saturate CPU cores.

### Measuring Performance

Use these patterns to benchmark your specific workload:

**CLI timing:**

```bash
time pdf2md sample.pdf > /dev/null

```

**Rust explicit timing:**

```rust
use std::time::Instant;
use pdf_inspector::process_pdf;

fn main() -> anyhow::Result<()> {
    let start = Instant::now();
    let result = process_pdf("sample.pdf")?;
    let elapsed = start.elapsed();
    println!("PDF type: {:?}", result.pdf_type);
    println!("Markdown length: {}", result.markdown.as_ref().map(|s| s.len()).unwrap_or(0));
    println!("Processing time: {:.2?}", elapsed);
    Ok(())
}

```

**Python batch processing:**

```python
import time
import pdf_inspector

def process_batch(paths):
    start = time.time()
    for p in paths:
        r = pdf_inspector.process_pdf(p)
        # Use r.pdf_type / r.markdown as needed

    print(f"Processed {len(paths)} PDFs in {time.time() - start:.2f}s")

process_batch(["doc1.pdf", "doc2.pdf", "doc3.pdf"])

```

## Summary

- **Single-pass loading** in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) eliminates duplicate parsing by sharing internal document representations between classification and extraction.
- **Fast stream sampling** in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) caches PDF type detection (10–50ms) to skip expensive OCR on native text documents.
- **Minimal dependencies** (only `lopdf` in [`Cargo.toml`](https://github.com/firecrawl/pdf-inspector/blob/main/Cargo.toml)) ensure fast startup and small binary size without ML framework overhead.
- **O(N) column detection** in [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs) uses histogram valleys instead of expensive layout engines.
- **Three-tiered table detection** short-circuits at the cheapest viable method (rectangles → heuristics → line grids).
- **Memory-efficient structs** (`TextItem`, `PdfRect` in [`src/types.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs)) avoid large intermediate allocations.
- **Sub-3ms latency** for text PDFs and **~400 PDFs/second** throughput on modern hardware make pdf-inspector suitable for high-volume pipelines.

## Frequently Asked Questions

### How does pdf-inspector achieve sub-3ms extraction times?

pdf-inspector achieves millisecond-level extraction by avoiding heavyweight operations entirely for native text PDFs. It uses fast stream sampling in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) to confirm a document contains extractable text, then processes that text directly using O(N) geometric algorithms in [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs) without invoking OCR engines or ML models.

### When should I use OCR with pdf-inspector?

You only need external OCR when the detector classifies a PDF as `Scanned` or `Mixed`. The classification logic in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) analyzes content streams to determine if text is natively available; if so, the extractor in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) bypasses OCR entirely. This selective approach prevents the multi-second latency typically associated with optical character recognition.

### What is the memory overhead for processing large PDFs?

The library maintains modest memory usage by using lightweight structs defined in [`src/types.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs) rather than allocating raster canvases or intermediate buffers. Since `TextItem` and `PdfRect` contain only essential metadata (coordinates, font indices), memory scales with the number of text elements rather than page resolution or file size.

### Can pdf-inspector handle high-throughput batch processing?

Yes. The library processes approximately 400 PDFs per second single-threaded on modern hardware (Apple M4 Pro) according to [`docs/benchmarking.md`](https://github.com/firecrawl/pdf-inspector/blob/main/docs/benchmarking.md), and scales linearly with CPU cores. Because it is pure Rust with no external services or model loading, you can run multi-process pipelines to achieve thousands of documents per second without network bottlenecks or GPU constraints.