Firecrawl pdf‑inspector Performance on Large PDFs: Benchmarks and Optimization Strategies

Firecrawl pdf‑inspector processes a 300‑page PDF in approximately 150–200 ms total, with PDF‑type detection alone taking just 10–50 ms, making it one of the fastest local PDF‑to‑Markdown solutions available.

The firecrawl/pdf-inspector repository delivers a Rust‑based PDF processing pipeline optimized for speed without OCR overhead. For developers building high‑throughput document ingestion systems, understanding how this tool handles large PDFs is critical to capacity planning and pipeline design.

How pdf‑inspector Achieves Sub‑Second Performance on Large PDFs

The architecture splits work into two distinct stages, each optimized for minimal latency:

Stage Typical Latency Implementation Detail
PDF‑type detection 10–50 ms Early‑exit content stream sampling; stops on first non‑text page
Full extraction + Markdown conversion ~150 ms (300‑page PDF) Single‑pass parsing with shared structure between stages

According to the project's benchmark suite on the opendataloader‑bench corpus—200 PDFs of varying lengths—the total runtime for the entire batch is 0.470 s, averaging ~2.35 ms per PDF for smaller documents. Larger PDFs dominate the timing and cluster in the 100–200 ms range.

Key Optimizations in the Source Code

Early‑Exit Detection Strategy

In src/detector.rs, the classifier implements smart sampling that avoids full document scans. The default behavior samples content streams and terminates immediately upon finding the first non‑text page. For even greater speed, users can enable a Sample(n) strategy that checks only n evenly‑spaced pages, keeping classification time bounded regardless of document length.

This design ensures that a 500‑page PDF takes roughly the same time to classify as a 50‑page PDF when sampling is enabled.

Shared Parsed Structure

The extractor in src/extractor/mod.rs loads the PDF once and shares the parsed structure between detection and extraction phases. This eliminates redundant I/O operations that would otherwise double‑ or triple‑wall‑clock time on large files.

Priority‑Chain Table Detection

Table detection runs through a fast‑fail sequence: rectangle‑based detection first, then heuristic fallback, stopping immediately when any method produces a valid result. This limits computational work on complex layouts without sacrificing accuracy.

Measuring Performance Yourself

CLI Timing (Linux/macOS)


# Install the binary once

cargo install pdf-inspector

# Time processing on a large PDF

time pdf2md report.pdf > out.md

On an Apple M4 Pro, typical output shows ~0.15 s total for a 300‑page document.

Python Binding Benchmark

import time
import pdf_inspector

pdf_path = "large_report.pdf"

# Stage 1: Classification only

t0 = time.time()
info = pdf_inspector.classify_pdf(pdf_path)      # fast detection

print(f"Classification: {info.pdf_type} in {(time.time()-t0)*1000:.1f} ms")

# Stage 2: Full pipeline

t0 = time.time()
result = pdf_inspector.process_pdf(pdf_path)     # extraction + markdown

print(f"Extraction: {len(result.markdown.splitlines())} lines")
print(f"Full run: {(time.time()-t0)*1000:.1f} ms")

Rust API Single‑Call Processing

use pdf_inspector::process_pdf;
use std::time::Instant;

let path = "large_report.pdf";
let start = Instant::now();
let res = process_pdf(path).expect("process failed");

println!("PDF type: {:?}", res.pdf_type);
println!("Markdown size: {} bytes", res.markdown.unwrap_or_default().len());
println!("Total time: {:.1} ms", start.elapsed().as_millis());

All three approaches demonstrate consistent sub‑second performance even on multi‑hundred‑page documents.

Critical Source Files for Performance Analysis

File Performance Role
src/detector.rs Fast classification with early‑exit and sampling strategies
src/extractor/mod.rs Coordinates font loading, content‑stream walking, and layout analysis with single‑pass parsing
src/markdown/convert.rs Converts TextItem structures to clean Markdown; handles tables and headings efficiently
src/bin/pdf2md.rs CLI entry point wiring detection → extraction → output
README.md (benchmark section) Documents real‑world performance on the 200‑PDF opendataloader‑bench corpus

Summary

  • Firecrawl pdf‑inspector achieves ~150 ms full‑pipeline latency on 300‑page PDFs via Rust‑optimized, OCR‑free processing
  • Early‑exit detection (10–50 ms) uses sampling strategies that scale independently of document size
  • Shared parsed structures eliminate redundant I/O between detection and extraction stages
  • Priority‑chain table detection fast‑fails on first valid result to limit computation
  • Benchmark hardware (Apple M4 Pro) demonstrates production‑ready throughput for high‑volume pipelines

Frequently Asked Questions

How does pdf‑inspector handle PDFs with thousands of pages?

The Sample(n) detection strategy in src/detector.rs bounds classification time by checking only n evenly‑spaced pages regardless of total page count. Extraction remains linear with document size but processes at roughly 0.5 ms per page on modern hardware, keeping even 1000‑page documents under one second.

Why is pdf‑inspector faster than OCR‑based solutions?

OCR requires rasterization and neural network inference, typically adding 1–10 seconds per page. Pdf‑inspector operates directly on PDF content streams and font data, bypassing image generation entirely. This architectural difference explains the 10–100x speed advantage on text‑based PDFs.

Can performance degrade on scanned or image‑heavy PDFs?

Yes. The early‑exit detector in src/detector.rs quickly identifies image‑dominated documents and flags them accordingly. While pdf‑inspector will still process these files, extraction time increases as it falls back to layout heuristics rather than direct text extraction. The tool is optimized for native PDFs with embedded text.

What hardware was used for the published benchmarks?

All documented performance figures come from an Apple M4 Pro CPU. The Rust implementation is single‑threaded for individual PDFs, so performance scales predictably with per‑core clock speed. Throughput scales linearly with available cores when processing batches.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →