Firecrawl pdf‑inspector Performance on Large PDFs: Benchmarks and Optimization Strategies
Firecrawl pdf‑inspector processes a 300‑page PDF in approximately 150–200 ms total, with PDF‑type detection alone taking just 10–50 ms, making it one of the fastest local PDF‑to‑Markdown solutions available.
The firecrawl/pdf-inspector repository delivers a Rust‑based PDF processing pipeline optimized for speed without OCR overhead. For developers building high‑throughput document ingestion systems, understanding how this tool handles large PDFs is critical to capacity planning and pipeline design.
How pdf‑inspector Achieves Sub‑Second Performance on Large PDFs
The architecture splits work into two distinct stages, each optimized for minimal latency:
| Stage | Typical Latency | Implementation Detail |
|---|---|---|
| PDF‑type detection | 10–50 ms | Early‑exit content stream sampling; stops on first non‑text page |
| Full extraction + Markdown conversion | ~150 ms (300‑page PDF) | Single‑pass parsing with shared structure between stages |
According to the project's benchmark suite on the opendataloader‑bench corpus—200 PDFs of varying lengths—the total runtime for the entire batch is 0.470 s, averaging ~2.35 ms per PDF for smaller documents. Larger PDFs dominate the timing and cluster in the 100–200 ms range.
Key Optimizations in the Source Code
Early‑Exit Detection Strategy
In src/detector.rs, the classifier implements smart sampling that avoids full document scans. The default behavior samples content streams and terminates immediately upon finding the first non‑text page. For even greater speed, users can enable a Sample(n) strategy that checks only n evenly‑spaced pages, keeping classification time bounded regardless of document length.
This design ensures that a 500‑page PDF takes roughly the same time to classify as a 50‑page PDF when sampling is enabled.
Shared Parsed Structure
The extractor in src/extractor/mod.rs loads the PDF once and shares the parsed structure between detection and extraction phases. This eliminates redundant I/O operations that would otherwise double‑ or triple‑wall‑clock time on large files.
Priority‑Chain Table Detection
Table detection runs through a fast‑fail sequence: rectangle‑based detection first, then heuristic fallback, stopping immediately when any method produces a valid result. This limits computational work on complex layouts without sacrificing accuracy.
Measuring Performance Yourself
CLI Timing (Linux/macOS)
# Install the binary once
cargo install pdf-inspector
# Time processing on a large PDF
time pdf2md report.pdf > out.md
On an Apple M4 Pro, typical output shows ~0.15 s total for a 300‑page document.
Python Binding Benchmark
import time
import pdf_inspector
pdf_path = "large_report.pdf"
# Stage 1: Classification only
t0 = time.time()
info = pdf_inspector.classify_pdf(pdf_path) # fast detection
print(f"Classification: {info.pdf_type} in {(time.time()-t0)*1000:.1f} ms")
# Stage 2: Full pipeline
t0 = time.time()
result = pdf_inspector.process_pdf(pdf_path) # extraction + markdown
print(f"Extraction: {len(result.markdown.splitlines())} lines")
print(f"Full run: {(time.time()-t0)*1000:.1f} ms")
Rust API Single‑Call Processing
use pdf_inspector::process_pdf;
use std::time::Instant;
let path = "large_report.pdf";
let start = Instant::now();
let res = process_pdf(path).expect("process failed");
println!("PDF type: {:?}", res.pdf_type);
println!("Markdown size: {} bytes", res.markdown.unwrap_or_default().len());
println!("Total time: {:.1} ms", start.elapsed().as_millis());
All three approaches demonstrate consistent sub‑second performance even on multi‑hundred‑page documents.
Critical Source Files for Performance Analysis
| File | Performance Role |
|---|---|
src/detector.rs |
Fast classification with early‑exit and sampling strategies |
src/extractor/mod.rs |
Coordinates font loading, content‑stream walking, and layout analysis with single‑pass parsing |
src/markdown/convert.rs |
Converts TextItem structures to clean Markdown; handles tables and headings efficiently |
src/bin/pdf2md.rs |
CLI entry point wiring detection → extraction → output |
README.md (benchmark section) |
Documents real‑world performance on the 200‑PDF opendataloader‑bench corpus |
Summary
- Firecrawl pdf‑inspector achieves ~150 ms full‑pipeline latency on 300‑page PDFs via Rust‑optimized, OCR‑free processing
- Early‑exit detection (10–50 ms) uses sampling strategies that scale independently of document size
- Shared parsed structures eliminate redundant I/O between detection and extraction stages
- Priority‑chain table detection fast‑fails on first valid result to limit computation
- Benchmark hardware (Apple M4 Pro) demonstrates production‑ready throughput for high‑volume pipelines
Frequently Asked Questions
How does pdf‑inspector handle PDFs with thousands of pages?
The Sample(n) detection strategy in src/detector.rs bounds classification time by checking only n evenly‑spaced pages regardless of total page count. Extraction remains linear with document size but processes at roughly 0.5 ms per page on modern hardware, keeping even 1000‑page documents under one second.
Why is pdf‑inspector faster than OCR‑based solutions?
OCR requires rasterization and neural network inference, typically adding 1–10 seconds per page. Pdf‑inspector operates directly on PDF content streams and font data, bypassing image generation entirely. This architectural difference explains the 10–100x speed advantage on text‑based PDFs.
Can performance degrade on scanned or image‑heavy PDFs?
Yes. The early‑exit detector in src/detector.rs quickly identifies image‑dominated documents and flags them accordingly. While pdf‑inspector will still process these files, extraction time increases as it falls back to layout heuristics rather than direct text extraction. The tool is optimized for native PDFs with embedded text.
What hardware was used for the published benchmarks?
All documented performance figures come from an Apple M4 Pro CPU. The Rust implementation is single‑threaded for individual PDFs, so performance scales predictably with per‑core clock speed. Throughput scales linearly with available cores when processing batches.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →