Firecrawl pdf-inspector Limitations: 9 Known Constraints Explained
Firecrawl pdf-inspector cannot extract text from scanned PDFs, has no OCR capability, caps table detection at 25 columns, and lacks support for password-protected documents or parallel processing.
Firecrawl pdf-inspector is a fast, pure-Rust PDF classifier and text-extraction library designed for speed and minimal dependencies. While it excels at converting native-text PDFs to clean Markdown, its lightweight architecture imposes specific boundaries that users should understand before adoption. This guide examines each limitation with direct references to the source code implementation.
No OCR Capability for Scanned Documents
The most significant limitation of firecrawl pdf-inspector is its complete absence of OCR functionality. The library can only classify PDFs—it cannot extract readable text from scanned pages.
The detection logic in src/detector.rs performs fast, sampling-only inspection of text operators (Tj, TJ) and image operators (Do). It returns a confidence score and a list of pages requiring OCR, but performs no actual text recognition:
// From src/lib.rs - document loading rejects encrypted files
pub fn process_pdf(path: &str) -> Result<PdfResult, Error> {
// Detection identifies scanned pages but cannot process them
let detection = detector::analyze_pdf(&doc)?;
if detection.pdf_type == PdfType::Scanned {
// Returns classification only, no extracted text for these pages
}
}
When result.pdf_type == "scanned", users must route pages to an external OCR service. The pages_needing_ocr field indicates exactly which pages require processing:
import pdf_inspector
result = pdf_inspector.process_pdf("sample.pdf")
if result.pdf_type == "scanned":
print("OCR required for pages:", result.pages_needing_ocr)
# External OCR step mandatory here
Hard-Capped Table Detection Width
Table extraction in firecrawl pdf-inspector is constrained by a 25-column maximum enforced in src/tables/grid.rs. The union-find clustering algorithm skips tables exceeding this width, and heuristic detection may fail on irregular layouts.
This architectural limit affects:
- Wide financial statements with many columns
- Complex multi-page tables with varying structures
- Irregular grid layouts that deviate from rectangular assumptions
The column detection uses histogram-valley heuristics with fixed thresholds, as implemented in src/extractor/layout.rs. Very wide tables are partially rendered or excluded entirely.
Dependency on lopdf Parser Reliability
pdf-inspector relies exclusively on the lopdf crate for low-level PDF parsing, as specified in Cargo.toml. This single-dependency design creates a single point of failure:
- Malformed PDFs with corrupted xref tables trigger unrecoverable errors
- Unusual object streams cause parse failures
- Encrypted documents abort the entire pipeline
There is no built-in fallback mechanism. When lopdf fails, extraction terminates immediately without partial results or recovery options.
No Password-Protected PDF Support
The library explicitly rejects encrypted documents. The document loading code in src/lib.rs checks for encryption and returns an error:
// Simplified from src/lib.rs
if doc.is_encrypted() {
return Err(Error::EncryptedPdf);
}
Users must decrypt PDFs before processing, adding a preprocessing step to any pipeline handling protected documents.
Limited Right-to-Left and CJK Script Handling
While src/text_utils.rs provides basic RTL detection and CJK utilities, the implementation lacks full Unicode support:
| Script Type | Support Level | Limitation |
|---|---|---|
| Basic RTL | Partial | Detection exists but no complex shaping |
| CJK | Helper functions only | No full line-breaking or substitution |
| Mixed directionality | Limited | Direction changes may garble output |
| Exotic scripts | None | Font substitution unavailable |
Complex scripts requiring glyph-level shaping or font substitution lose fidelity during extraction.
Static Column-Detection Thresholds
The layout analysis in src/extractor/layout.rs employs histogram-valley column detection with fixed thresholds. This heuristic approach struggles with:
- Newspaper-style PDFs with irregular column widths
- Non-rectilinear column layouts
- Highly variable page designs within single documents
Misdetected columns produce incorrect reading order, corrupting downstream Markdown structure.
No Actual Image Content Extraction
Images and vector graphics receive placeholder treatment only. The src/extractor/xobjects.rs module extracts markers indicating image presence, but never exports actual pixel data:
// From src/extractor/xobjects.rs - placeholder extraction only
pub fn extract_images(page: &Page) -> Vec<ImagePlaceholder> {
// Returns metadata, not image bytes
vec![ImagePlaceholder { rect, .. }]
}
Use cases requiring embedded figures, diagrams, or base64-encoded images must implement separate image handling pipelines.
Single-Threaded Processing Architecture
The process_pdf function in src/lib.rs operates synchronously without internal parallelism:
pub fn process_pdf(path: &str) -> Result<PdfResult, Error> {
// Sequential single-pass processing
let doc = Document::load(path)?;
let detection = detector::analyze_pdf(&doc)?;
let markdown = extractor::extract_markdown(&doc, &options)?;
// No Tokio, no Rayon, no thread spawning
}
Large documents exceeding 500 pages cannot leverage multi-core processing. Users must implement external parallelism at the file level if batch throughput is critical.
Limited Configurability via PdfOptions
The PdfOptions builder in src/lib.rs exposes coarse controls like scan strategy, but fine-grained tuning is unavailable:
- Table heuristic thresholds are hardcoded in
src/tables/grid.rs - Heading detection parameters fixed in layout engine
- Font-size clustering boundaries not user-adjustable
Advanced customization requires source modification and recompilation.
Summary
- No OCR capability — scanned pages require external processing; the library only classifies PDF types
- 25-column table limit — wide tables skipped by union-find clustering in
src/tables/grid.rs - lopdf dependency fragility — malformed or encrypted PDFs cause complete pipeline failure
- Password-protected PDF rejection — encrypted documents return errors from
src/lib.rs - Incomplete script support — RTL and CJK handling lacks full shaping and line-breaking
- Fixed layout heuristics — histogram-valley column detection fails on irregular layouts
- Placeholder-only images —
src/extractor/xobjects.rsnever exports actual image data - Synchronous single-threading — no internal parallelism for large document processing
- Restricted customization —
PdfOptionslacks fine-grained control over extraction parameters
Firecrawl pdf-inspector prioritizes speed, determinism, and minimal dependencies over universal PDF handling. It excels with native-text documents but requires complementary tools for OCR, decryption, image extraction, and complex layout analysis.
Frequently Asked Questions
Does firecrawl pdf-inspector support OCR for scanned PDFs?
No. The library can detect that a PDF is scanned and identify which pages need OCR via pages_needing_ocr, but it cannot perform text extraction on those pages itself. Users must integrate with external OCR services like Tesseract, AWS Textract, or Azure AI Document Intelligence for actual text recognition.
What happens when pdf-inspector encounters a password-protected PDF?
The library rejects encrypted documents immediately. The document loading code in src/lib.rs checks for the encryption flag and returns an Error::EncryptedPdf before any processing begins. Users must decrypt files using qpdf, pdftk, or similar tools before passing them to pdf-inspector.
Why are some tables missing or malformed in the Markdown output?
Table detection uses union-find clustering with a hard limit of 25 columns defined in src/tables/grid.rs. Tables exceeding this width are skipped. Additionally, the histogram-valley heuristics in src/extractor/layout.rs may misidentify irregular or non-rectangular table structures. Very wide financial tables and complex multi-page layouts are most commonly affected.
Can I extract embedded images from PDFs using pdf-inspector?
No. The src/extractor/xobjects.rs module extracts only image placeholders—metadata rectangles indicating where images appear on pages. Actual pixel data, base64 encoding, and separate image file export are not implemented. For image extraction, use dedicated tools like pdf2image, Poppler, or PyMuPDF.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →