Firecrawl pdf-inspector Limitations: 9 Known Constraints Explained

Firecrawl pdf-inspector cannot extract text from scanned PDFs, has no OCR capability, caps table detection at 25 columns, and lacks support for password-protected documents or parallel processing.

Firecrawl pdf-inspector is a fast, pure-Rust PDF classifier and text-extraction library designed for speed and minimal dependencies. While it excels at converting native-text PDFs to clean Markdown, its lightweight architecture imposes specific boundaries that users should understand before adoption. This guide examines each limitation with direct references to the source code implementation.

No OCR Capability for Scanned Documents

The most significant limitation of firecrawl pdf-inspector is its complete absence of OCR functionality. The library can only classify PDFs—it cannot extract readable text from scanned pages.

The detection logic in src/detector.rs performs fast, sampling-only inspection of text operators (Tj, TJ) and image operators (Do). It returns a confidence score and a list of pages requiring OCR, but performs no actual text recognition:

// From src/lib.rs - document loading rejects encrypted files
pub fn process_pdf(path: &str) -> Result<PdfResult, Error> {
    // Detection identifies scanned pages but cannot process them
    let detection = detector::analyze_pdf(&doc)?;
    if detection.pdf_type == PdfType::Scanned {
        // Returns classification only, no extracted text for these pages
    }
}

When result.pdf_type == "scanned", users must route pages to an external OCR service. The pages_needing_ocr field indicates exactly which pages require processing:

import pdf_inspector

result = pdf_inspector.process_pdf("sample.pdf")

if result.pdf_type == "scanned":
    print("OCR required for pages:", result.pages_needing_ocr)
    # External OCR step mandatory here

Hard-Capped Table Detection Width

Table extraction in firecrawl pdf-inspector is constrained by a 25-column maximum enforced in src/tables/grid.rs. The union-find clustering algorithm skips tables exceeding this width, and heuristic detection may fail on irregular layouts.

This architectural limit affects:

  • Wide financial statements with many columns
  • Complex multi-page tables with varying structures
  • Irregular grid layouts that deviate from rectangular assumptions

The column detection uses histogram-valley heuristics with fixed thresholds, as implemented in src/extractor/layout.rs. Very wide tables are partially rendered or excluded entirely.

Dependency on lopdf Parser Reliability

pdf-inspector relies exclusively on the lopdf crate for low-level PDF parsing, as specified in Cargo.toml. This single-dependency design creates a single point of failure:

  • Malformed PDFs with corrupted xref tables trigger unrecoverable errors
  • Unusual object streams cause parse failures
  • Encrypted documents abort the entire pipeline

There is no built-in fallback mechanism. When lopdf fails, extraction terminates immediately without partial results or recovery options.

No Password-Protected PDF Support

The library explicitly rejects encrypted documents. The document loading code in src/lib.rs checks for encryption and returns an error:

// Simplified from src/lib.rs
if doc.is_encrypted() {
    return Err(Error::EncryptedPdf);
}

Users must decrypt PDFs before processing, adding a preprocessing step to any pipeline handling protected documents.

Limited Right-to-Left and CJK Script Handling

While src/text_utils.rs provides basic RTL detection and CJK utilities, the implementation lacks full Unicode support:

Script Type Support Level Limitation
Basic RTL Partial Detection exists but no complex shaping
CJK Helper functions only No full line-breaking or substitution
Mixed directionality Limited Direction changes may garble output
Exotic scripts None Font substitution unavailable

Complex scripts requiring glyph-level shaping or font substitution lose fidelity during extraction.

Static Column-Detection Thresholds

The layout analysis in src/extractor/layout.rs employs histogram-valley column detection with fixed thresholds. This heuristic approach struggles with:

  • Newspaper-style PDFs with irregular column widths
  • Non-rectilinear column layouts
  • Highly variable page designs within single documents

Misdetected columns produce incorrect reading order, corrupting downstream Markdown structure.

No Actual Image Content Extraction

Images and vector graphics receive placeholder treatment only. The src/extractor/xobjects.rs module extracts markers indicating image presence, but never exports actual pixel data:

// From src/extractor/xobjects.rs - placeholder extraction only
pub fn extract_images(page: &Page) -> Vec<ImagePlaceholder> {
    // Returns metadata, not image bytes
    vec![ImagePlaceholder { rect, .. }]
}

Use cases requiring embedded figures, diagrams, or base64-encoded images must implement separate image handling pipelines.

Single-Threaded Processing Architecture

The process_pdf function in src/lib.rs operates synchronously without internal parallelism:

pub fn process_pdf(path: &str) -> Result<PdfResult, Error> {
    // Sequential single-pass processing
    let doc = Document::load(path)?;
    let detection = detector::analyze_pdf(&doc)?;
    let markdown = extractor::extract_markdown(&doc, &options)?;
    // No Tokio, no Rayon, no thread spawning
}

Large documents exceeding 500 pages cannot leverage multi-core processing. Users must implement external parallelism at the file level if batch throughput is critical.

Limited Configurability via PdfOptions

The PdfOptions builder in src/lib.rs exposes coarse controls like scan strategy, but fine-grained tuning is unavailable:

  • Table heuristic thresholds are hardcoded in src/tables/grid.rs
  • Heading detection parameters fixed in layout engine
  • Font-size clustering boundaries not user-adjustable

Advanced customization requires source modification and recompilation.

Summary

  • No OCR capability — scanned pages require external processing; the library only classifies PDF types
  • 25-column table limit — wide tables skipped by union-find clustering in src/tables/grid.rs
  • lopdf dependency fragility — malformed or encrypted PDFs cause complete pipeline failure
  • Password-protected PDF rejection — encrypted documents return errors from src/lib.rs
  • Incomplete script support — RTL and CJK handling lacks full shaping and line-breaking
  • Fixed layout heuristics — histogram-valley column detection fails on irregular layouts
  • Placeholder-only images — src/extractor/xobjects.rs never exports actual image data
  • Synchronous single-threading — no internal parallelism for large document processing
  • Restricted customization — PdfOptions lacks fine-grained control over extraction parameters

Firecrawl pdf-inspector prioritizes speed, determinism, and minimal dependencies over universal PDF handling. It excels with native-text documents but requires complementary tools for OCR, decryption, image extraction, and complex layout analysis.

Frequently Asked Questions

Does firecrawl pdf-inspector support OCR for scanned PDFs?

No. The library can detect that a PDF is scanned and identify which pages need OCR via pages_needing_ocr, but it cannot perform text extraction on those pages itself. Users must integrate with external OCR services like Tesseract, AWS Textract, or Azure AI Document Intelligence for actual text recognition.

What happens when pdf-inspector encounters a password-protected PDF?

The library rejects encrypted documents immediately. The document loading code in src/lib.rs checks for the encryption flag and returns an Error::EncryptedPdf before any processing begins. Users must decrypt files using qpdf, pdftk, or similar tools before passing them to pdf-inspector.

Why are some tables missing or malformed in the Markdown output?

Table detection uses union-find clustering with a hard limit of 25 columns defined in src/tables/grid.rs. Tables exceeding this width are skipped. Additionally, the histogram-valley heuristics in src/extractor/layout.rs may misidentify irregular or non-rectangular table structures. Very wide financial tables and complex multi-page layouts are most commonly affected.

Can I extract embedded images from PDFs using pdf-inspector?

No. The src/extractor/xobjects.rs module extracts only image placeholders—metadata rectangles indicating where images appear on pages. Actual pixel data, base64 encoding, and separate image file export are not implemented. For image extraction, use dedicated tools like pdf2image, Poppler, or PyMuPDF.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →