# Firecrawl pdf-inspector Limitations: 9 Known Constraints Explained

> Discover the limitations of Firecrawl pdf-inspector. Learn about its constraints with scanned PDFs, OCR, table detection, password protection, and parallel processing.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: limitations
- Published: 2026-08-07

---

**Firecrawl pdf-inspector cannot extract text from scanned PDFs, has no OCR capability, caps table detection at 25 columns, and lacks support for password-protected documents or parallel processing.**

Firecrawl pdf-inspector is a fast, pure-Rust PDF classifier and text-extraction library designed for speed and minimal dependencies. While it excels at converting native-text PDFs to clean Markdown, its lightweight architecture imposes specific boundaries that users should understand before adoption. This guide examines each limitation with direct references to the source code implementation.

## No OCR Capability for Scanned Documents

The most significant limitation of firecrawl pdf-inspector is its **complete absence of OCR functionality**. The library can only classify PDFs—it cannot extract readable text from scanned pages.

The detection logic in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) performs fast, sampling-only inspection of text operators (`Tj`, `TJ`) and image operators (`Do`). It returns a confidence score and a list of pages requiring OCR, but performs no actual text recognition:

```rust
// From src/lib.rs - document loading rejects encrypted files
pub fn process_pdf(path: &str) -> Result<PdfResult, Error> {
    // Detection identifies scanned pages but cannot process them
    let detection = detector::analyze_pdf(&doc)?;
    if detection.pdf_type == PdfType::Scanned {
        // Returns classification only, no extracted text for these pages
    }
}

```

When `result.pdf_type == "scanned"`, users must route pages to an external OCR service. The `pages_needing_ocr` field indicates exactly which pages require processing:

```python
import pdf_inspector

result = pdf_inspector.process_pdf("sample.pdf")

if result.pdf_type == "scanned":
    print("OCR required for pages:", result.pages_needing_ocr)
    # External OCR step mandatory here

```

## Hard-Capped Table Detection Width

Table extraction in firecrawl pdf-inspector is constrained by a **25-column maximum** enforced in [`src/tables/grid.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/grid.rs). The union-find clustering algorithm skips tables exceeding this width, and heuristic detection may fail on irregular layouts.

This architectural limit affects:

- Wide financial statements with many columns
- Complex multi-page tables with varying structures
- Irregular grid layouts that deviate from rectangular assumptions

The column detection uses histogram-valley heuristics with fixed thresholds, as implemented in [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs). Very wide tables are partially rendered or excluded entirely.

## Dependency on lopdf Parser Reliability

pdf-inspector relies exclusively on the `lopdf` crate for low-level PDF parsing, as specified in [`Cargo.toml`](https://github.com/firecrawl/pdf-inspector/blob/main/Cargo.toml). This single-dependency design creates a **single point of failure**:

- Malformed PDFs with corrupted xref tables trigger unrecoverable errors
- Unusual object streams cause parse failures
- Encrypted documents abort the entire pipeline

There is no built-in fallback mechanism. When `lopdf` fails, extraction terminates immediately without partial results or recovery options.

## No Password-Protected PDF Support

The library explicitly rejects encrypted documents. The document loading code in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) checks for encryption and returns an error:

```rust
// Simplified from src/lib.rs
if doc.is_encrypted() {
    return Err(Error::EncryptedPdf);
}

```

Users must decrypt PDFs before processing, adding a preprocessing step to any pipeline handling protected documents.

## Limited Right-to-Left and CJK Script Handling

While [`src/text_utils.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_utils.rs) provides basic RTL detection and CJK utilities, the implementation lacks full Unicode support:

| Script Type | Support Level | Limitation |
|-------------|-------------|------------|
| Basic RTL | Partial | Detection exists but no complex shaping |
| CJK | Helper functions only | No full line-breaking or substitution |
| Mixed directionality | Limited | Direction changes may garble output |
| Exotic scripts | None | Font substitution unavailable |

Complex scripts requiring glyph-level shaping or font substitution lose fidelity during extraction.

## Static Column-Detection Thresholds

The layout analysis in [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs) employs histogram-valley column detection with **fixed thresholds**. This heuristic approach struggles with:

- Newspaper-style PDFs with irregular column widths
- Non-rectilinear column layouts
- Highly variable page designs within single documents

Misdetected columns produce incorrect reading order, corrupting downstream Markdown structure.

## No Actual Image Content Extraction

Images and vector graphics receive placeholder treatment only. The [`src/extractor/xobjects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/xobjects.rs) module extracts markers indicating image presence, but **never exports actual pixel data**:

```rust
// From src/extractor/xobjects.rs - placeholder extraction only
pub fn extract_images(page: &Page) -> Vec<ImagePlaceholder> {
    // Returns metadata, not image bytes
    vec![ImagePlaceholder { rect, .. }]
}

```

Use cases requiring embedded figures, diagrams, or base64-encoded images must implement separate image handling pipelines.

## Single-Threaded Processing Architecture

The `process_pdf` function in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) operates **synchronously without internal parallelism**:

```rust
pub fn process_pdf(path: &str) -> Result<PdfResult, Error> {
    // Sequential single-pass processing
    let doc = Document::load(path)?;
    let detection = detector::analyze_pdf(&doc)?;
    let markdown = extractor::extract_markdown(&doc, &options)?;
    // No Tokio, no Rayon, no thread spawning
}

```

Large documents exceeding 500 pages cannot leverage multi-core processing. Users must implement external parallelism at the file level if batch throughput is critical.

## Limited Configurability via PdfOptions

The `PdfOptions` builder in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) exposes coarse controls like scan strategy, but **fine-grained tuning is unavailable**:

- Table heuristic thresholds are hardcoded in [`src/tables/grid.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/grid.rs)
- Heading detection parameters fixed in layout engine
- Font-size clustering boundaries not user-adjustable

Advanced customization requires source modification and recompilation.

## Summary

- **No OCR capability** — scanned pages require external processing; the library only classifies PDF types
- **25-column table limit** — wide tables skipped by union-find clustering in [`src/tables/grid.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/grid.rs)
- **lopdf dependency fragility** — malformed or encrypted PDFs cause complete pipeline failure
- **Password-protected PDF rejection** — encrypted documents return errors from [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)
- **Incomplete script support** — RTL and CJK handling lacks full shaping and line-breaking
- **Fixed layout heuristics** — histogram-valley column detection fails on irregular layouts
- **Placeholder-only images** — [`src/extractor/xobjects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/xobjects.rs) never exports actual image data
- **Synchronous single-threading** — no internal parallelism for large document processing
- **Restricted customization** — `PdfOptions` lacks fine-grained control over extraction parameters

Firecrawl pdf-inspector prioritizes speed, determinism, and minimal dependencies over universal PDF handling. It excels with native-text documents but requires complementary tools for OCR, decryption, image extraction, and complex layout analysis.

## Frequently Asked Questions

### Does firecrawl pdf-inspector support OCR for scanned PDFs?

No. The library can detect that a PDF is scanned and identify which pages need OCR via `pages_needing_ocr`, but it cannot perform text extraction on those pages itself. Users must integrate with external OCR services like Tesseract, AWS Textract, or Azure AI Document Intelligence for actual text recognition.

### What happens when pdf-inspector encounters a password-protected PDF?

The library rejects encrypted documents immediately. The document loading code in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) checks for the encryption flag and returns an `Error::EncryptedPdf` before any processing begins. Users must decrypt files using `qpdf`, `pdftk`, or similar tools before passing them to pdf-inspector.

### Why are some tables missing or malformed in the Markdown output?

Table detection uses union-find clustering with a hard limit of 25 columns defined in [`src/tables/grid.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/grid.rs). Tables exceeding this width are skipped. Additionally, the histogram-valley heuristics in [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs) may misidentify irregular or non-rectangular table structures. Very wide financial tables and complex multi-page layouts are most commonly affected.

### Can I extract embedded images from PDFs using pdf-inspector?

No. The [`src/extractor/xobjects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/xobjects.rs) module extracts only image placeholders—metadata rectangles indicating where images appear on pages. Actual pixel data, base64 encoding, and separate image file export are not implemented. For image extraction, use dedicated tools like `pdf2image`, `Poppler`, or `PyMuPDF`.