# Limitations of pdf-inspector's PDF Analysis: Architecture, Constraints, and Workarounds

> Discover the limitations of pdf-inspector's PDF analysis. Learn about its architectural constraints, like the 1M operation limit and 25-column cap, and explore available workarounds for faster edge-computing.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: architecture
- Published: 2026-08-04

---

**pdf-inspector deliberately enforces strict architectural boundaries—including a 1,000,000 operation limit per page, a 25-column cap on table detection, and mandatory external OCR for scanned documents—trading universal compatibility for edge-computing speed.**

The `firecrawl/pdf-inspector` repository provides a fast, Rust-based PDF classification and extraction engine optimized for CLI, WebAssembly, and language bindings. While it efficiently converts text-based PDFs to structured Markdown, understanding the **limitations of pdf-inspector's PDF analysis** is critical for production deployments, as these constraints are intentional trade-offs that prioritize performance and dependency-free operation over handling every edge case.

## Hard Architectural Limits in the Rust Core

### Operation Count Ceiling (MAX_OPERATIONS)

To prevent denial-of-service from malicious or pathologically complex PDFs, the extractor hard-codes a maximum operation threshold. In [[`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs#L56-L63), the constant `MAX_OPERATIONS: usize = 1_000_000` defines the upper bound of PDF drawing and text operators processed per page. When a page exceeds this limit, the engine logs a warning and skips extraction for that specific page, returning `None` for its markdown content rather than hanging or crashing.

This safeguard means extremely dense vector graphics, complex CAD drawings, or obfuscated PDFs with millions of redundant operators will have gaps in the extracted text output.

### Table Column Constraints (25-Column Limit)

The library explicitly refuses to parse tables wider than 25 columns. This constraint is documented in the project design notes within [[`AGENTS.md`](https://github.com/firecrawl/pdf-inspector/blob/main/AGENTS.md)](https://github.com/firecrawl/pdf-inspector/blob/main/AGENTS.md#L78-L80), which states: *"Column limit for tables: 25 (wide statistical tables)"*. The heuristic table detection logic in [[`src/tables/grid.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/grid.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/grid.rs#L75-L82) enforces this boundary during the column-boundary analysis phase.

Consequently, wide statistical spreadsheets or financial reports with dozens of columns will not be detected as tables, falling back to plain text extraction that loses the tabular structure.

## Detection and Extraction Boundaries

### No Built-in OCR for Scanned PDFs

pdf-inspector is strictly a classifier and text extractor, not an OCR engine. When it detects rasterized or scanned content, it returns a `pages_needing_ocr` field in the result object rather than performing text recognition itself. As noted in the [README](https://github.com/firecrawl/pdf-inspector/blob/main/README.md#L14-L22), the detector *"detects whether a PDF is text-based or scanned … returns a confidence score and per-page OCR routing"*, leaving the actual OCR step to external services like Tesseract or cloud vision APIs.

This architectural split means scanned PDFs will not produce Markdown text unless you implement a secondary OCR pipeline to handle the flagged pages.

### Early-Exit Scan Strategy Trade-offs

By default, the detector uses `ScanStrategy::EarlyExit`, which stops analyzing pages at the first non-text (image or scan) detection. While this optimizes speed for pure-text documents, it can misclassify mixed PDFs. The [README](https://github.com/firecrawl/pdf-inspector/blob/main/README.md#L24-L31) documents this behavior in the `ScanStrategy` table, showing that `EarlyExit` favors speed over accuracy for mixed content.

To analyze every page regardless of initial findings, you must explicitly switch to `ScanStrategy::Full`, incurring a performance penalty on large documents.

## Layout and Formatting Restrictions

### Heuristic Table Detection Gaps

Table detection relies primarily on rectangle-based analysis, falling back to heuristic column detection only when rectangles fail. The heuristic mode uses gap-histogram thresholds defined in [[`src/tables/grid.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/grid.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/grid.rs#L30-L33), specifically a hard-coded range of 8–25 points for column spacing. PDFs with unconventional layouts—such as tables with extremely narrow gutters or irregular spacing—will bypass detection or be parsed with incorrect column boundaries, resulting in jumbled reading order.

### Markdown Output Constraints

The Markdown converter supports only a fixed subset of elements: headings, lists, code blocks, tables, URLs, and basic formatting. Complex semantic structures like footnotes, endnotes, multi-column captions, or nested tables are flattened to plain text or simplified Markdown. This restriction ensures deterministic output but loses sophisticated document semantics present in the original PDF.

## Code Examples: Handling Limitations Programmatically

The following examples demonstrate how to detect and work around pdf-inspector's limitations using the Rust API and Python bindings.

```rust
use pdf_inspector::{process_pdf_with_options, PdfOptions, ScanStrategy};

fn main() -> Result<(), Box<dyn std::error::Error>> {
    // Avoid early-exit to ensure all pages are analyzed for mixed content
    let opts = PdfOptions::default()
        .scan_strategy(ScanStrategy::Full);

    let result = process_pdf_with_options("complex-document.pdf", opts)?;

    // Handle OCR requirements
    if !result.pages_needing_ocr.is_empty() {
        eprintln!(
            "OCR required for pages: {:?}. Route to external OCR service.",
            result.pages_needing_ocr
        );
    }

    // Detect operation-limit skips (wide tables or complex graphics)
    if result.markdown.is_none() {
        eprintln!("Warning: Page skipped due to MAX_OPERATIONS limit (1,000,000)");
    } else {
        // Check for wide tables that may have exceeded the 25-column limit
        let md = result.markdown.unwrap();
        let pipe_count = md.matches('|').count();
        if pipe_count > 50 { // Rough heuristic: >25 columns = >50 pipes
            eprintln!("Warning: Detected table possibly exceeding 25-column limit");
        }
        println!("{}", md);
    }

    Ok(())
}

```

```python
import pdf_inspector

# Process with full scan to avoid early-exit limitations

result = pdf_inspector.process_pdf(
    "scanned-mix.pdf",
    scan_strategy="Full"
)

# Check for OCR routing

if result.pages_needing_ocr:
    print(f"External OCR needed for pages: {result.pages_needing_ocr}")

# Check for extraction gaps (operation limit exceeded)

if result.markdown is None:
    print("Extraction failed: Page exceeded 1,000,000 operation limit")
else:
    print(result.markdown)

```

## Summary

- **Operation Limit**: Hard cap of 1,000,000 PDF operators per page in [[`content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/content_stream.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs#L56-L63); exceeding pages return no text.
- **Table Width**: Maximum 25 columns enforced in [[`AGENTS.md`](https://github.com/firecrawl/pdf-inspector/blob/main/AGENTS.md)](https://github.com/firecrawl/pdf-inspector/blob/main/AGENTS.md#L78-L80) and [[`grid.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/grid.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/grid.rs#L75-L82); wider tables parse as plain text.
- **OCR Requirement**: Scanned pages are flagged in `pages_needing_ocr` but not processed; external OCR is mandatory.
- **Scan Strategy**: Default `EarlyExit` may miss mixed content; use `Full` for complete analysis at a performance cost.
- **Layout Heuristics**: Column detection relies on 8–25 point gap thresholds; atypical spacing causes mis-detection.
- **Markdown Scope**: Limited to basic structural elements; complex formatting is flattened.

## Frequently Asked Questions

### Does pdf-inspector support OCR for scanned documents?

No. pdf-inspector only detects whether pages require OCR and returns their indices in the `pages_needing_ocr` field. You must implement a separate OCR pipeline (e.g., Tesseract, AWS Textract) to process these pages and insert the resulting text back into your workflow.

### What happens when a PDF exceeds the operation limit?

When a single page contains more than 1,000,000 PDF operators, the extractor logs a warning and skips that page entirely, returning `None` for its Markdown content. This prevents resource exhaustion but creates gaps in the output for extremely complex vector graphics or obfuscated PDFs.

### Why is there a 25-column limit for tables?

The 25-column limit is an intentional design constraint documented in [[`AGENTS.md`](https://github.com/firecrawl/pdf-inspector/blob/main/AGENTS.md)](https://github.com/firecrawl/pdf-inspector/blob/main/AGENTS.md#L78-L80) to prevent performance degradation on wide statistical tables. The heuristic column-detection algorithms in [[`src/tables/grid.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/grid.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/grid.rs#L75-L82) use this cap to maintain predictable memory usage and processing time.

### Can pdf-inspector extract images from PDFs?

No. While the detector identifies pages containing images (using the `Do` operator) for classification purposes, it does not extract or output raster image data. The library focuses exclusively on text extraction; users requiring image extraction must use complementary PDF libraries.