Limitations of pdf-inspector's PDF Analysis: Architecture, Constraints, and Workarounds

pdf-inspector deliberately enforces strict architectural boundaries—including a 1,000,000 operation limit per page, a 25-column cap on table detection, and mandatory external OCR for scanned documents—trading universal compatibility for edge-computing speed.

The firecrawl/pdf-inspector repository provides a fast, Rust-based PDF classification and extraction engine optimized for CLI, WebAssembly, and language bindings. While it efficiently converts text-based PDFs to structured Markdown, understanding the limitations of pdf-inspector's PDF analysis is critical for production deployments, as these constraints are intentional trade-offs that prioritize performance and dependency-free operation over handling every edge case.

Hard Architectural Limits in the Rust Core

Operation Count Ceiling (MAX_OPERATIONS)

To prevent denial-of-service from malicious or pathologically complex PDFs, the extractor hard-codes a maximum operation threshold. In [src/extractor/content_stream.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs#L56-L63), the constant MAX_OPERATIONS: usize = 1_000_000 defines the upper bound of PDF drawing and text operators processed per page. When a page exceeds this limit, the engine logs a warning and skips extraction for that specific page, returning None for its markdown content rather than hanging or crashing.

This safeguard means extremely dense vector graphics, complex CAD drawings, or obfuscated PDFs with millions of redundant operators will have gaps in the extracted text output.

Table Column Constraints (25-Column Limit)

The library explicitly refuses to parse tables wider than 25 columns. This constraint is documented in the project design notes within [AGENTS.md](https://github.com/firecrawl/pdf-inspector/blob/main/AGENTS.md#L78-L80), which states: "Column limit for tables: 25 (wide statistical tables)". The heuristic table detection logic in [src/tables/grid.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/grid.rs#L75-L82) enforces this boundary during the column-boundary analysis phase.

Consequently, wide statistical spreadsheets or financial reports with dozens of columns will not be detected as tables, falling back to plain text extraction that loses the tabular structure.

Detection and Extraction Boundaries

No Built-in OCR for Scanned PDFs

pdf-inspector is strictly a classifier and text extractor, not an OCR engine. When it detects rasterized or scanned content, it returns a pages_needing_ocr field in the result object rather than performing text recognition itself. As noted in the README, the detector "detects whether a PDF is text-based or scanned … returns a confidence score and per-page OCR routing", leaving the actual OCR step to external services like Tesseract or cloud vision APIs.

This architectural split means scanned PDFs will not produce Markdown text unless you implement a secondary OCR pipeline to handle the flagged pages.

Early-Exit Scan Strategy Trade-offs

By default, the detector uses ScanStrategy::EarlyExit, which stops analyzing pages at the first non-text (image or scan) detection. While this optimizes speed for pure-text documents, it can misclassify mixed PDFs. The README documents this behavior in the ScanStrategy table, showing that EarlyExit favors speed over accuracy for mixed content.

To analyze every page regardless of initial findings, you must explicitly switch to ScanStrategy::Full, incurring a performance penalty on large documents.

Layout and Formatting Restrictions

Heuristic Table Detection Gaps

Table detection relies primarily on rectangle-based analysis, falling back to heuristic column detection only when rectangles fail. The heuristic mode uses gap-histogram thresholds defined in [src/tables/grid.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/grid.rs#L30-L33), specifically a hard-coded range of 8–25 points for column spacing. PDFs with unconventional layouts—such as tables with extremely narrow gutters or irregular spacing—will bypass detection or be parsed with incorrect column boundaries, resulting in jumbled reading order.

Markdown Output Constraints

The Markdown converter supports only a fixed subset of elements: headings, lists, code blocks, tables, URLs, and basic formatting. Complex semantic structures like footnotes, endnotes, multi-column captions, or nested tables are flattened to plain text or simplified Markdown. This restriction ensures deterministic output but loses sophisticated document semantics present in the original PDF.

Code Examples: Handling Limitations Programmatically

The following examples demonstrate how to detect and work around pdf-inspector's limitations using the Rust API and Python bindings.

use pdf_inspector::{process_pdf_with_options, PdfOptions, ScanStrategy};

fn main() -> Result<(), Box<dyn std::error::Error>> {
    // Avoid early-exit to ensure all pages are analyzed for mixed content
    let opts = PdfOptions::default()
        .scan_strategy(ScanStrategy::Full);

    let result = process_pdf_with_options("complex-document.pdf", opts)?;

    // Handle OCR requirements
    if !result.pages_needing_ocr.is_empty() {
        eprintln!(
            "OCR required for pages: {:?}. Route to external OCR service.",
            result.pages_needing_ocr
        );
    }

    // Detect operation-limit skips (wide tables or complex graphics)
    if result.markdown.is_none() {
        eprintln!("Warning: Page skipped due to MAX_OPERATIONS limit (1,000,000)");
    } else {
        // Check for wide tables that may have exceeded the 25-column limit
        let md = result.markdown.unwrap();
        let pipe_count = md.matches('|').count();
        if pipe_count > 50 { // Rough heuristic: >25 columns = >50 pipes
            eprintln!("Warning: Detected table possibly exceeding 25-column limit");
        }
        println!("{}", md);
    }

    Ok(())
}
import pdf_inspector

# Process with full scan to avoid early-exit limitations

result = pdf_inspector.process_pdf(
    "scanned-mix.pdf",
    scan_strategy="Full"
)

# Check for OCR routing

if result.pages_needing_ocr:
    print(f"External OCR needed for pages: {result.pages_needing_ocr}")

# Check for extraction gaps (operation limit exceeded)

if result.markdown is None:
    print("Extraction failed: Page exceeded 1,000,000 operation limit")
else:
    print(result.markdown)

Summary

Frequently Asked Questions

Does pdf-inspector support OCR for scanned documents?

No. pdf-inspector only detects whether pages require OCR and returns their indices in the pages_needing_ocr field. You must implement a separate OCR pipeline (e.g., Tesseract, AWS Textract) to process these pages and insert the resulting text back into your workflow.

What happens when a PDF exceeds the operation limit?

When a single page contains more than 1,000,000 PDF operators, the extractor logs a warning and skips that page entirely, returning None for its Markdown content. This prevents resource exhaustion but creates gaps in the output for extremely complex vector graphics or obfuscated PDFs.

Why is there a 25-column limit for tables?

The 25-column limit is an intentional design constraint documented in [AGENTS.md](https://github.com/firecrawl/pdf-inspector/blob/main/AGENTS.md#L78-L80) to prevent performance degradation on wide statistical tables. The heuristic column-detection algorithms in [src/tables/grid.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/grid.rs#L75-L82) use this cap to maintain predictable memory usage and processing time.

Can pdf-inspector extract images from PDFs?

No. While the detector identifies pages containing images (using the Do operator) for classification purposes, it does not extract or output raster image data. The library focuses exclusively on text extraction; users requiring image extraction must use complementary PDF libraries.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →