# Does Firecrawl PDF-Inspector Support OCR for Image-Based PDFs?

> Firecrawl PDF-Inspector detects image-based PDFs and flags pages needing OCR for external processing. Learn how it handles OCR-requiring documents.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-07

---

**No, firecrawl pdf-inspector does not perform OCR itself—it detects when image-based PDFs need OCR and flags those pages for external processing.**

The firecrawl/pdf-inspector repository is a Rust-based PDF text extraction library designed to identify when pages contain unreliable or unextractable text. Rather than bundling OCR functionality, it implements **text-quality detection heuristics** that determine which pages require OCR and exposes this information through a clean API for downstream processing.

## How PDF-Inspector Handles Image-Based PDFs

PDF-inspector's architecture treats OCR as an external concern. The core library focuses exclusively on detecting problematic content and signaling where human-readable text cannot be directly extracted.

### Text-Quality Detection Engine

The detection logic lives in [`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs). This module analyzes extracted text for statistical and structural anomalies that indicate extraction failure:

- **Replacement characters (`U+FFFD`)** — Unicode replacement glyphs signaling decode errors
- **"Dollar-as-space" patterns** — Common artifact when ToUnicode mappings fail
- **Letter frequency anomalies** — Statistical deviations suggesting substitution-cipher garbling

When any heuristic triggers, the code invokes `add_ocr_reason()` and appends the page index to `pages_needing_ocr`.

### Broken ToUnicode Detection

File [`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs) handles font-level failures. Fonts with missing or broken ToUnicode character maps cannot provide proper Unicode output—the module flags these with `needs_ocr` so the extractor knows to abandon text extraction for affected regions.

### Result Propagation Through the Pipeline

The `needs_ocr` flag flows through [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) and [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs), ultimately attaching to each `Page` struct in the final output. Consumers receive structured data indicating exactly which pages failed extraction and why.

## API Output: What You Receive

The public API defined in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) returns a `ProcessResult` containing:

| Field | Type | Description |
|-------|------|-------------|
| `pages_needing_ocr` | `Vec<u32>` | 1-based page numbers requiring OCR |
| `ocr_reasons_by_page` | optional mapping | Human-readable explanations per page |

This design lets callers make intelligent routing decisions—sending flagged pages to Tesseract, Google Vision, AWS Textract, or any preferred OCR engine.

## Practical Usage Examples

### Detect Pages Requiring OCR

```bash

# Run the standalone detector

detect-pdf --json my-document.pdf

```

Example output:

```json
{
  "pages_needing_ocr": [2, 5],
  "ocr_reasons_by_page": {
    "2": ["Identity-H font without ToUnicode"],
    "5": ["Suspected garbled text"]
  }
}

```

### Extract with OCR Flags

```bash

# Convert to Markdown—flagged pages yield empty content

pdf2md --json my-document.pdf > out.json

```

### Post-Process with External OCR

```bash

# Route flagged pages to Tesseract

for page in $(jq -r '.pages_needing_ocr[]' out.json); do
    tesseract my-document.pdf[${page}] -l eng page-${page}.txt
done

```

## Key Implementation Files

- **[`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs)** — Heuristics engine for garbled-text detection
- **[`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs)** — ToUnicode map validation and OCR flagging
- **[`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)** — Public API with `ProcessResult` structures
- **[`src/bin/detect_pdf.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs)** — CLI detector entry point
- **[`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs)** — Markdown extractor entry point

## Summary

- **PDF-inspector does not perform OCR**—it detects when OCR is necessary
- Detection covers broken fonts, replacement characters, and statistical text anomalies
- Flagged pages are exposed via `pages_needing_ocr` in the API response
- External OCR tools (Tesseract, cloud APIs) handle the actual image-to-text conversion
- This separation keeps the core library lightweight and engine-agnostic

## Frequently Asked Questions

### Does pdf-inspector include built-in OCR?

No. According to the firecrawl/pdf-inspector source code, the library deliberately excludes OCR functionality. It identifies problematic pages through heuristics in [`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs) and leaves OCR execution to external tools.

### What triggers the `needs_ocr` flag?

Three primary conditions: fonts lacking ToUnicode maps (handled in [`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs)), Unicode replacement characters in output, and statistical letter-frequency anomalies suggesting garbled encoding. Each triggers `add_ocr_reason()` and populates `pages_needing_ocr`.

### How do I actually OCR pages flagged by pdf-inspector?

Extract the `pages_needing_ocr` array from the `ProcessResult`, then route those page numbers to any OCR engine. The repository provides no preference—Tesseract, Google Vision, Azure AI, and AWS Textract are all valid choices depending on your accuracy and latency requirements.