Does Firecrawl PDF-Inspector Support OCR for Image-Based PDFs?
No, firecrawl pdf-inspector does not perform OCR itself—it detects when image-based PDFs need OCR and flags those pages for external processing.
The firecrawl/pdf-inspector repository is a Rust-based PDF text extraction library designed to identify when pages contain unreliable or unextractable text. Rather than bundling OCR functionality, it implements text-quality detection heuristics that determine which pages require OCR and exposes this information through a clean API for downstream processing.
How PDF-Inspector Handles Image-Based PDFs
PDF-inspector's architecture treats OCR as an external concern. The core library focuses exclusively on detecting problematic content and signaling where human-readable text cannot be directly extracted.
Text-Quality Detection Engine
The detection logic lives in src/text_quality.rs. This module analyzes extracted text for statistical and structural anomalies that indicate extraction failure:
- Replacement characters (
U+FFFD) — Unicode replacement glyphs signaling decode errors - "Dollar-as-space" patterns — Common artifact when ToUnicode mappings fail
- Letter frequency anomalies — Statistical deviations suggesting substitution-cipher garbling
When any heuristic triggers, the code invokes add_ocr_reason() and appends the page index to pages_needing_ocr.
Broken ToUnicode Detection
File src/tounicode.rs handles font-level failures. Fonts with missing or broken ToUnicode character maps cannot provide proper Unicode output—the module flags these with needs_ocr so the extractor knows to abandon text extraction for affected regions.
Result Propagation Through the Pipeline
The needs_ocr flag flows through src/extractor/mod.rs and src/extractor/layout.rs, ultimately attaching to each Page struct in the final output. Consumers receive structured data indicating exactly which pages failed extraction and why.
API Output: What You Receive
The public API defined in src/lib.rs returns a ProcessResult containing:
| Field | Type | Description |
|---|---|---|
pages_needing_ocr |
Vec<u32> |
1-based page numbers requiring OCR |
ocr_reasons_by_page |
optional mapping | Human-readable explanations per page |
This design lets callers make intelligent routing decisions—sending flagged pages to Tesseract, Google Vision, AWS Textract, or any preferred OCR engine.
Practical Usage Examples
Detect Pages Requiring OCR
# Run the standalone detector
detect-pdf --json my-document.pdf
Example output:
{
"pages_needing_ocr": [2, 5],
"ocr_reasons_by_page": {
"2": ["Identity-H font without ToUnicode"],
"5": ["Suspected garbled text"]
}
}
Extract with OCR Flags
# Convert to Markdown—flagged pages yield empty content
pdf2md --json my-document.pdf > out.json
Post-Process with External OCR
# Route flagged pages to Tesseract
for page in $(jq -r '.pages_needing_ocr[]' out.json); do
tesseract my-document.pdf[${page}] -l eng page-${page}.txt
done
Key Implementation Files
src/text_quality.rs— Heuristics engine for garbled-text detectionsrc/tounicode.rs— ToUnicode map validation and OCR flaggingsrc/lib.rs— Public API withProcessResultstructuressrc/bin/detect_pdf.rs— CLI detector entry pointsrc/bin/pdf2md.rs— Markdown extractor entry point
Summary
- PDF-inspector does not perform OCR—it detects when OCR is necessary
- Detection covers broken fonts, replacement characters, and statistical text anomalies
- Flagged pages are exposed via
pages_needing_ocrin the API response - External OCR tools (Tesseract, cloud APIs) handle the actual image-to-text conversion
- This separation keeps the core library lightweight and engine-agnostic
Frequently Asked Questions
Does pdf-inspector include built-in OCR?
No. According to the firecrawl/pdf-inspector source code, the library deliberately excludes OCR functionality. It identifies problematic pages through heuristics in src/text_quality.rs and leaves OCR execution to external tools.
What triggers the needs_ocr flag?
Three primary conditions: fonts lacking ToUnicode maps (handled in src/tounicode.rs), Unicode replacement characters in output, and statistical letter-frequency anomalies suggesting garbled encoding. Each triggers add_ocr_reason() and populates pages_needing_ocr.
How do I actually OCR pages flagged by pdf-inspector?
Extract the pages_needing_ocr array from the ProcessResult, then route those page numbers to any OCR engine. The repository provides no preference—Tesseract, Google Vision, Azure AI, and AWS Textract are all valid choices depending on your accuracy and latency requirements.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →