How Per-Page OCR Routing Works with `pages_needing_ocr` in PDF Inspector

PDF Inspector uses a boolean vector produced by the detector to route only image-based pages to OCR while processing searchable text directly from pages that already contain extractable Unicode content.

The firecrawl/pdf-inspector repository implements intelligent per-page OCR routing to minimize processing costs when converting mixed PDFs to Markdown. By analyzing each page individually during the detection phase, the system generates a pages_needing_ocr bitmap that drives conditional routing in the extraction pipeline. This architecture ensures that computationally expensive OCR operations are only invoked for pages that actually contain scanned images rather than searchable text.

Detection Phase: Generating the pages_needing_ocr Bitmap

The detection logic resides in [src/detector.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs), where the detect_pdf_type function scans every page to classify the document as TextBased, Scanned, or Mixed. For mixed documents, the detector returns PdfType::Mixed(Vec<bool>), where the vector serves as the pages_needing_ocr flag—each boolean indicates whether the corresponding page lacks extractable text and requires OCR processing.

The PdfType Enum Classification

According to the source code, the detector categorizes documents into three distinct types. The Mixed variant specifically carries a Vec<bool> payload that maps directly to the pages_needing_ocr concept, allowing downstream components to make per-page routing decisions without re-analyzing the content.

Extraction Orchestration: Conditional Routing Logic

The extraction orchestrator in [src/extractor/mod.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) consumes the detection result and implements the routing logic. When processing a Mixed PDF, the system iterates over the pages_needing_ocr vector: pages marked false are processed through the standard content stream parser, while pages marked true are routed to the OCR pipeline.

Normal Text Extraction vs. OCR Pathways

Direct extraction leverages [src/extractor/content_stream.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs) to parse PDF content streams when the boolean flag is false. OCR routing occurs when the flag is true, triggering the OCR pipeline (via external tools or the optional pdf-ocr crate) to extract text from raster images before integrating the results back into the document flow.

Command-Line and Programmatic Usage

The pdf2md binary in [src/bin/pdf2md.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) exposes this functionality through a simple CLI interface that automatically handles per-page OCR routing based on the detector's analysis.

Run the converter from the command line:

pdf2md mixed_document.pdf --output out.md

Access the routing flags programmatically via the Rust API:

use pdf_inspector::{process_pdf_with_options, ProcessOptions, PdfType};

let opts = ProcessOptions::default();
let result = process_pdf_with_options("mixed.pdf", opts).unwrap();

if let PdfType::Mixed(pages_needing_ocr) = result.pdf_type {
    for (i, needs_ocr) in pages_needing_ocr.iter().enumerate() {
        println!("Page {} needs OCR: {}", i + 1, needs_ocr);
    }
}

Summary

Frequently Asked Questions

What data structure does pages_needing_ocr use?

The system implements pages_needing_ocr as a Vec<bool> (boolean vector) inside the PdfType::Mixed variant returned by the detector. Each index corresponds to a page number (0-indexed), with true values indicating pages that contain only raster images and require OCR processing.

How does the extractor decide whether to run OCR on a specific page?

The extraction orchestrator checks the boolean value at the corresponding index in the pages_needing_ocr vector. If the value is true, the page is routed to the OCR pipeline; if false, the system processes the page using the standard content stream extractor to pull text directly from the PDF structure.

Can I access the pages_needing_ocr flags when using the CLI tool?

While the pdf2md CLI in [src/bin/pdf2md.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) handles routing automatically without exposing the internal bitmap, you can access the flags programmatically through the Rust API by matching on PdfType::Mixed(pages_needing_ocr) after calling process_pdf_with_options.

Which source files handle the routing decision logic?

The routing logic spans two primary files: [src/detector.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) generates the classification and bitmap, while [src/extractor/mod.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) implements the conditional logic that directs pages to either [src/extractor/content_stream.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs) (normal extraction) or the OCR pipeline based on the bitmap values.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →