# How Per-Page OCR Routing Works with `pages_needing_ocr` in PDF Inspector

> Learn how PDF Inspector optimizes OCR with pages_needing_ocr, routing only image-based pages for efficient text extraction and processing searchable content directly.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-08

---

**PDF Inspector uses a boolean vector produced by the detector to route only image-based pages to OCR while processing searchable text directly from pages that already contain extractable Unicode content.**

The `firecrawl/pdf-inspector` repository implements intelligent per-page OCR routing to minimize processing costs when converting mixed PDFs to Markdown. By analyzing each page individually during the detection phase, the system generates a `pages_needing_ocr` bitmap that drives conditional routing in the extraction pipeline. This architecture ensures that computationally expensive OCR operations are only invoked for pages that actually contain scanned images rather than searchable text.

## Detection Phase: Generating the `pages_needing_ocr` Bitmap

The detection logic resides in [[`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs), where the `detect_pdf_type` function scans every page to classify the document as **TextBased**, **Scanned**, or **Mixed**. For mixed documents, the detector returns `PdfType::Mixed(Vec<bool>)`, where the vector serves as the `pages_needing_ocr` flag—each boolean indicates whether the corresponding page lacks extractable text and requires OCR processing.

### The `PdfType` Enum Classification

According to the source code, the detector categorizes documents into three distinct types. The **Mixed** variant specifically carries a `Vec<bool>` payload that maps directly to the `pages_needing_ocr` concept, allowing downstream components to make per-page routing decisions without re-analyzing the content.

## Extraction Orchestration: Conditional Routing Logic

The extraction orchestrator in [[`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) consumes the detection result and implements the routing logic. When processing a **Mixed** PDF, the system iterates over the `pages_needing_ocr` vector: pages marked `false` are processed through the standard content stream parser, while pages marked `true` are routed to the OCR pipeline.

### Normal Text Extraction vs. OCR Pathways

**Direct extraction** leverages [[`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs) to parse PDF content streams when the boolean flag is `false`. **OCR routing** occurs when the flag is `true`, triggering the OCR pipeline (via external tools or the optional `pdf-ocr` crate) to extract text from raster images before integrating the results back into the document flow.

## Command-Line and Programmatic Usage

The `pdf2md` binary in [[`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) exposes this functionality through a simple CLI interface that automatically handles per-page OCR routing based on the detector's analysis.

Run the converter from the command line:

```bash
pdf2md mixed_document.pdf --output out.md

```

Access the routing flags programmatically via the Rust API:

```rust
use pdf_inspector::{process_pdf_with_options, ProcessOptions, PdfType};

let opts = ProcessOptions::default();
let result = process_pdf_with_options("mixed.pdf", opts).unwrap();

if let PdfType::Mixed(pages_needing_ocr) = result.pdf_type {
    for (i, needs_ocr) in pages_needing_ocr.iter().enumerate() {
        println!("Page {} needs OCR: {}", i + 1, needs_ocr);
    }
}

```

## Summary

- Detection occurs in [[`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs), which returns a `PdfType::Mixed(Vec<bool>)` structure representing the `pages_needing_ocr` bitmap.
- The extraction orchestrator in [[`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) uses this bitmap to route pages to either direct text extraction or OCR processing.
- This per-page granularity minimizes OCR costs by only processing pages that contain scanned images.
- Final output generation in [[`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) seamlessly merges OCR-derived text with standard extraction results.

## Frequently Asked Questions

### What data structure does `pages_needing_ocr` use?

The system implements `pages_needing_ocr` as a `Vec<bool>` (boolean vector) inside the `PdfType::Mixed` variant returned by the detector. Each index corresponds to a page number (0-indexed), with `true` values indicating pages that contain only raster images and require OCR processing.

### How does the extractor decide whether to run OCR on a specific page?

The extraction orchestrator checks the boolean value at the corresponding index in the `pages_needing_ocr` vector. If the value is `true`, the page is routed to the OCR pipeline; if `false`, the system processes the page using the standard content stream extractor to pull text directly from the PDF structure.

### Can I access the `pages_needing_ocr` flags when using the CLI tool?

While the `pdf2md` CLI in [[`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) handles routing automatically without exposing the internal bitmap, you can access the flags programmatically through the Rust API by matching on `PdfType::Mixed(pages_needing_ocr)` after calling `process_pdf_with_options`.

### Which source files handle the routing decision logic?

The routing logic spans two primary files: [[`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) generates the classification and bitmap, while [[`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) implements the conditional logic that directs pages to either [[`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs) (normal extraction) or the OCR pipeline based on the bitmap values.