# Region-Based Text Extraction in Hybrid OCR Pipelines: A Technical Guide to pdf-inspector

> Learn about region-based text extraction in hybrid OCR pipelines. This technical guide explains how pdf-inspector processes specific PDF areas for efficient OCR and preserves native text.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-08

---

**Region-based text extraction enables OCR systems to process only specific bounding-box areas of a PDF page, allowing hybrid pipelines to fall back to OCR solely for corrupted regions while preserving clean native text elsewhere.**

The `firecrawl/pdf-inspector` crate implements sophisticated region-based text extraction that serves as the foundation for efficient hybrid OCR pipelines. Unlike traditional approaches that process entire pages uniformly, this Rust library extracts text from user-defined rectangular regions and independently assesses text quality within each zone. This granular approach minimizes unnecessary OCR processing while maximizing accuracy when dealing with PDFs containing mixed or corrupted text layers.

## What Is Region-Based Text Extraction?

Region-based text extraction restricts PDF parsing to specific bounding-box coordinates rather than scanning entire pages. In pdf-inspector, the `extract_text_in_regions_mem` function accepts a slice of page-region pairs, allowing developers to target precise areas such as headers, tables, or specific paragraphs while ignoring irrelevant content.

The public API signature defined in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) exposes this capability:

```rust
pub fn extract_text_in_regions_mem(
    buffer: &[u8],
    page_regions: &[(u32, Vec<[f32; 4]>)],
) -> Result<Vec<PageRegionResult>, PdfError>

```

Each region is defined as `[x1, y1, x2, y2]` in PDF units, representing lower-left to upper-right coordinates.

### Region Assignment Logic

When processing regions, the engine iterates over extracted `TextItem`s and assigns each to the region with the **largest overlap area** (`region_item_overlap_area`). This logic, implemented in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) (lines 792-815), produces a `Vec<Vec<TextItem>>` where each inner vector contains items belonging to a specific region. If a text item overlaps multiple regions, it joins only the one where it occupies the greatest area, ensuring clean separation of content.

## How Hybrid OCR Pipelines Leverage Region-Based Extraction

The true power of region-based extraction emerges in **hybrid OCR workflows** that combine native PDF text extraction with selective OCR fallback. Rather than applying OCR to entire pages when corruption is detected, the pipeline treats each region independently.

### Per-Region Quality Assessment

After grouping text items by region, pdf-inspector runs three independent quality detectors on each zone in [`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs):

- **`region_items_have_decoding_issue`** – Checks for broken CID→Unicode mappings at the `TextItem` level (line 303)
- **`is_cid_garbage`** – Identifies unresolvable glyph-ID fonts that indicate font encoding failures
- **`detect_encoding_issues`** – Scans rendered markdown for replacement characters () and "dollar-as-space" patterns that signal encoding problems

If any detector flags issues, the region receives an `ocr_reason` and is marked with `needs_ocr: true`.

### Selective OCR Workflow

A typical hybrid pipeline follows three distinct steps:

1. **Region extraction** – Call `extract_text_in_regions_mem` with target rectangles
2. **Quality assessment** – The library automatically evaluates text trustworthiness per region using the three heuristics
3. **Conditional OCR** – Invoke OCR engines (e.g., GPU-accelerated Tesseract) only on regions that failed quality checks, preserving clean text from other areas

This selective strategy dramatically reduces processing time and computational costs while handling PDFs with mixed text layers.

## Implementation in pdf-inspector

### Core API Methods

The primary entry points for region-based extraction are `extract_text_in_regions_mem` for text and `extract_tables_in_regions_mem` for structured data. Both functions accept PDF bytes and structured page-region definitions, returning `PageRegionResult` structs that include extracted content and quality metadata.

### Table Detection Integration

When the caller requests tables via `extract_tables_in_regions_mem`, the same region-assignment logic applies before invoking detection modules in [`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs). This ensures table extraction also benefits from the hybrid OCR fallback if the text inside a specific region is garbled, maintaining consistency across text and table processing pipelines.

### Source File Architecture

| File | Role |
|------|------|
| [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) | Coordinates region-wise item assignment, invokes per-region quality detectors, builds `PageRegionResult` |
| [`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs) | Implements the three text-quality heuristics (`region_items_have_decoding_issue`, `detect_encoding_issues`, `is_cid_garbage`) |
| [`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs) | Provides table-region detection that works on the same bounding-box infrastructure |
| [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) | Core PDF content-stream extraction that produces the raw `TextItem`s later grouped by region |

## Practical Code Example

The following Rust example demonstrates extracting text from two specific regions on page 1, then conditionally applying OCR based on quality flags:

```rust
use pdf_inspector::extract_text_in_regions_mem;
use pdf_inspector::PdfError;

fn main() -> Result<(), PdfError> {
    // Load a PDF file into memory
    let pdf_bytes = std::fs::read("sample.pdf")?;

    // Define two rectangular regions on page 1 (coordinates are PDF units)
    // Format: (x1, y1, x2, y2) – lower‑left → upper‑right.
    let regions = vec![
        (1, vec![[50.0, 700.0, 300.0, 750.0],   // Region A
                 [50.0, 600.0, 300.0, 650.0]]), // Region B
    ];

    // Extract text only from those rectangles
    let results = extract_text_in_regions_mem(&pdf_bytes, &regions)?;

    for page in results {
        for (idx, region) in page.regions.iter().enumerate() {
            println!("Page {}, Region {}:", page.page + 1, idx + 1);
            println!("  Text: {}", region.text);
            if region.needs_ocr {
                println!("  → OCR required (reason: {:?})", region.ocr_reason);
            }
        }
    }

    Ok(())
}

```

For regions flagged as needing OCR, you would typically rasterize only that specific area and process it through your OCR engine:

```rust
if region.needs_ocr {
    let image = rasterize_page_region(pdf_bytes, page.page, region_bbox)?;
    let ocr_text = run_gpu_ocr(&image)?;
    // Merge OCR result back into the final markdown.
}

```

## Summary

- **Region-based text extraction** targets specific bounding boxes rather than entire pages, reducing processing overhead and enabling surgical text recovery
- The `extract_text_in_regions_mem` API in pdf-inspector enables precise extraction from user-defined rectangles with coordinates in PDF units
- Three quality heuristics (`region_items_have_decoding_issue`, `is_cid_garbage`, `detect_encoding_issues`) determine OCR necessity independently for each region
- **Hybrid pipelines** avoid unnecessary OCR costs by preserving clean native text and only processing corrupted regions through expensive OCR engines
- Table extraction via `extract_tables_in_regions_mem` inherits these per-region quality checks for consistent hybrid processing across content types

## Frequently Asked Questions

### How does region-based extraction differ from full-page OCR?

Region-based extraction processes only user-defined rectangular areas within a PDF page, whereas full-page OCR rasterizes and recognizes text from the entire page. According to the pdf-inspector source code, this approach allows hybrid pipelines to preserve high-quality native text in clean regions while applying OCR only to specific zones where the native text layer is corrupted, significantly reducing computational costs.

### What triggers the OCR fallback in pdf-inspector's hybrid pipeline?

The OCR fallback triggers when any of three quality detectors flag issues within a specific region: `region_items_have_decoding_issue` identifies broken CID→Unicode mappings, `is_cid_garbage` detects unresolvable glyph-ID fonts, and `detect_encoding_issues` finds replacement characters or encoding artifacts in the rendered markdown. These checks occur in [`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs) and operate independently on each region.

### Can region-based extraction handle multiple regions on the same page?

Yes, the API accepts multiple bounding boxes per page through the `page_regions` parameter, which takes a slice of `(u32, Vec<[f32; 4]>)` tuples. The assignment logic in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) automatically distributes `TextItem`s to their respective regions based on largest overlap area, ensuring each text element belongs to exactly one region even when multiple regions exist on the same page.

### How does table extraction integrate with region-based quality checks?

The `extract_tables_in_regions_mem` function applies the same region-assignment and quality-assessment logic as text extraction before invoking table detection modules. This means tables located in regions with encoding issues will be flagged for OCR processing, while clean regions proceed with native text extraction, ensuring consistent hybrid OCR behavior across both text and table content.