How to Extract Text from Specific Bounding Box Regions Using `extract_text_in_regions_mem` in pdf-inspector

Use the extract_text_in_regions_mem function in firecrawl/pdf-inspector to extract text from user-defined rectangular regions by passing PDF bytes and a list of page-specific bounding boxes, returning a vector of strings without filesystem I/O.

The extract_text_in_regions_mem function provides a zero-copy, in-memory API for precise text extraction from PDF documents. According to the firecrawl/pdf-inspector source code, this low-level function powers the crate's region-based extraction capabilities while operating entirely on byte slices—making it ideal for serverless deployments, WebAssembly environments, and streaming pipelines where disk access is prohibited or undesirable.

What extract_text_in_regions_mem Does

extract_text_in_regions_mem serves as the core engine behind pdf-inspector's bounding-box extraction. Unlike the high-level pdf2md binary, which converts entire documents to Markdown, this function targets specific rectangular regions on specific pages and returns only the text contained within those boundaries.

The function signature in src/lib.rs accepts:

  • pdf_bytes: &[u8] — The complete PDF document as a byte slice
  • regions: &[(PdfPageNum, PdfRect)] — A slice of tuples pairing page numbers with bounding rectangles

It returns Result<Vec<String>, PdfInspectorError> where each string corresponds to the input region in order.

Internal Extraction Pipeline

The function orchestrates four specialized stages:

  1. PDF parsing — Opens the byte stream with lopdf and decodes content streams (src/extractor/content_stream.rs)
  2. Font resolution — Maps glyphs to Unicode via tounicode tables with fallback logic (src/extractor/fonts.rs, src/tounicode.rs)
  3. Region clipping — Intersects text rendering matrices with user-supplied rectangles during content stream walking
  4. Text normalization — Applies ligature expansion and NFKC normalization (src/text_utils.rs)

Calling extract_text_in_regions_mem from Rust

The native Rust API provides full type safety through the PdfRect and PdfPageNum newtypes defined in src/types.rs.

use pdf_inspector::{
    extract_text_in_regions_mem,
    types::{PdfRect, PdfPageNum},
};

fn main() -> Result<(), Box<dyn std::error::Error>> {
    // Load PDF into memory from any source (file, network, embedded bytes)
    let pdf_bytes = std::fs::read("sample.pdf")?;

    // Define bounding boxes in PDF user space coordinates (points)
    // Format: (page_number, left, bottom, right, top)
    let regions = vec![
        (PdfPageNum::new(1), PdfRect::new(100.0, 200.0, 300.0, 250.0)),
        (PdfPageNum::new(2), PdfRect::new(50.0, 400.0, 400.0, 500.0)),
    ];

    // Execute extraction—returns Vec<String> aligned with input order
    let extracted = extract_text_in_regions_mem(&pdf_bytes, &regions)?;

    for (i, text) in extracted.iter().enumerate() {
        println!("Region {}: {}", i + 1, text);
    }

    Ok(())
}

Key implementation details from src/lib.rs:

  • The function wraps extractor::extract_regions with public visibility
  • Errors propagate from PdfInspectorError with context about parsing failures or invalid regions
  • No temporary files are created; all operations occur in heap-allocated buffers

Python API via N-API Bindings

The napi/src/lib.rs module exposes the same functionality to Python and Node.js environments with automatic type conversion for the region tuples.

import pdf_inspector

# PDF bytes from any source—files, HTTP responses, databases

with open("sample.pdf", "rb") as f:
    pdf_data = f.read()

# Regions as simple tuples: (page, left, bottom, right, top)

# Coordinates are float values in PDF user space

regions = [
    (1, 100.0, 200.0, 300.0, 250.0),
    (2, 50.0, 400.0, 400.0, 500.0),
]

# Returns list[str] with preserved ordering

texts = pdf_inspector.extract_text_in_regions_mem(pdf_data, regions)

for i, txt in enumerate(texts):
    print(f"Region {i+1}: {txt}")

Conversion behavior in napi/src/lib.rs:

  • Python int values for page numbers convert to PdfPageNum::new()
  • Python float or int coordinates convert to PdfRect::new(left, bottom, right, top)
  • The function raises PdfInspectorError exceptions on parsing failures

CLI Usage with pdf2md

The pdf2md binary in src/bin/pdf2md.rs provides a JSON interface for testing and scripting.

pdf2md \
  --json \
  --extract-regions '[
    {"page":1,"bbox":[100.0,200.0,300.0,250.0]},
    {"page":2,"bbox":[50.0,400.0,400.0,500.0]}
  ]' \
  sample.pdf

Output format:

{
  "regions": [
    "Text from first region on page 1...",
    "Text from second region on page 2..."
  ]
}

Understanding Bounding Box Coordinates

PDF user space coordinates require careful interpretation. The PdfRect struct in src/types.rs uses the conventional PDF coordinate system:

Parameter Description Typical Range
left X-coordinate of rectangle's left edge 0 to page width
bottom Y-coordinate of rectangle's bottom edge 0 to page height
right X-coordinate of rectangle's right edge > left
top Y-coordinate of rectangle's top edge > bottom

Critical: PDF coordinates originate at the bottom-left corner of the page, unlike many image formats where (0,0) is top-left. When converting from screen coordinates or image processing libraries, you must flip the Y-axis using the page's /CropBox or /MediaBox height.

Performance Characteristics

extract_text_in_regions_mem optimizes for the specific case of partial document extraction:

  • Memory: O(P + R) where P is page content stream size and R is region count—only pages containing target regions are fully processed
  • Time: Linear in content stream operators; clipping occurs during parsing to avoid post-processing
  • Zero-allocation parsing: The content stream walker in src/extractor/content_stream.rs reuses buffers for operator operands

For documents with many regions across many pages, batch regions by page to minimize repeated content stream decoding.

Error Handling Patterns

The function returns distinct error variants from src/error.rs:

match extract_text_in_regions_mem(&pdf_bytes, &regions) {
    Ok(texts) => println!("Extracted {} regions", texts.len()),
    Err(PdfInspectorError::InvalidPageNumber(p)) => {
        eprintln!("Page {} does not exist in document", p);
    }
    Err(PdfInspectorError::ParsingFailed(msg)) => {
        eprintln!("Malformed PDF: {}", msg);
    }
    Err(e) => eprintln!("Extraction failed: {:?}", e),
}

Summary

  • extract_text_in_regions_mem operates on &[u8] without filesystem dependencies, as implemented in src/lib.rs
  • Coordinate system: PDF user space with bottom-left origin, expressed through PdfRect in src/types.rs
  • Pipeline stages: PDF parsing → font resolution → region clipping → text normalization across src/extractor/ modules
  • Multi-language support: Native Rust, Python/Node via napi/src/lib.rs, and CLI through src/bin/pdf2md.rs
  • Performance: Streaming content stream processing with early clipping to minimize unnecessary work

Frequently Asked Questions

Why does my extracted text contain garbled characters?

Garbled output typically indicates missing or malformed ToUnicode CMaps in the source PDF. The extractor falls back to encoding-based glyph mapping from src/extractor/fonts.rs, but some proprietary fonts lack extractable Unicode data. Try the -repair flag in pdf2md or pre-process with qpdf --qdf to normalize font structures.

How do I convert screen coordinates to PDF coordinates?

Screen coordinates (top-left origin) require Y-axis inversion. Retrieve the page's MediaBox height from the PDF catalog, then calculate: pdf_y = media_box_height - screen_y. The lopdf integration in src/extractor/mod.rs provides page dictionary access for this transformation.

Can I extract text from rotated pages?

Yes. The extractor in src/extractor/content_stream.rs applies the complete text rendering matrix including rotation components when testing bounding box intersections. Specify coordinates in the final, rendered user space—no manual transformation required.

Is there a streaming API for very large PDFs?

extract_text_in_regions_mem currently requires the complete document in memory. For streaming scenarios, the internal extractor module supports incremental page processing, but this interface is not yet public. For production streaming, consider memory-mapping the file or using the memmap2 crate with extract_text_in_regions_mem.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →