# How to Extract Text from Specific Bounding Box Regions Using `extract_text_in_regions_mem` in pdf-inspector

> Extract text from specific bounding box regions using pdf-inspectors extract_text_in_regions_mem function. Pass PDF bytes and bounding boxes for precise data extraction without file I/O.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-06

---

**Use the `extract_text_in_regions_mem` function in firecrawl/pdf-inspector to extract text from user-defined rectangular regions by passing PDF bytes and a list of page-specific bounding boxes, returning a vector of strings without filesystem I/O.**

The `extract_text_in_regions_mem` function provides a zero-copy, in-memory API for precise text extraction from PDF documents. According to the firecrawl/pdf-inspector source code, this low-level function powers the crate's region-based extraction capabilities while operating entirely on byte slices—making it ideal for serverless deployments, WebAssembly environments, and streaming pipelines where disk access is prohibited or undesirable.

## What `extract_text_in_regions_mem` Does

`extract_text_in_regions_mem` serves as the core engine behind pdf-inspector's bounding-box extraction. Unlike the high-level `pdf2md` binary, which converts entire documents to Markdown, this function targets specific rectangular regions on specific pages and returns only the text contained within those boundaries.

The function signature in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) accepts:

- **`pdf_bytes: &[u8]`** — The complete PDF document as a byte slice
- **`regions: &[(PdfPageNum, PdfRect)]`** — A slice of tuples pairing page numbers with bounding rectangles

It returns `Result<Vec<String>, PdfInspectorError>` where each string corresponds to the input region in order.

### Internal Extraction Pipeline

The function orchestrates four specialized stages:

1. **PDF parsing** — Opens the byte stream with `lopdf` and decodes content streams ([`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs))
2. **Font resolution** — Maps glyphs to Unicode via `tounicode` tables with fallback logic ([`src/extractor/fonts.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/fonts.rs), [`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs))
3. **Region clipping** — Intersects text rendering matrices with user-supplied rectangles during content stream walking
4. **Text normalization** — Applies ligature expansion and NFKC normalization ([`src/text_utils.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_utils.rs))

## Calling `extract_text_in_regions_mem` from Rust

The native Rust API provides full type safety through the `PdfRect` and `PdfPageNum` newtypes defined in [`src/types.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs).

```rust
use pdf_inspector::{
    extract_text_in_regions_mem,
    types::{PdfRect, PdfPageNum},
};

fn main() -> Result<(), Box<dyn std::error::Error>> {
    // Load PDF into memory from any source (file, network, embedded bytes)
    let pdf_bytes = std::fs::read("sample.pdf")?;

    // Define bounding boxes in PDF user space coordinates (points)
    // Format: (page_number, left, bottom, right, top)
    let regions = vec![
        (PdfPageNum::new(1), PdfRect::new(100.0, 200.0, 300.0, 250.0)),
        (PdfPageNum::new(2), PdfRect::new(50.0, 400.0, 400.0, 500.0)),
    ];

    // Execute extraction—returns Vec<String> aligned with input order
    let extracted = extract_text_in_regions_mem(&pdf_bytes, &regions)?;

    for (i, text) in extracted.iter().enumerate() {
        println!("Region {}: {}", i + 1, text);
    }

    Ok(())
}

```

**Key implementation details from [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs):**
- The function wraps `extractor::extract_regions` with public visibility
- Errors propagate from `PdfInspectorError` with context about parsing failures or invalid regions
- No temporary files are created; all operations occur in heap-allocated buffers

## Python API via N-API Bindings

The [`napi/src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/napi/src/lib.rs) module exposes the same functionality to Python and Node.js environments with automatic type conversion for the region tuples.

```python
import pdf_inspector

# PDF bytes from any source—files, HTTP responses, databases

with open("sample.pdf", "rb") as f:
    pdf_data = f.read()

# Regions as simple tuples: (page, left, bottom, right, top)

# Coordinates are float values in PDF user space

regions = [
    (1, 100.0, 200.0, 300.0, 250.0),
    (2, 50.0, 400.0, 400.0, 500.0),
]

# Returns list[str] with preserved ordering

texts = pdf_inspector.extract_text_in_regions_mem(pdf_data, regions)

for i, txt in enumerate(texts):
    print(f"Region {i+1}: {txt}")

```

**Conversion behavior in [`napi/src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/napi/src/lib.rs):**
- Python `int` values for page numbers convert to `PdfPageNum::new()`
- Python `float` or `int` coordinates convert to `PdfRect::new(left, bottom, right, top)`
- The function raises `PdfInspectorError` exceptions on parsing failures

## CLI Usage with `pdf2md`

The `pdf2md` binary in [`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) provides a JSON interface for testing and scripting.

```bash
pdf2md \
  --json \
  --extract-regions '[
    {"page":1,"bbox":[100.0,200.0,300.0,250.0]},
    {"page":2,"bbox":[50.0,400.0,400.0,500.0]}
  ]' \
  sample.pdf

```

Output format:

```json
{
  "regions": [
    "Text from first region on page 1...",
    "Text from second region on page 2..."
  ]
}

```

## Understanding Bounding Box Coordinates

PDF user space coordinates require careful interpretation. The `PdfRect` struct in [`src/types.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs) uses the conventional PDF coordinate system:

| Parameter | Description | Typical Range |
|-----------|-------------|---------------|
| `left` | X-coordinate of rectangle's left edge | 0 to page width |
| `bottom` | Y-coordinate of rectangle's bottom edge | 0 to page height |
| `right` | X-coordinate of rectangle's right edge | > `left` |
| `top` | Y-coordinate of rectangle's top edge | > `bottom` |

**Critical:** PDF coordinates originate at the bottom-left corner of the page, unlike many image formats where (0,0) is top-left. When converting from screen coordinates or image processing libraries, you must flip the Y-axis using the page's `/CropBox` or `/MediaBox` height.

## Performance Characteristics

`extract_text_in_regions_mem` optimizes for the specific case of partial document extraction:

- **Memory:** O(P + R) where P is page content stream size and R is region count—only pages containing target regions are fully processed
- **Time:** Linear in content stream operators; clipping occurs during parsing to avoid post-processing
- **Zero-allocation parsing:** The content stream walker in [`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs) reuses buffers for operator operands

For documents with many regions across many pages, batch regions by page to minimize repeated content stream decoding.

## Error Handling Patterns

The function returns distinct error variants from [`src/error.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/error.rs):

```rust
match extract_text_in_regions_mem(&pdf_bytes, &regions) {
    Ok(texts) => println!("Extracted {} regions", texts.len()),
    Err(PdfInspectorError::InvalidPageNumber(p)) => {
        eprintln!("Page {} does not exist in document", p);
    }
    Err(PdfInspectorError::ParsingFailed(msg)) => {
        eprintln!("Malformed PDF: {}", msg);
    }
    Err(e) => eprintln!("Extraction failed: {:?}", e),
}

```

## Summary

- **`extract_text_in_regions_mem`** operates on `&[u8]` without filesystem dependencies, as implemented in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)
- **Coordinate system:** PDF user space with bottom-left origin, expressed through `PdfRect` in [`src/types.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs)
- **Pipeline stages:** PDF parsing → font resolution → region clipping → text normalization across `src/extractor/` modules
- **Multi-language support:** Native Rust, Python/Node via [`napi/src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/napi/src/lib.rs), and CLI through [`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs)
- **Performance:** Streaming content stream processing with early clipping to minimize unnecessary work

## Frequently Asked Questions

### Why does my extracted text contain garbled characters?

Garbled output typically indicates missing or malformed **ToUnicode CMaps** in the source PDF. The extractor falls back to **encoding-based glyph mapping** from [`src/extractor/fonts.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/fonts.rs), but some proprietary fonts lack extractable Unicode data. Try the `-repair` flag in `pdf2md` or pre-process with `qpdf --qdf` to normalize font structures.

### How do I convert screen coordinates to PDF coordinates?

Screen coordinates (top-left origin) require Y-axis inversion. Retrieve the page's **MediaBox height** from the PDF catalog, then calculate: `pdf_y = media_box_height - screen_y`. The `lopdf` integration in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) provides page dictionary access for this transformation.

### Can I extract text from rotated pages?

Yes. The extractor in [`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs) applies the complete **text rendering matrix** including rotation components when testing bounding box intersections. Specify coordinates in the final, rendered user space—no manual transformation required.

### Is there a streaming API for very large PDFs?

`extract_text_in_regions_mem` currently requires the complete document in memory. For streaming scenarios, the internal `extractor` module supports incremental page processing, but this interface is not yet public. For production streaming, consider memory-mapping the file or using the `memmap2` crate with `extract_text_in_regions_mem`.