How to Extract Text from Specific Bounding Box Regions Using `extract_text_in_regions_mem` in pdf-inspector
Use the extract_text_in_regions_mem function in firecrawl/pdf-inspector to extract text from user-defined rectangular regions by passing PDF bytes and a list of page-specific bounding boxes, returning a vector of strings without filesystem I/O.
The extract_text_in_regions_mem function provides a zero-copy, in-memory API for precise text extraction from PDF documents. According to the firecrawl/pdf-inspector source code, this low-level function powers the crate's region-based extraction capabilities while operating entirely on byte slices—making it ideal for serverless deployments, WebAssembly environments, and streaming pipelines where disk access is prohibited or undesirable.
What extract_text_in_regions_mem Does
extract_text_in_regions_mem serves as the core engine behind pdf-inspector's bounding-box extraction. Unlike the high-level pdf2md binary, which converts entire documents to Markdown, this function targets specific rectangular regions on specific pages and returns only the text contained within those boundaries.
The function signature in src/lib.rs accepts:
pdf_bytes: &[u8]— The complete PDF document as a byte sliceregions: &[(PdfPageNum, PdfRect)]— A slice of tuples pairing page numbers with bounding rectangles
It returns Result<Vec<String>, PdfInspectorError> where each string corresponds to the input region in order.
Internal Extraction Pipeline
The function orchestrates four specialized stages:
- PDF parsing — Opens the byte stream with
lopdfand decodes content streams (src/extractor/content_stream.rs) - Font resolution — Maps glyphs to Unicode via
tounicodetables with fallback logic (src/extractor/fonts.rs,src/tounicode.rs) - Region clipping — Intersects text rendering matrices with user-supplied rectangles during content stream walking
- Text normalization — Applies ligature expansion and NFKC normalization (
src/text_utils.rs)
Calling extract_text_in_regions_mem from Rust
The native Rust API provides full type safety through the PdfRect and PdfPageNum newtypes defined in src/types.rs.
use pdf_inspector::{
extract_text_in_regions_mem,
types::{PdfRect, PdfPageNum},
};
fn main() -> Result<(), Box<dyn std::error::Error>> {
// Load PDF into memory from any source (file, network, embedded bytes)
let pdf_bytes = std::fs::read("sample.pdf")?;
// Define bounding boxes in PDF user space coordinates (points)
// Format: (page_number, left, bottom, right, top)
let regions = vec![
(PdfPageNum::new(1), PdfRect::new(100.0, 200.0, 300.0, 250.0)),
(PdfPageNum::new(2), PdfRect::new(50.0, 400.0, 400.0, 500.0)),
];
// Execute extraction—returns Vec<String> aligned with input order
let extracted = extract_text_in_regions_mem(&pdf_bytes, ®ions)?;
for (i, text) in extracted.iter().enumerate() {
println!("Region {}: {}", i + 1, text);
}
Ok(())
}
Key implementation details from src/lib.rs:
- The function wraps
extractor::extract_regionswith public visibility - Errors propagate from
PdfInspectorErrorwith context about parsing failures or invalid regions - No temporary files are created; all operations occur in heap-allocated buffers
Python API via N-API Bindings
The napi/src/lib.rs module exposes the same functionality to Python and Node.js environments with automatic type conversion for the region tuples.
import pdf_inspector
# PDF bytes from any source—files, HTTP responses, databases
with open("sample.pdf", "rb") as f:
pdf_data = f.read()
# Regions as simple tuples: (page, left, bottom, right, top)
# Coordinates are float values in PDF user space
regions = [
(1, 100.0, 200.0, 300.0, 250.0),
(2, 50.0, 400.0, 400.0, 500.0),
]
# Returns list[str] with preserved ordering
texts = pdf_inspector.extract_text_in_regions_mem(pdf_data, regions)
for i, txt in enumerate(texts):
print(f"Region {i+1}: {txt}")
Conversion behavior in napi/src/lib.rs:
- Python
intvalues for page numbers convert toPdfPageNum::new() - Python
floatorintcoordinates convert toPdfRect::new(left, bottom, right, top) - The function raises
PdfInspectorErrorexceptions on parsing failures
CLI Usage with pdf2md
The pdf2md binary in src/bin/pdf2md.rs provides a JSON interface for testing and scripting.
pdf2md \
--json \
--extract-regions '[
{"page":1,"bbox":[100.0,200.0,300.0,250.0]},
{"page":2,"bbox":[50.0,400.0,400.0,500.0]}
]' \
sample.pdf
Output format:
{
"regions": [
"Text from first region on page 1...",
"Text from second region on page 2..."
]
}
Understanding Bounding Box Coordinates
PDF user space coordinates require careful interpretation. The PdfRect struct in src/types.rs uses the conventional PDF coordinate system:
| Parameter | Description | Typical Range |
|---|---|---|
left |
X-coordinate of rectangle's left edge | 0 to page width |
bottom |
Y-coordinate of rectangle's bottom edge | 0 to page height |
right |
X-coordinate of rectangle's right edge | > left |
top |
Y-coordinate of rectangle's top edge | > bottom |
Critical: PDF coordinates originate at the bottom-left corner of the page, unlike many image formats where (0,0) is top-left. When converting from screen coordinates or image processing libraries, you must flip the Y-axis using the page's /CropBox or /MediaBox height.
Performance Characteristics
extract_text_in_regions_mem optimizes for the specific case of partial document extraction:
- Memory: O(P + R) where P is page content stream size and R is region count—only pages containing target regions are fully processed
- Time: Linear in content stream operators; clipping occurs during parsing to avoid post-processing
- Zero-allocation parsing: The content stream walker in
src/extractor/content_stream.rsreuses buffers for operator operands
For documents with many regions across many pages, batch regions by page to minimize repeated content stream decoding.
Error Handling Patterns
The function returns distinct error variants from src/error.rs:
match extract_text_in_regions_mem(&pdf_bytes, ®ions) {
Ok(texts) => println!("Extracted {} regions", texts.len()),
Err(PdfInspectorError::InvalidPageNumber(p)) => {
eprintln!("Page {} does not exist in document", p);
}
Err(PdfInspectorError::ParsingFailed(msg)) => {
eprintln!("Malformed PDF: {}", msg);
}
Err(e) => eprintln!("Extraction failed: {:?}", e),
}
Summary
extract_text_in_regions_memoperates on&[u8]without filesystem dependencies, as implemented insrc/lib.rs- Coordinate system: PDF user space with bottom-left origin, expressed through
PdfRectinsrc/types.rs - Pipeline stages: PDF parsing → font resolution → region clipping → text normalization across
src/extractor/modules - Multi-language support: Native Rust, Python/Node via
napi/src/lib.rs, and CLI throughsrc/bin/pdf2md.rs - Performance: Streaming content stream processing with early clipping to minimize unnecessary work
Frequently Asked Questions
Why does my extracted text contain garbled characters?
Garbled output typically indicates missing or malformed ToUnicode CMaps in the source PDF. The extractor falls back to encoding-based glyph mapping from src/extractor/fonts.rs, but some proprietary fonts lack extractable Unicode data. Try the -repair flag in pdf2md or pre-process with qpdf --qdf to normalize font structures.
How do I convert screen coordinates to PDF coordinates?
Screen coordinates (top-left origin) require Y-axis inversion. Retrieve the page's MediaBox height from the PDF catalog, then calculate: pdf_y = media_box_height - screen_y. The lopdf integration in src/extractor/mod.rs provides page dictionary access for this transformation.
Can I extract text from rotated pages?
Yes. The extractor in src/extractor/content_stream.rs applies the complete text rendering matrix including rotation components when testing bounding box intersections. Specify coordinates in the final, rendered user space—no manual transformation required.
Is there a streaming API for very large PDFs?
extract_text_in_regions_mem currently requires the complete document in memory. For streaming scenarios, the internal extractor module supports incremental page processing, but this interface is not yet public. For production streaming, consider memory-mapping the file or using the memmap2 crate with extract_text_in_regions_mem.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →