How to Use `extract_positioned_text_from_doc` to Extract Raw `TextItem` Objects with Coordinates in pdf-inspector
Call extract_positioned_text_from_doc from src/extractor/mod.rs and access the Vec<TextItem> in the returned PageExtraction tuple to get raw text with x/y coordinates, dimensions, and font metadata.
The pdf-inspector crate provides a low-level API for extracting precisely positioned text from PDF documents. Unlike high-level extractors that return plain strings, the extract_positioned_text_from_doc function preserves every glyph's geometric placement and styling information. This guide covers how to use this extractor in both Rust and Python, with direct reference to the source implementation in the firecrawl/pdf-inspector repository.
Understanding extract_positioned_text_from_doc
The core extractor lives in [src/extractor/mod.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs#L158-L165) and exposes the following signature:
pub fn extract_positioned_text_from_doc(
doc: &Document,
font_cmaps: &FontCMaps,
page_filter: Option<&HashSet<u32>>,
) -> Result<(PageExtraction, PageThresholds, HashSet<u32>), PdfError>
The function returns a three-element tuple. For coordinate extraction, focus on the first element:
PageExtraction=(Vec<TextItem>, Vec<PdfRect>, Vec<PdfLine>)— theVec<TextItem>contains your raw positioned text objects
What Each TextItem Contains
The TextItem struct is defined in [src/types.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs#L96-L132):
| Field | Type | Purpose |
|---|---|---|
text |
String |
Unicode content from the glyph run |
x, y |
f64 |
Bottom-left coordinates in PDF space (origin at lower-left corner) |
width, height |
f64 |
Approximate bounding box from font metrics |
font |
String |
Font identifier or name |
font_size |
f64 |
Text size in points |
page |
u32 |
1-based page number |
is_bold, is_italic, is_underline, is_strikeout |
bool |
Style flags detected geometrically |
item_type |
ItemType |
Enum: Text, Image, Link, or FormField |
mcid |
Option<i32> |
Marked-content ID for tagged PDF accessibility |
Rust Implementation
Use extract_positioned_text_from_doc directly from the crate's public API in [src/lib.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs):
use pdf_inspector::{
extractor::{extract_positioned_text_from_doc, FontCMaps},
load_document_from_mem, PdfError,
};
fn main() -> Result<(), PdfError> {
// Step 1: Load PDF bytes into the internal Document structure
let pdf_bytes = std::fs::read("example.pdf")?;
let (doc, _) = load_document_from_mem(&pdf_bytes)?;
// Step 2: Build the font-to-Unicode cache
// This caches CMap and ToUnicode tables for proper glyph decoding
let font_cmaps = FontCMaps::from_doc(&doc);
// Step 3: Extract all positioned text items (None = all pages)
let (extraction, _thresholds, _gid_pages) =
extract_positioned_text_from_doc(&doc, &font_cmaps, None)?;
// Step 4: Destructure to access the raw TextItem vector
let (items, _rects, _lines) = extraction;
// Step 5: Iterate over positioned text with full coordinate data
for item in items {
println!(
"Page {} | '{}' | x:{:.2} y:{:.2} w:{:.2} h:{:.2} | font:{} size:{:.1}",
item.page,
item.text,
item.x,
item.y,
item.width,
item.height,
item.font,
item.font_size,
);
}
Ok(())
}
Required Setup
The FontCMaps cache is mandatory. Without it, glyph IDs cannot be mapped to Unicode, and you'll receive garbled text or hex-encoded character codes.
Python Implementation via PyO3 Bindings
The Python interface in [src/python.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs) mirrors the Rust API. Install with:
pip install pdf-inspector
import pdf_inspector
# Load PDF document
doc = pdf_inspector.load_document("example.pdf")
# Build font cache (automatic in Python wrapper)
font_cmaps = pdf_inspector.FontCMaps.from_doc(doc)
# Extract positioned text items
items, _thresholds, _gid_pages = pdf_inspector.extract_positioned_text_from_doc(
doc, font_cmaps, None # None extracts all pages
)
# Access coordinates and metadata for each TextItem
for item in items:
print(
f"Page {item.page} | '{item.text}' | "
f"x={item.x:.2f} y={item.y:.2f} w={item.width:.2f} h={item.height:.2f} | "
f"font={item.font} size={item.font_size:.1f}"
)
# Check style flags
if item.is_bold:
print(f" → Bold text detected")
if item.is_underline:
print(f" → Underlined text detected")
The Python TextItem exposes identical fields to the Rust struct through the PyO3 binding layer.
Filtering by Specific Pages
Pass a page filter set to limit extraction instead of processing the entire document:
Rust
use std::collections::HashSet;
let mut target_pages = HashSet::new();
target_pages.insert(5);
target_pages.insert(7);
let (extraction, _, _) = extract_positioned_text_from_doc(
&doc,
&font_cmaps,
Some(&target_pages) // Only pages 5 and 7
)?;
let (items, _, _) = extraction;
// items now contains only TextItems from pages 5 and 7
Python
target_pages = {5, 7}
items, _, _ = pdf_inspector.extract_positioned_text_from_doc(
doc, font_cmaps, target_pages
)
Coordinate System Reference
PDF coordinates follow a specific convention that affects how you interpret x and y:
- Origin: Bottom-left corner of the page
- Units: Points (1/72 inch)
- Y-axis: Increases upward (opposite of many screen coordinate systems)
- Page rotation: Applied in user space; raw coordinates reflect the stored transformation
For layout analysis, combine x, y, width, and height to reconstruct bounding boxes. The y coordinate represents the baseline position adjusted by font descent.
Working with Extraction Metadata
The full return tuple from extract_positioned_text_from_doc provides additional pipeline data:
| Position | Type | Use Case |
|---|---|---|
0 |
PageExtraction |
Text items, rectangles, and lines for document reconstruction |
1 |
PageThresholds |
Font size thresholds for header/footer detection algorithms |
2 |
HashSet<u32> |
Pages containing GID-font glyphs (diagnostic for font handling) |
Most applications only need the first tuple element. The second and third are reserved for advanced layout analysis in downstream processing.
Summary
extract_positioned_text_from_docinsrc/extractor/mod.rsreturns rawTextItemobjects with complete coordinate data- Always build
FontCMapsfrom the document first for proper Unicode decoding - Access positioned text through the
PageExtractiontuple:(items, rects, lines) - Each
TextItemdefined insrc/types.rsprovidesx,y,width,height,font,font_size, and geometric style flags - Use
page_filterwith aHashSet<u32>to extract from specific pages only - Both Rust and Python APIs expose identical functionality through PyO3 bindings in
src/python.rs
Frequently Asked Questions
What is the coordinate system for TextItem x and y values?
PDF user space: origin at the bottom-left corner, Y increasing upward, units in points. The x and y fields in TextItem represent the bottom-left position of the glyph run's bounding box. This differs from screen coordinates where Y typically increases downward. For layout reconstruction, you may need to transform these values based on page MediaBox or CropBox dimensions.
Why must I create FontCMaps before extracting text?
Glyph IDs require font-specific CMap or ToUnicode tables to become readable Unicode. PDFs often store text as font-specific glyph identifiers rather than Unicode codepoints. FontCMaps::from_doc caches these mapping tables from the PDF's font dictionaries. Without this cache, extract_positioned_text_from_doc cannot decode embedded fonts—particularly problematic with GID fonts and subsetted Type1C streams.
How do I distinguish headers from body text using TextItem data?
Compare font_size against PageThresholds or apply heuristics to y position. The second tuple element (PageThresholds) contains computed font size thresholds that separate headers, body text, and footnotes. Alternatively, analyze the y coordinate distribution: headers cluster near the top of the page (high Y values), footers near the bottom (low Y values). The is_bold flag also correlates with header styling in many document templates.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →