How to Use `extract_positioned_text_from_doc` to Extract Raw `TextItem` Objects with Coordinates in pdf-inspector

Call extract_positioned_text_from_doc from src/extractor/mod.rs and access the Vec<TextItem> in the returned PageExtraction tuple to get raw text with x/y coordinates, dimensions, and font metadata.

The pdf-inspector crate provides a low-level API for extracting precisely positioned text from PDF documents. Unlike high-level extractors that return plain strings, the extract_positioned_text_from_doc function preserves every glyph's geometric placement and styling information. This guide covers how to use this extractor in both Rust and Python, with direct reference to the source implementation in the firecrawl/pdf-inspector repository.

Understanding extract_positioned_text_from_doc

The core extractor lives in [src/extractor/mod.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs#L158-L165) and exposes the following signature:

pub fn extract_positioned_text_from_doc(
    doc: &Document,
    font_cmaps: &FontCMaps,
    page_filter: Option<&HashSet<u32>>,
) -> Result<(PageExtraction, PageThresholds, HashSet<u32>), PdfError>

The function returns a three-element tuple. For coordinate extraction, focus on the first element:

  • PageExtraction = (Vec<TextItem>, Vec<PdfRect>, Vec<PdfLine>) — the Vec<TextItem> contains your raw positioned text objects

What Each TextItem Contains

The TextItem struct is defined in [src/types.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs#L96-L132):

Field Type Purpose
text String Unicode content from the glyph run
x, y f64 Bottom-left coordinates in PDF space (origin at lower-left corner)
width, height f64 Approximate bounding box from font metrics
font String Font identifier or name
font_size f64 Text size in points
page u32 1-based page number
is_bold, is_italic, is_underline, is_strikeout bool Style flags detected geometrically
item_type ItemType Enum: Text, Image, Link, or FormField
mcid Option<i32> Marked-content ID for tagged PDF accessibility

Rust Implementation

Use extract_positioned_text_from_doc directly from the crate's public API in [src/lib.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs):

use pdf_inspector::{
    extractor::{extract_positioned_text_from_doc, FontCMaps},
    load_document_from_mem, PdfError,
};

fn main() -> Result<(), PdfError> {
    // Step 1: Load PDF bytes into the internal Document structure
    let pdf_bytes = std::fs::read("example.pdf")?;
    let (doc, _) = load_document_from_mem(&pdf_bytes)?;

    // Step 2: Build the font-to-Unicode cache
    // This caches CMap and ToUnicode tables for proper glyph decoding
    let font_cmaps = FontCMaps::from_doc(&doc);

    // Step 3: Extract all positioned text items (None = all pages)
    let (extraction, _thresholds, _gid_pages) =
        extract_positioned_text_from_doc(&doc, &font_cmaps, None)?;

    // Step 4: Destructure to access the raw TextItem vector
    let (items, _rects, _lines) = extraction;

    // Step 5: Iterate over positioned text with full coordinate data
    for item in items {
        println!(
            "Page {} | '{}' | x:{:.2} y:{:.2} w:{:.2} h:{:.2} | font:{} size:{:.1}",
            item.page,
            item.text,
            item.x,
            item.y,
            item.width,
            item.height,
            item.font,
            item.font_size,
        );
    }

    Ok(())
}

Required Setup

The FontCMaps cache is mandatory. Without it, glyph IDs cannot be mapped to Unicode, and you'll receive garbled text or hex-encoded character codes.

Python Implementation via PyO3 Bindings

The Python interface in [src/python.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs) mirrors the Rust API. Install with:

pip install pdf-inspector
import pdf_inspector

# Load PDF document

doc = pdf_inspector.load_document("example.pdf")

# Build font cache (automatic in Python wrapper)

font_cmaps = pdf_inspector.FontCMaps.from_doc(doc)

# Extract positioned text items

items, _thresholds, _gid_pages = pdf_inspector.extract_positioned_text_from_doc(
    doc, font_cmaps, None  # None extracts all pages

)

# Access coordinates and metadata for each TextItem

for item in items:
    print(
        f"Page {item.page} | '{item.text}' | "
        f"x={item.x:.2f} y={item.y:.2f} w={item.width:.2f} h={item.height:.2f} | "
        f"font={item.font} size={item.font_size:.1f}"
    )
    
    # Check style flags

    if item.is_bold:
        print(f"  → Bold text detected")
    if item.is_underline:
        print(f"  → Underlined text detected")

The Python TextItem exposes identical fields to the Rust struct through the PyO3 binding layer.

Filtering by Specific Pages

Pass a page filter set to limit extraction instead of processing the entire document:

Rust

use std::collections::HashSet;

let mut target_pages = HashSet::new();
target_pages.insert(5);
target_pages.insert(7);

let (extraction, _, _) = extract_positioned_text_from_doc(
    &doc,
    &font_cmaps,
    Some(&target_pages)  // Only pages 5 and 7
)?;

let (items, _, _) = extraction;
// items now contains only TextItems from pages 5 and 7

Python

target_pages = {5, 7}

items, _, _ = pdf_inspector.extract_positioned_text_from_doc(
    doc, font_cmaps, target_pages
)

Coordinate System Reference

PDF coordinates follow a specific convention that affects how you interpret x and y:

  • Origin: Bottom-left corner of the page
  • Units: Points (1/72 inch)
  • Y-axis: Increases upward (opposite of many screen coordinate systems)
  • Page rotation: Applied in user space; raw coordinates reflect the stored transformation

For layout analysis, combine x, y, width, and height to reconstruct bounding boxes. The y coordinate represents the baseline position adjusted by font descent.

Working with Extraction Metadata

The full return tuple from extract_positioned_text_from_doc provides additional pipeline data:

Position Type Use Case
0 PageExtraction Text items, rectangles, and lines for document reconstruction
1 PageThresholds Font size thresholds for header/footer detection algorithms
2 HashSet<u32> Pages containing GID-font glyphs (diagnostic for font handling)

Most applications only need the first tuple element. The second and third are reserved for advanced layout analysis in downstream processing.

Summary

  • extract_positioned_text_from_doc in src/extractor/mod.rs returns raw TextItem objects with complete coordinate data
  • Always build FontCMaps from the document first for proper Unicode decoding
  • Access positioned text through the PageExtraction tuple: (items, rects, lines)
  • Each TextItem defined in src/types.rs provides x, y, width, height, font, font_size, and geometric style flags
  • Use page_filter with a HashSet<u32> to extract from specific pages only
  • Both Rust and Python APIs expose identical functionality through PyO3 bindings in src/python.rs

Frequently Asked Questions

What is the coordinate system for TextItem x and y values?

PDF user space: origin at the bottom-left corner, Y increasing upward, units in points. The x and y fields in TextItem represent the bottom-left position of the glyph run's bounding box. This differs from screen coordinates where Y typically increases downward. For layout reconstruction, you may need to transform these values based on page MediaBox or CropBox dimensions.

Why must I create FontCMaps before extracting text?

Glyph IDs require font-specific CMap or ToUnicode tables to become readable Unicode. PDFs often store text as font-specific glyph identifiers rather than Unicode codepoints. FontCMaps::from_doc caches these mapping tables from the PDF's font dictionaries. Without this cache, extract_positioned_text_from_doc cannot decode embedded fonts—particularly problematic with GID fonts and subsetted Type1C streams.

How do I distinguish headers from body text using TextItem data?

Compare font_size against PageThresholds or apply heuristics to y position. The second tuple element (PageThresholds) contains computed font size thresholds that separate headers, body text, and footnotes. Alternatively, analyze the y coordinate distribution: headers cluster near the top of the page (high Y values), footers near the bottom (low Y values). The is_bold flag also correlates with header styling in many document templates.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →