# How to Use `extract_positioned_text_from_doc` to Extract Raw `TextItem` Objects with Coordinates in pdf-inspector

> Learn to extract raw TextItem objects with coordinates using extract_positioned_text_from_doc in pdf-inspector. Get text, coordinates, dimensions, and font details efficiently.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-06

---

**Call `extract_positioned_text_from_doc` from [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) and access the `Vec<TextItem>` in the returned `PageExtraction` tuple to get raw text with x/y coordinates, dimensions, and font metadata.**

The `pdf-inspector` crate provides a low-level API for extracting precisely positioned text from PDF documents. Unlike high-level extractors that return plain strings, the **`extract_positioned_text_from_doc`** function preserves every glyph's geometric placement and styling information. This guide covers how to use this extractor in both Rust and Python, with direct reference to the source implementation in the [firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector) repository.

## Understanding `extract_positioned_text_from_doc`

The core extractor lives in [[`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs#L158-L165) and exposes the following signature:

```rust
pub fn extract_positioned_text_from_doc(
    doc: &Document,
    font_cmaps: &FontCMaps,
    page_filter: Option<&HashSet<u32>>,
) -> Result<(PageExtraction, PageThresholds, HashSet<u32>), PdfError>

```

The function returns a three-element tuple. For coordinate extraction, focus on the first element:

- **`PageExtraction`** = `(Vec<TextItem>, Vec<PdfRect>, Vec<PdfLine>)` — the `Vec<TextItem>` contains your raw positioned text objects

### What Each `TextItem` Contains

The **`TextItem`** struct is defined in [[`src/types.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs#L96-L132):

| Field | Type | Purpose |
|-------|------|---------|
| `text` | `String` | Unicode content from the glyph run |
| `x`, `y` | `f64` | Bottom-left coordinates in PDF space (origin at lower-left corner) |
| `width`, `height` | `f64` | Approximate bounding box from font metrics |
| `font` | `String` | Font identifier or name |
| `font_size` | `f64` | Text size in points |
| `page` | `u32` | 1-based page number |
| `is_bold`, `is_italic`, `is_underline`, `is_strikeout` | `bool` | Style flags detected geometrically |
| `item_type` | `ItemType` | Enum: `Text`, `Image`, `Link`, or `FormField` |
| `mcid` | `Option<i32>` | Marked-content ID for tagged PDF accessibility |

## Rust Implementation

Use `extract_positioned_text_from_doc` directly from the crate's public API in [[`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs):

```rust
use pdf_inspector::{
    extractor::{extract_positioned_text_from_doc, FontCMaps},
    load_document_from_mem, PdfError,
};

fn main() -> Result<(), PdfError> {
    // Step 1: Load PDF bytes into the internal Document structure
    let pdf_bytes = std::fs::read("example.pdf")?;
    let (doc, _) = load_document_from_mem(&pdf_bytes)?;

    // Step 2: Build the font-to-Unicode cache
    // This caches CMap and ToUnicode tables for proper glyph decoding
    let font_cmaps = FontCMaps::from_doc(&doc);

    // Step 3: Extract all positioned text items (None = all pages)
    let (extraction, _thresholds, _gid_pages) =
        extract_positioned_text_from_doc(&doc, &font_cmaps, None)?;

    // Step 4: Destructure to access the raw TextItem vector
    let (items, _rects, _lines) = extraction;

    // Step 5: Iterate over positioned text with full coordinate data
    for item in items {
        println!(
            "Page {} | '{}' | x:{:.2} y:{:.2} w:{:.2} h:{:.2} | font:{} size:{:.1}",
            item.page,
            item.text,
            item.x,
            item.y,
            item.width,
            item.height,
            item.font,
            item.font_size,
        );
    }

    Ok(())
}

```

### Required Setup

The `FontCMaps` cache is mandatory. Without it, glyph IDs cannot be mapped to Unicode, and you'll receive garbled text or hex-encoded character codes.

## Python Implementation via PyO3 Bindings

The Python interface in [[`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs) mirrors the Rust API. Install with:

```bash
pip install pdf-inspector

```

```python
import pdf_inspector

# Load PDF document

doc = pdf_inspector.load_document("example.pdf")

# Build font cache (automatic in Python wrapper)

font_cmaps = pdf_inspector.FontCMaps.from_doc(doc)

# Extract positioned text items

items, _thresholds, _gid_pages = pdf_inspector.extract_positioned_text_from_doc(
    doc, font_cmaps, None  # None extracts all pages

)

# Access coordinates and metadata for each TextItem

for item in items:
    print(
        f"Page {item.page} | '{item.text}' | "
        f"x={item.x:.2f} y={item.y:.2f} w={item.width:.2f} h={item.height:.2f} | "
        f"font={item.font} size={item.font_size:.1f}"
    )
    
    # Check style flags

    if item.is_bold:
        print(f"  → Bold text detected")
    if item.is_underline:
        print(f"  → Underlined text detected")

```

The Python `TextItem` exposes identical fields to the Rust struct through the PyO3 binding layer.

## Filtering by Specific Pages

Pass a page filter set to limit extraction instead of processing the entire document:

### Rust

```rust
use std::collections::HashSet;

let mut target_pages = HashSet::new();
target_pages.insert(5);
target_pages.insert(7);

let (extraction, _, _) = extract_positioned_text_from_doc(
    &doc,
    &font_cmaps,
    Some(&target_pages)  // Only pages 5 and 7
)?;

let (items, _, _) = extraction;
// items now contains only TextItems from pages 5 and 7

```

### Python

```python
target_pages = {5, 7}

items, _, _ = pdf_inspector.extract_positioned_text_from_doc(
    doc, font_cmaps, target_pages
)

```

## Coordinate System Reference

PDF coordinates follow a specific convention that affects how you interpret `x` and `y`:

- **Origin**: Bottom-left corner of the page
- **Units**: Points (1/72 inch)
- **Y-axis**: Increases upward (opposite of many screen coordinate systems)
- **Page rotation**: Applied in user space; raw coordinates reflect the stored transformation

For layout analysis, combine `x`, `y`, `width`, and `height` to reconstruct bounding boxes. The `y` coordinate represents the **baseline** position adjusted by font descent.

## Working with Extraction Metadata

The full return tuple from `extract_positioned_text_from_doc` provides additional pipeline data:

| Position | Type | Use Case |
|----------|------|----------|
| `0` | `PageExtraction` | Text items, rectangles, and lines for document reconstruction |
| `1` | `PageThresholds` | Font size thresholds for header/footer detection algorithms |
| `2` | `HashSet<u32>` | Pages containing GID-font glyphs (diagnostic for font handling) |

Most applications only need the first tuple element. The second and third are reserved for advanced layout analysis in downstream processing.

## Summary

- **`extract_positioned_text_from_doc`** in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) returns raw `TextItem` objects with complete coordinate data
- Always build **`FontCMaps`** from the document first for proper Unicode decoding
- Access positioned text through the `PageExtraction` tuple: `(items, rects, lines)`
- Each `TextItem` defined in [`src/types.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs) provides `x`, `y`, `width`, `height`, `font`, `font_size`, and geometric style flags
- Use `page_filter` with a `HashSet<u32>` to extract from specific pages only
- Both Rust and Python APIs expose identical functionality through PyO3 bindings in [`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs)

## Frequently Asked Questions

### What is the coordinate system for TextItem x and y values?

**PDF user space: origin at the bottom-left corner, Y increasing upward, units in points.** The `x` and `y` fields in `TextItem` represent the bottom-left position of the glyph run's bounding box. This differs from screen coordinates where Y typically increases downward. For layout reconstruction, you may need to transform these values based on page `MediaBox` or `CropBox` dimensions.

### Why must I create FontCMaps before extracting text?

**Glyph IDs require font-specific CMap or ToUnicode tables to become readable Unicode.** PDFs often store text as font-specific glyph identifiers rather than Unicode codepoints. `FontCMaps::from_doc` caches these mapping tables from the PDF's font dictionaries. Without this cache, `extract_positioned_text_from_doc` cannot decode embedded fonts—particularly problematic with GID fonts and subsetted Type1C streams.

### How do I distinguish headers from body text using TextItem data?

**Compare `font_size` against `PageThresholds` or apply heuristics to `y` position.** The second tuple element (`PageThresholds`) contains computed font size thresholds that separate headers, body text, and footnotes. Alternatively, analyze the `y` coordinate distribution: headers cluster near the top of the page (high Y values), footers near the bottom (low Y values). The `is_bold` flag also correlates with header styling in many document templates.