What is the TextItem Structure in pdf-inspector? Complete Field Reference
The TextItem struct in pdf-inspector is a flat, copy-friendly data structure that captures every piece of extracted PDF text with its position, typography, and semantic flags, defined in src/types.rs and serving as the foundation for the extraction-to-markdown pipeline.
The TextItem structure is the universal currency of the pdf-inspector extraction engine. Every glyph run pulled from a PDF content stream becomes one TextItem, making it essential to understand this struct when building custom pipelines, debugging extraction issues, or extending the library's functionality.
TextItem Definition and Location
The TextItem struct lives in [src/types.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs) alongside related types like ItemType and TextLine. Its flat design prioritizes performance and interop—all fields are public, enabling direct construction from tests, N-API bindings, or external Rust crates.
Here's the complete field breakdown:
| Field | Type | Purpose |
|---|---|---|
text |
String |
Raw character content extracted from the PDF |
x |
f32 |
Horizontal position in PDF user space (left edge) |
y |
f32 |
Vertical position with origin at bottom-left |
width |
f32 |
Rendered width from glyph metrics and font size |
height |
f32 |
Approximate height, typically font size |
font |
String |
BaseFont family name (e.g., "ABCDEF+CMMI10") |
font_tag |
String |
Resource tag distinguishing font programs (e.g., "F2") |
font_size |
f32 |
Size in points |
page |
u32 |
1-indexed page number |
is_bold |
bool |
Bold detection via font metrics or style flags |
is_italic |
bool |
Italic detection via font metrics |
is_underline |
bool |
Geometric underline detection |
is_strikeout |
bool |
Geometric strike-through detection |
item_type |
ItemType |
Discriminates text, images, or hyperlinks |
mcid |
Option<i64> |
Marked-Content ID for tagged PDF structure trees |
Positional Fields: x, y, width, height
The spatial coordinates in TextItem follow PDF conventions exactly. The x and y values place the item's left-most glyph in user space, with y increasing upward from the page bottom. These fields enable:
- Line grouping via vertical proximity in
extractor::layout - Word spacing analysis through
text_utils::should_join_items - Precise Markdown layout preservation
The width and height fields support geometric detection algorithms. In src/extractor/underline.rs, these dimensions help determine whether a horizontal line beneath text qualifies as an underline or strikeout.
Font Identification: font vs. font_tag
The dual font fields solve a common PDF extraction problem. Multiple font programs can share the same BaseFont name within a single page—font_tag provides the unique resource identifier (like "T22") that distinguishes them.
The extractor::fonts::item_font_name function uses font_tag to resolve the actual font program, critical for handling CID fonts where character encoding varies between programs with identical family names.
Semantic Flags: Bold, Italic, Underline, Strikeout
The four boolean flags drive pdf-inspector's high-fidelity Markdown output:
is_bold/is_italic: Detected from font metrics or embedded style flags; trigger**…**or*…*wrapping inmarkdown::convert::text_with_formattingis_underline/is_strikeout: Geometrically detected since PDFs lack explicit style flags; can emit<u>or<s>HTML tags
This geometric detection happens in src/extractor/underline.rs after the content-stream walk completes, comparing line positions against text baselines.
The item_type and mcid Fields
The item_type enum (defined in src/types.rs) extends TextItem beyond pure text to handle:
ItemType::Text— standard text runsItemType::Image— placeholder items for image regionsItemType::Link— hyperlink items with associated URIs
This unified representation lets downstream modules process diverse content uniformly.
The mcid field stores the Marked-Content ID from PDF BDC/BMC operators. For tagged PDFs, this links items to the structure tree, enabling accessibility-aware extraction and semantic role preservation.
Creating and Using TextItem
Manual Construction
For tests or custom pipelines, construct TextItem directly:
use pdf_inspector::types::{TextItem, ItemType};
let item = TextItem {
text: "Hello, world!".to_string(),
x: 72.0,
y: 720.0,
width: 120.0,
height: 12.0,
font: "ABCDEF+Helvetica".to_string(),
font_tag: "F1".to_string(),
font_size: 12.0,
page: 1,
is_bold: false,
is_italic: false,
is_underline: false,
is_strikeout: false,
item_type: ItemType::Text,
mcid: None,
};
Extracting from PDFs
Process a document and filter items by properties:
use pdf_inspector::lib::process_pdf_with_options;
use pdf_inspector::types::ItemType;
let options = Default::default();
let (items, _rects, _lines) = process_pdf_with_options("sample.pdf", options)
.expect("PDF processing failed");
let bold_items: Vec<_> = items
.into_iter()
.filter(|it| it.is_bold && matches!(it.item_type, ItemType::Text))
.collect();
println!("Found {} bold pieces of text", bold_items.len());
Converting to Markdown
TextItem collections become TextLine groups before Markdown conversion:
use pdf_inspector::markdown::convert::LineToMarkdown;
use pdf_inspector::types::TextLine;
let line = TextLine {
items: vec![item],
y: 720.0,
page: 1,
adaptive_threshold: 0.10,
};
let md = line.text_with_formatting(true, true, true);
println!("Markdown: {}", md);
Key Source Files for TextItem
| File | Role |
|---|---|
[src/types.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs) |
TextItem definition and ItemType enum |
[src/extractor/content_stream.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs) |
Populates TextItems during PDF parsing |
[src/markdown/convert.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) |
Reads flags to produce styled Markdown |
[src/extractor/fonts.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/fonts.rs) |
Uses font_tag for CID font resolution |
[src/extractor/underline.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/underline.rs) |
Detects underline/strikeout geometry |
[src/text_utils.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_utils.rs) |
Spacing utilities like should_join_items |
Summary
TextItemis the single shared representation for all extracted PDF content in pdf-inspector- Its 14 fields capture position, typography, styling, and semantic role in a flat, public-struct design
- Geometric detection (underline, strikeout) compensates for PDF's lack of explicit style flags
- Dual font identifiers (
fontandfont_tag) handle complex CID font scenarios - All fields are public and copy-friendly, enabling direct construction from tests, N-API bindings, or external crates
Frequently Asked Questions
What is the difference between font and font_tag in TextItem?
The font field stores the BaseFont family name (e.g., "Helvetica"), while font_tag holds the PDF resource tag (e.g., "F3"). Multiple font programs can share the same family name within a page, so font_tag uniquely identifies which program rendered specific glyphs. The extractor::fonts module uses font_tag to resolve correct character mappings, especially for CID fonts.
How does pdf-inspector detect bold and italic text?
Detection happens through font metrics analysis and embedded style flags in the PDF font descriptor. When the font's weight class or flags indicate bold, is_bold is set to true. Italic detection uses the italic angle metric. These flags then trigger Markdown formatting in text_with_formatting(), wrapping text in ** or * delimiters.
Why does TextItem use geometric detection for underline and strikeout?
PDFs do not encode underline or strikeout as text attributes—they draw separate lines as vector graphics. The extractor::underline module analyzes horizontal line positions relative to text baselines after the content stream finishes. Lines near the baseline become is_underline, lines through the glyph area become is_strikeout. This approach achieves semantic extraction without requiring visual PDF rendering.
When would I use mcid in TextItem processing?
Use mcid when working with tagged PDFs that contain accessibility structure trees. The Marked-Content ID links extracted TextItems to their semantic roles (headings, lists, paragraphs) in the PDF's structure tree. This enables extraction that preserves document semantics rather than just visual layout, critical for screen-reader compatibility and structured data extraction.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →