font_size vs font_height in LiteParse TextItem Objects: Understanding the Distinction
In LiteParse, font_size stores the raw font size declared in the PDF text state, while font_height represents the effective vertical dimension after applying the transformation matrix, calculated as font_size * scale_y.
When extracting text from PDFs using the run-llama/liteparse crate, each text element is represented as a TextItem struct defined in crates/liteparse/src/types.rs. Understanding the distinction between these two dimensional fields is essential for accurate layout analysis, multi-column detection, and OCR alignment tasks.
What font_size Represents
The font_size field contains the nominal font size reported by the PDF’s text state. This value reflects the size that the PDF author originally declared (for example, “12 pt”) before any transformation matrix is applied to the text.
In crates/liteparse/src/extract.rs, the extractor populates this field directly from the TextChar object using ch.font_size(). When the PDF reports a zero size, LiteParse derives the value from the character’s bounding box as a fallback. This raw value preserves the intended typography regardless of how the text might be stretched or compressed visually on the page.
What font_height Represents
The font_height field stores the effective vertical size after the current text matrix has been applied. It equals font_size * scale_y, where scale_y represents the vertical scaling component extracted from the character’s transformation matrix.
This calculation occurs in crates/liteparse/src/extract.rs around lines 87-91, where the implementation executes self.font_height = Some(self.font_size * sy);. The font_height value accounts for non-uniform scaling operations—such as vertical stretching or compression—that occur at the PDF level. When the transformation matrix is missing from the PDF data, this field is omitted (set to None).
Implementation in the Source Code
The differentiation between these values happens during the text extraction pipeline in crates/liteparse/src/extract.rs. As the parser iterates through characters provided by PDFium, it captures both the declared metrics and the transformed dimensions:
- font_size is taken directly from
ch.font_size()or computed from the bounding box when the PDF reports zero. - font_height is calculated by multiplying the raw
font_sizeby thesycomponent of the character’s transformation matrix.
This distinction allows downstream components to access either the semantic typography (font_size) or the actual rendered geometry (font_height) depending on their specific layout requirements.
Practical Usage Examples
When processing parsed PDF content, choose the appropriate field based on whether you need the declared type size or the visual dimensions:
use liteparse::LiteParse;
use liteparse::types::TextItem;
// Parse a PDF document
let parser = LiteParse::default();
let parsed = parser.parse_path("document.pdf").unwrap();
// Inspect metrics from the first text element
let first_item: &TextItem = &parsed.pages[0].text_items[0];
println!("Raw font size: {:?}", first_item.font_size); // → Some(12.0)
println!("Visual height: {:?}", first_item.font_height); // → Some(13.5) when vertically scaled
For layout algorithms that require the actual occupied space, implement a robust fallback pattern:
fn visual_height(item: &TextItem) -> f32 {
// Prefer the transformed height; fall back to raw size when unavailable
item.font_height.unwrap_or_else(|| item.font_size.unwrap_or(0.0))
}
Use font_size when grouping text by intended typography or comparing against style guidelines. Use font_height when calculating precise vertical positioning, detecting multi-column layouts, or aligning OCR results where the visual appearance matters more than the declared point size.
Summary
font_sizecontains the raw font size value from the PDF text state, representing the author’s declared size before any geometric transformations.font_heightprovides the effective vertical dimension after applying the transformation matrix, calculated asfont_size * scale_yincrates/liteparse/src/extract.rs.- The transformation matrix determines the relationship between these values, with
font_heightreflecting non-uniform scaling that stretches or compresses text vertically. - Always handle
Nonevalues forfont_heightwhen the matrix is unavailable, falling back tofont_sizefor robust text processing pipelines.
Frequently Asked Questions
Why would font_height differ from font_size in a TextItem?
When a PDF applies non-uniform scaling through its transformation matrix, the visual height of text can stretch or compress independently of the declared font size. The font_height field captures this effective dimension by multiplying font_size by the vertical scale factor (sy) from the character's transformation matrix, while font_size preserves the original declared value set by the PDF author.
Which field should I use for calculating line spacing?
For accurate line spacing calculations, use font_height because it represents the actual vertical space occupied by the glyph on the page. The font_size value may not reflect the true dimensions if the PDF contains vertically stretched text, leading to incorrect spacing estimates in layout reconstruction tasks.
Can font_height be None while font_size has a value?
Yes, font_height is optional and will be None when the character's transformation matrix is missing or cannot be extracted from the PDF by the PDFium backend. In such cases, you should fall back to using font_size as a reasonable approximation of the text height, though it may not account for any scaling transformations applied to the document.
Where are these fields defined in the LiteParse source code?
Both fields are defined in the TextItem struct within crates/liteparse/src/types.rs. The population logic resides in crates/liteparse/src/extract.rs, where the extractor processes TextChar objects and calculates font_height using the vertical scaling component of the transformation matrix at approximately lines 87-91.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →