font_size vs font_height in LiteParse TextItem Objects: Understanding the Distinction

In LiteParse, font_size stores the raw font size declared in the PDF text state, while font_height represents the effective vertical dimension after applying the transformation matrix, calculated as font_size * scale_y.

When extracting text from PDFs using the run-llama/liteparse crate, each text element is represented as a TextItem struct defined in crates/liteparse/src/types.rs. Understanding the distinction between these two dimensional fields is essential for accurate layout analysis, multi-column detection, and OCR alignment tasks.

What font_size Represents

The font_size field contains the nominal font size reported by the PDF’s text state. This value reflects the size that the PDF author originally declared (for example, “12 pt”) before any transformation matrix is applied to the text.

In crates/liteparse/src/extract.rs, the extractor populates this field directly from the TextChar object using ch.font_size(). When the PDF reports a zero size, LiteParse derives the value from the character’s bounding box as a fallback. This raw value preserves the intended typography regardless of how the text might be stretched or compressed visually on the page.

What font_height Represents

The font_height field stores the effective vertical size after the current text matrix has been applied. It equals font_size * scale_y, where scale_y represents the vertical scaling component extracted from the character’s transformation matrix.

This calculation occurs in crates/liteparse/src/extract.rs around lines 87-91, where the implementation executes self.font_height = Some(self.font_size * sy);. The font_height value accounts for non-uniform scaling operations—such as vertical stretching or compression—that occur at the PDF level. When the transformation matrix is missing from the PDF data, this field is omitted (set to None).

Implementation in the Source Code

The differentiation between these values happens during the text extraction pipeline in crates/liteparse/src/extract.rs. As the parser iterates through characters provided by PDFium, it captures both the declared metrics and the transformed dimensions:

  • font_size is taken directly from ch.font_size() or computed from the bounding box when the PDF reports zero.
  • font_height is calculated by multiplying the raw font_size by the sy component of the character’s transformation matrix.

This distinction allows downstream components to access either the semantic typography (font_size) or the actual rendered geometry (font_height) depending on their specific layout requirements.

Practical Usage Examples

When processing parsed PDF content, choose the appropriate field based on whether you need the declared type size or the visual dimensions:

use liteparse::LiteParse;
use liteparse::types::TextItem;

// Parse a PDF document
let parser = LiteParse::default();
let parsed = parser.parse_path("document.pdf").unwrap();

// Inspect metrics from the first text element
let first_item: &TextItem = &parsed.pages[0].text_items[0];

println!("Raw font size: {:?}", first_item.font_size);      // → Some(12.0)
println!("Visual height: {:?}", first_item.font_height);    // → Some(13.5) when vertically scaled

For layout algorithms that require the actual occupied space, implement a robust fallback pattern:

fn visual_height(item: &TextItem) -> f32 {
    // Prefer the transformed height; fall back to raw size when unavailable
    item.font_height.unwrap_or_else(|| item.font_size.unwrap_or(0.0))
}

Use font_size when grouping text by intended typography or comparing against style guidelines. Use font_height when calculating precise vertical positioning, detecting multi-column layouts, or aligning OCR results where the visual appearance matters more than the declared point size.

Summary

  • font_size contains the raw font size value from the PDF text state, representing the author’s declared size before any geometric transformations.
  • font_height provides the effective vertical dimension after applying the transformation matrix, calculated as font_size * scale_y in crates/liteparse/src/extract.rs.
  • The transformation matrix determines the relationship between these values, with font_height reflecting non-uniform scaling that stretches or compresses text vertically.
  • Always handle None values for font_height when the matrix is unavailable, falling back to font_size for robust text processing pipelines.

Frequently Asked Questions

Why would font_height differ from font_size in a TextItem?

When a PDF applies non-uniform scaling through its transformation matrix, the visual height of text can stretch or compress independently of the declared font size. The font_height field captures this effective dimension by multiplying font_size by the vertical scale factor (sy) from the character's transformation matrix, while font_size preserves the original declared value set by the PDF author.

Which field should I use for calculating line spacing?

For accurate line spacing calculations, use font_height because it represents the actual vertical space occupied by the glyph on the page. The font_size value may not reflect the true dimensions if the PDF contains vertically stretched text, leading to incorrect spacing estimates in layout reconstruction tasks.

Can font_height be None while font_size has a value?

Yes, font_height is optional and will be None when the character's transformation matrix is missing or cannot be extracted from the PDF by the PDFium backend. In such cases, you should fall back to using font_size as a reasonable approximation of the text height, though it may not account for any scaling transformations applied to the document.

Where are these fields defined in the LiteParse source code?

Both fields are defined in the TextItem struct within crates/liteparse/src/types.rs. The population logic resides in crates/liteparse/src/extract.rs, where the extractor processes TextChar objects and calculates font_height using the vertical scaling component of the transformation matrix at approximately lines 87-91.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →