# font_size vs font_height in LiteParse TextItem Objects: Understanding the Distinction

> Explore the difference between font_size and font_height in LiteParse TextItem objects. Understand how font_height is calculated from font_size and the transformation matrix for accurate text dimension analysis.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: deep-dive
- Published: 2026-06-06

---

**In LiteParse, `font_size` stores the raw font size declared in the PDF text state, while `font_height` represents the effective vertical dimension after applying the transformation matrix, calculated as `font_size * scale_y`.**

When extracting text from PDFs using the `run-llama/liteparse` crate, each text element is represented as a **TextItem** struct defined in [`crates/liteparse/src/types.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/types.rs). Understanding the distinction between these two dimensional fields is essential for accurate layout analysis, multi-column detection, and OCR alignment tasks.

## What font_size Represents

The **`font_size`** field contains the **nominal font size** reported by the PDF’s text state. This value reflects the size that the PDF author originally declared (for example, “12 pt”) before any transformation matrix is applied to the text.

In [`crates/liteparse/src/extract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/extract.rs), the extractor populates this field directly from the `TextChar` object using `ch.font_size()`. When the PDF reports a zero size, LiteParse derives the value from the character’s bounding box as a fallback. This raw value preserves the intended typography regardless of how the text might be stretched or compressed visually on the page.

## What font_height Represents

The **`font_height`** field stores the **effective vertical size** after the current text matrix has been applied. It equals `font_size * scale_y`, where `scale_y` represents the vertical scaling component extracted from the character’s transformation matrix.

This calculation occurs in [`crates/liteparse/src/extract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/extract.rs) around lines 87-91, where the implementation executes `self.font_height = Some(self.font_size * sy);`. The `font_height` value accounts for non-uniform scaling operations—such as vertical stretching or compression—that occur at the PDF level. When the transformation matrix is missing from the PDF data, this field is omitted (set to `None`).

## Implementation in the Source Code

The differentiation between these values happens during the text extraction pipeline in [`crates/liteparse/src/extract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/extract.rs). As the parser iterates through characters provided by PDFium, it captures both the declared metrics and the transformed dimensions:

- **font_size** is taken directly from `ch.font_size()` or computed from the bounding box when the PDF reports zero.
- **font_height** is calculated by multiplying the raw `font_size` by the `sy` component of the character’s transformation matrix.

This distinction allows downstream components to access either the semantic typography (`font_size`) or the actual rendered geometry (`font_height`) depending on their specific layout requirements.

## Practical Usage Examples

When processing parsed PDF content, choose the appropriate field based on whether you need the declared type size or the visual dimensions:

```rust
use liteparse::LiteParse;
use liteparse::types::TextItem;

// Parse a PDF document
let parser = LiteParse::default();
let parsed = parser.parse_path("document.pdf").unwrap();

// Inspect metrics from the first text element
let first_item: &TextItem = &parsed.pages[0].text_items[0];

println!("Raw font size: {:?}", first_item.font_size);      // → Some(12.0)
println!("Visual height: {:?}", first_item.font_height);    // → Some(13.5) when vertically scaled

```

For layout algorithms that require the actual occupied space, implement a robust fallback pattern:

```rust
fn visual_height(item: &TextItem) -> f32 {
    // Prefer the transformed height; fall back to raw size when unavailable
    item.font_height.unwrap_or_else(|| item.font_size.unwrap_or(0.0))
}

```

Use **`font_size`** when grouping text by intended typography or comparing against style guidelines. Use **`font_height`** when calculating precise vertical positioning, detecting multi-column layouts, or aligning OCR results where the visual appearance matters more than the declared point size.

## Summary

- **`font_size`** contains the raw font size value from the PDF text state, representing the author’s declared size before any geometric transformations.
- **`font_height`** provides the effective vertical dimension after applying the transformation matrix, calculated as `font_size * scale_y` in [`crates/liteparse/src/extract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/extract.rs).
- The **transformation matrix** determines the relationship between these values, with `font_height` reflecting non-uniform scaling that stretches or compresses text vertically.
- Always handle `None` values for `font_height` when the matrix is unavailable, falling back to `font_size` for robust text processing pipelines.

## Frequently Asked Questions

### Why would font_height differ from font_size in a TextItem?

When a PDF applies non-uniform scaling through its transformation matrix, the visual height of text can stretch or compress independently of the declared font size. The `font_height` field captures this effective dimension by multiplying `font_size` by the vertical scale factor (`sy`) from the character's transformation matrix, while `font_size` preserves the original declared value set by the PDF author.

### Which field should I use for calculating line spacing?

For accurate line spacing calculations, use `font_height` because it represents the actual vertical space occupied by the glyph on the page. The `font_size` value may not reflect the true dimensions if the PDF contains vertically stretched text, leading to incorrect spacing estimates in layout reconstruction tasks.

### Can font_height be None while font_size has a value?

Yes, `font_height` is optional and will be `None` when the character's transformation matrix is missing or cannot be extracted from the PDF by the PDFium backend. In such cases, you should fall back to using `font_size` as a reasonable approximation of the text height, though it may not account for any scaling transformations applied to the document.

### Where are these fields defined in the LiteParse source code?

Both fields are defined in the `TextItem` struct within [`crates/liteparse/src/types.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/types.rs). The population logic resides in [`crates/liteparse/src/extract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/extract.rs), where the extractor processes `TextChar` objects and calculates `font_height` using the vertical scaling component of the transformation matrix at approximately lines 87-91.