LiteParse Rotation Adjustment: How It Normalizes Counter-Clockwise Page Rotation for TextItem Objects

LiteParse handles counter-clockwise page rotation by detecting each TextItem's rotation angle, grouping items by their canonical orientation, transforming their coordinates to an upright coordinate space, and flagging them as rotated so downstream layout logic maintains correct reading order.

When extracting text from PDFs, LiteParse represents every piece of content as a ProjectedTextItem. Physical pages may contain text rotated at 90°, 180°, or 270° counter-clockwise, which would break natural left-to-right reading flow if left unadjusted. The normalization logic lives primarily in crates/liteparse/src/projection.rs, where the pipeline canonicalizes rotation values, clusters related items, and performs geometric transformations to present all text in a unified, upright coordinate system.

Detecting and Canonicalizing Rotation Values

The rotation adjustment process begins by standardizing the raw rotation values reported by the PDF extraction layer. PDFium may report slight variations in rotation angles due to floating-point imprecision or subtle page transformations.

Canonical Rotation Calculation

The canonical_rotation function in crates/liteparse/src/projection.rs rounds any raw rotation to the nearest 90° increment (0°, 90°, 180°, or 270°), tolerating a ±2° error margin:

// https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs#L89-L111
fn canonical_rotation(rotation: f32) -> i32 {
    // Rounds to nearest 0, 90, 180, 270 with ±2° tolerance
    // Implementation rounds to nearest 90° quadrant
}

This canonicalization ensures that text items with nearly identical rotations (e.g., 89° vs. 91°) are treated as the same 90° group, preventing fragmentation during the grouping phase.

Early Exit for Unrotated Content

The entry point handle_rotation_reading_order first checks whether any processing is necessary. If no items exhibit non-zero rotation, the function returns immediately to avoid unnecessary computation:

// https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs#L113-L119
fn handle_rotation_reading_order(items: &mut [ProjectedTextItem], page_height: f32) {
    if !items.iter().any(|b| canonical_rotation(b.item.rotation) != 0) {
        return;
    }
    // ... proceed with rotation handling
}

Grouping and Clustering Rotated Text Items

Once the presence of rotated content is confirmed, LiteParse organizes items to preserve spatial relationships during transformation.

Rotation-Based Grouping

All items are categorized into a HashMap<i32, Vec<usize>> keyed by their canonical rotation angle. This separation ensures that 90° text (counter-clockwise) and 270° text (clockwise) undergo different transformation logic:

// https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs#L121-L126
let mut groups_by_rotation = HashMap::new();
for (idx, bbox) in items.iter().enumerate() {
    let r = canonical_rotation(bbox.item.rotation);
    groups_by_rotation.entry(r).or_default().push(idx);
}

Spatial Clustering for Vertical Labels

When processing 90° or 270° groups, LiteParse further subdivides items into spatial clusters based on y-gap thresholds. This step is critical for documents like schematics that contain multiple vertical text columns—ensuring that labels at the top of the page remain distinct from those at the bottom after rotation:

// https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs#L131-L152
// Cluster splitting logic based on y-gap size

Coordinate Transformation Logic

For each identified cluster, LiteParse applies geometric transformations to convert rotated coordinates into the unrotated page space. The algorithm calculates a delta_y offset to align rotated clusters with existing content rows.

90° Counter-Clockwise Adjustment

For text rotated 90° counter-clockwise (appearing as vertical text reading bottom-to-top), the coordinates undergo a swap and shift operation:

// https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs#L263-L274
let new_x = items[*idx].item.y.round();                     // swap X←Y
let new_y = items[*idx].item.x + delta_y;                   // shift Y
let new_w = items[*idx].item.height;                        // swap width/height
let new_h = items[*idx].item.width;

items[*idx].item.rotation = 0.0;
items[*idx].rotated = true;

After transformation, the item's rotation flag is cleared to 0.0 while the boolean rotated flag is set to true, signaling to downstream layout algorithms that this item originated from rotated source material.

270° Clockwise Adjustment

For 270° rotation (vertical text reading top-to-bottom), the transformation accounts for the cluster's bottom boundary:

// https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs#L278-L292
let max_y = // bottom of the cluster
let new_x = (max_y - items[*idx].item.y - items[*idx].item.height).round();
let new_y = items[*idx].item.x + delta_y;
let new_w = items[*idx].item.height;
let new_h = items[*idx].item.width;

180° Upside-Down Handling

For upside-down text (180° rotation), LiteParse preserves the original x-ordering, clears the rotation value, and marks the items as rotated without coordinate swapping:

// https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs#L311-L327
// 180° handling preserves x-ordering, marks as rotated

Final Reading Order Normalization

After all geometric transformations complete, the entire items vector undergoes a final sort by the normalized y coordinate. This sort guarantees that subsequent layout logic—such as line formation and paragraph detection—processes content in correct top-to-bottom reading order:

// https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs#L329-L330
items.sort_by(|a, b| a.item.y.total_cmp(&b.item.y));

Practical Implementation Example

The following example demonstrates parsing a PDF containing rotated labels and inspecting the normalized output:

use liteparse::LiteParse;
use liteparse::config::LiteParseConfig;

let cfg = LiteParseConfig {
    // Configure extraction options...
    ..Default::default()
};

let parser = LiteParse::new(cfg)?;
let result = parser.parse_file("samples/rotated_label.pdf")?;

// Access normalized text items
for item in result.text_items {
    println!("Text: {} (rotated: {})", item.text, item.rotated);
}

For debugging purposes, you can invoke the rotation handler directly on a mutable slice of items:

liteparse::projection::handle_rotation_reading_order(&mut items, page_height);

Summary

  • Canonicalization: LiteParse rounds raw rotation values to the nearest 90° increment using canonical_rotation, tolerating ±2° of error.
  • Selective Processing: The handle_rotation_reading_order function exits early if no rotation is detected, optimizing performance for standard documents.
  • Intelligent Grouping: Items are grouped by rotation angle and spatially clustered to preserve relationships between vertically stacked labels.
  • Geometric Transformation: 90° and 270° rotations trigger coordinate swapping (x↔y) and dimension swapping (width↔height), while 180° rotations preserve x-ordering.
  • Metadata Preservation: After transformation, items retain a rotated flag but reset their rotation angle to 0.0, maintaining compatibility with downstream layout algorithms.
  • Reading Order Guarantee: A final y-coordinate sort ensures content flows logically from top to bottom regardless of original page orientation.

Frequently Asked Questions

How does LiteParse determine if a TextItem needs rotation adjustment?

LiteParse examines the rotation field of each ProjectedTextItem and passes it through the canonical_rotation function in crates/liteparse/src/projection.rs. This function rounds the raw angle to the nearest 90° increment (0°, 90°, 180°, or 270°). If any item returns a non-zero canonical value, the system triggers the full rotation handling pipeline; otherwise, it returns early to save processing cycles.

What happens to the dimensions of a TextItem during 90-degree rotation adjustment?

During 90° counter-clockwise rotation handling, LiteParse swaps the width and height values of the TextItem while transforming its coordinates. Specifically, new_w becomes the original height and new_h becomes the original width, reflecting the physical reality that a vertical line of text becomes horizontal after rotation normalization. This swap occurs in the coordinate transformation block at lines 263-274 of projection.rs.

Does LiteParse support PDFs with mixed rotation angles on the same page?

Yes, the rotation adjustment system explicitly supports mixed-orientation pages. The groups_by_rotation HashMap separates items into distinct buckets (0°, 90°, 180°, 270°) before processing. Furthermore, the spatial clustering logic splits 90° and 270° groups into sub-clusters based on vertical gaps, ensuring that multiple columns of rotated text—such as pin labels on opposite sides of a schematic—maintain their distinct identities through the transformation process.

How can I detect if a TextItem was originally rotated after parsing?

After the rotation adjustment completes, each ProjectedTextItem contains a boolean rotated field set to true if the item originated from non-zero rotation. Simultaneously, the rotation float value resets to 0.0 to reflect the normalized coordinate space. You can inspect this flag in the parsed results to apply special formatting or handling logic to originally-vertical text.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →