How LiteParse Manages Rotated Text in PDF Documents: Grid Projection and Reading Order

LiteParse normalizes rotated text in PDF documents during its grid-projection phase by detecting cardinal angles, grouping items by rotation, transforming their coordinates into axis-aligned boxes, and marking them as rotated so downstream stages preserve logical reading order.

LiteParse is an open-source Rust library for extracting structured text from PDFs. When it encounters rotated text in PDF documents—such as vertical axis labels, sideways page numbers, or inverted annotations—the library automatically reorients those items before layout analysis begins. This rotation normalization happens inside crates/liteparse/src/projection.rs and ensures the final output maintains a natural top-to-bottom, left-to-right flow regardless of how glyphs are stored in the source file.

Detecting and Canonicalizing Rotated Text in PDF Documents

The first step is determining whether a text item is actually rotated. LiteParse maps every raw angle to the nearest cardinal direction—0°, 90°, 180°, or 270°—using a tolerance of ±2°. Angles outside this window are left untouched.

This logic lives in the canonical_rotation function in crates/liteparse/src/projection.rs (lines 89–111). Before any work begins, handle_rotation_reading_order checks whether any items need rotation handling; if not, it exits immediately via a guard clause at lines 13–19.

Grouping and Clustering by Rotation Angle

Once canonical angles are known, LiteParse groups item indices by their rotation value into a HashMap<i32, Vec<usize>> (projection.rs, lines 22–27).

For groups rotated at 90° or 270°, the library performs extra spatial clustering. Items are sorted by their y coordinate and split into separate clusters whenever the vertical gap between neighbors exceeds three times the maximum item height (3 × max_item_height). This prevents unrelated vertical labels—such as pin labels at the top and bottom of a diagram—from being merged into a single logical block (lines 30–48).

Coordinate Transformations for Axis-Aligned Output

LiteParse transforms the bounding boxes of rotated groups so they become axis-aligned before grid projection continues.

90° rotation. Each item’s x, y, width, and height are swapped and repositioned. The new x becomes the old y (rounded), while the new y is computed from the old x plus a running offset called delta_y. This lays the text out left-to-right (projection.rs, lines 63–75).

270° rotation. The transformation is similar to 90°, but the new x is derived from the group’s bottom edge so the natural reading direction remains left-to-right (projection.rs, lines 77–95).

180° rotation. Items are reordered by their x coordinate and the rotation flag is cleared. No geometry transformation is required because a full flip does not alter line ordering once the items are sorted horizontally (projection.rs, lines 111–124).

Handling Inline Overlaps with Non-Rotated Text

If a rotated group visually overlaps any non-rotated items, LiteParse treats it as inline content rather than a floating object. The group is flattened onto a common baseline calculated as the average vertical center of its members, and every item is marked rotated = true.

This collision detection is driven by the global_overlap check in projection.rs (lines 84–128). It ensures that vertical tick marks or small inline labels stay aligned with surrounding text instead of distorting column calculations.

Marking Items and Final Reading-Order Sort

During every transformation branch, LiteParse sets items[idx].rotated = true. This Boolean flag—defined on ProjectedTextItem in crates/liteparse/src/types.rs—signals later pipeline stages, such as anchor detection and flowing-text detection, to exclude these boxes from column-anchor calculations.

After all rotations are normalized, the full item list is sorted by the new y coordinate to enforce a proper top-to-bottom reading order (projection.rs, lines 129–130):

items.sort_by(|a, b| a.item.y.total_cmp(&b.item.y));

Integration in the Parsing Pipeline

Rotation handling is not an optional post-processing step. The handle_rotation_reading_order function is invoked from project_to_grid, which is part of the public project_pages_to_grid pipeline. According to crates/liteparse/src/parser.rs (lines 65–70), this pipeline is called by LiteParse::parse_input on every page before layout projection and text-line formation:

let (projected_items, text) = project_to_grid(&page, projection_boxes);

Because the logic lives in the Rust core, all official language bindings—Node.js, Python, and WebAssembly—inherit the same behavior automatically.

Parsing Rotated PDFs in Rust, Node.js, and Python

You do not need to enable a special flag to handle rotated text in PDF documents. The following examples demonstrate the default behavior across the official wrappers.

Rust

use liteparse::LiteParse;
use liteparse::config::LiteParseConfig;
use liteparse::error::LiteParseError;

#[tokio::main]
async fn main() -> Result<(), LiteParseError> {
    let cfg = LiteParseConfig::default();
    let parser = LiteParse::new(cfg);

    let result = parser.parse_input(
        liteparse::PdfInput::Path("rotated.pdf".into())
    ).await?;

    println!("--- Full document text ---\n{}", result.text);
    Ok(())
}

Node.js

import { LiteParse } from "liteparse-node";

(async () => {
  const lp = new LiteParse();
  const res = await lp.parse("rotated.pdf");
  console.log(res.text);
})();

Python

from liteparse import LiteParse

lp = LiteParse()
res = lp.parse("rotated.pdf")
print(res.text)

In each case, vertically oriented labels are reordered into logical reading order automatically.

Summary

  • LiteParse detects rotated text in PDF documents using canonical_rotation in crates/liteparse/src/projection.rs, snapping angles within ±2° to the nearest cardinal direction.
  • Items are grouped by rotation angle, and 90°/270° groups are split into spatial clusters to keep unrelated labels separate.
  • Coordinates are swapped and repositioned so rotated boxes become axis-aligned, while 180° items are simply re-sorted horizontally.
  • Overlapping rotated groups are flattened onto a common baseline and marked rotated = true to preserve inline flow.
  • The entire process runs automatically inside project_to_grid before the final top-to-bottom sort, ensuring consistent output across Rust, Node.js, and Python bindings.

Frequently Asked Questions

Does LiteParse require manual configuration to handle rotated text in PDF documents?

No. Rotation normalization is automatic and runs inside project_to_grid for every page. There is no user-visible flag in LiteParseConfig to enable or disable it; the behavior is inherited by all language bindings.

What happens to text rotated at an odd angle, such as 45°?

The canonical_rotation function only snaps angles within ±2° of 0°, 90°, 180°, or 270°. Angles outside this tolerance are left untouched and pass through the projection pipeline in their original orientation.

How does LiteParse prevent vertically separated labels from merging into one line?

For 90° and 270° groups, LiteParse sorts items by y and breaks them into distinct clusters whenever the gap exceeds 3 × max_item_height. This threshold ensures spatially distant vertical labels remain independent logical objects.

Why does LiteParse mark rotated items with the rotated flag?

The rotated boolean on ProjectedTextItem tells downstream stages—such as anchor extraction and floating-item alignment—to exclude these boxes from column calculations. This prevents vertical labels from distorting the reading-order inference of normal horizontal paragraphs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →