# How LiteParse Handles Rotated Text in PDFs: A Deep Technical Guide

> Learn how LiteParse handles rotated text in PDFs. Discover its technical approach to detecting, reorienting, and transforming rotated text for logical reading order.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: deep-dive
- Published: 2026-06-24

---

**LiteParse automatically detects, canonicalizes, and reorients rotated text elements during its grid-projection phase, snapping angles to cardinal directions and transforming coordinates so that vertical labels appear in logical reading order.**

LiteParse is a Rust-based PDF parsing library developed by Run Llama that extracts structured text while preserving reading order. When processing documents containing rotated elements—such as vertical axis labels, page numbers, or diagram annotations—the library handles rotated text in PDFs through a sophisticated normalization pipeline that executes before layout analysis.

## The Five-Step Rotation Normalization Pipeline

The `handle_rotation_reading_order` function in [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs) orchestrates the entire process. This pipeline ensures that text stored at 90°, 180°, or 270° in the PDF renders in the correct logical sequence in the final output.

### Step 1: Canonical Rotation Detection

First, the system canonicalizes raw rotation angles using a **±2° tolerance**. The `canonical_rotation` function (lines 89–111 in [`projection.rs`](https://github.com/run-llama/liteparse/blob/main/projection.rs)) computes the nearest cardinal direction using circular distance, mapping near-cardinal angles to exactly 0, 90, 180, or 270. Angles outside this tolerance remain untouched to preserve intentionally skewed elements.

### Step 2: Grouping by Rotation Angle

The algorithm creates a `HashMap<i32, Vec<usize>>` where keys represent canonical rotation angles and values are vectors of item indices. This grouping allows the system to process each rotation class separately. An early exit guard clause (lines 13–19) immediately returns if no items require rotation handling, optimizing performance for straightforward documents.

### Step 3: Spatial Clustering for Vertical Text

For items rotated **90° or 270°**, the system splits groups into spatial clusters to prevent unrelated labels from merging. The logic (lines 30–48) sorts items by their Y-coordinate and breaks clusters whenever the vertical gap exceeds **`3 × max_item_height`**. This ensures that a vertical label at the top of a diagram remains separate from one at the bottom, even when both share the same rotation angle.

### Step 4: Coordinate Transformation and Inline Handling

The library applies geometric transformations based on rotation type:

- **90° rotation**: The system swaps `(x, y)` coordinates and repositions items using a running `delta_y` offset so text reads left-to-right (lines 63–75).
- **270° rotation**: Similar to 90°, but calculates the new X-coordinate from the group's bottom edge to maintain natural orientation (lines 77–95).
- **180° rotation**: Items are simply reordered by X-coordinate and the rotation flag is cleared; no geometry changes are required because a 180° flip does not affect line ordering after the X-sort (lines 111–124).

If a rotated group visually overlaps non-rotated items, the system detects `global_overlap` and **flattens** the group onto a common baseline (average vertical center), marking each item as `rotated = true` (lines 84–128). This keeps inline labels (like vertical tick marks) aligned with surrounding text flow.

### Step 5: Reading Order Sorting and Flagging

After transformation, the system marks processed items with `items[idx].rotated = true` to signal later stages (anchor detection, floating-text alignment) to skip these boxes for column-anchor calculations. Finally, a sort by the new Y-coordinate (lines 129–130) ensures proper top-to-bottom reading order:

```rust
items.sort_by(|a, b| a.item.y.total_cmp(&b.item.y));

```

## Integration with the Parsing Pipeline

The rotation handler integrates seamlessly into the main extraction flow. In [`crates/liteparse/src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs) (lines 65–70), the `project_pages_to_grid` function invokes `project_to_grid`, which internally calls `handle_rotation_reading_order`. This guarantees that **every** page undergoes rotation normalization before layout projection and text-line formation.

The `TextItem` struct in [`types.rs`](https://github.com/run-llama/liteparse/blob/main/types.rs) holds the raw rotation value, while `ProjectedTextItem` adds the boolean `rotated` flag that downstream algorithms consume.

## Code Examples: Parsing Rotated PDFs

### Rust (Core Library)

```rust
use liteparse::LiteParse;
use liteparse::config::LiteParseConfig;
use liteparse::error::LiteParseError;

#[tokio::main]
async fn main() -> Result<(), LiteParseError> {
    // Default configuration – OCR disabled for speed.
    let cfg = LiteParseConfig::default();
    let parser = LiteParse::new(cfg);

    // `rotated.pdf` contains text boxes at 90°, 180° and 270°.
    let result = parser.parse_input(
        liteparse::PdfInput::Path("rotated.pdf".into())
    ).await?;

    // The text is already in logical reading order.
    println!("--- Full document text ---\n{}", result.text);
    Ok(())
}

```

*Result:* The printed text contains vertically-oriented labels in the correct logical order (top-to-bottom, left-to-right).

### Node.js (npm wrapper)

```typescript
import { LiteParse } from "liteparse-node";

(async () => {
  const lp = new LiteParse();          // uses default config
  const res = await lp.parse("rotated.pdf");
  console.log(res.text);               // rotated labels appear in reading order
})();

```

### Python (PyPI wrapper)

```python
from liteparse import LiteParse

lp = LiteParse()                       # defaults, OCR off

res = lp.parse("rotated.pdf")        # parses synchronously

print(res.text)                      # logical order, rotation handled

```

All three bindings delegate to the same Rust core in [`projection.rs`](https://github.com/run-llama/liteparse/blob/main/projection.rs), ensuring consistent rotation handling across languages.

## Key Source Files

| File | Purpose |
|------|---------|
| [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs) | Core projection logic, including `canonical_rotation` and `handle_rotation_reading_order`. |
| [`crates/liteparse/src/types.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/types.rs) | Definitions of `TextItem` (raw rotation) and `ProjectedTextItem` (`rotated` flag). |
| [`crates/liteparse/src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs) | Orchestrates extraction and calls `project_pages_to_grid`. |
| [`crates/liteparse/src/config.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/config.rs) | Configuration; rotation handling requires no user-visible flags. |

## Summary

- **Automatic detection**: LiteParse identifies rotated text within ±2° of cardinal angles during the grid-projection phase.
- **Canonical grouping**: Items are grouped by rotation (0, 90, 180, 270) and spatially clustered to prevent merging unrelated labels.
- **Geometric transformation**: 90° and 270° items undergo coordinate swapping and repositioning; 180° items are simply reordered.
- **Inline preservation**: Overlapping rotated groups are flattened to common baselines while maintaining the `rotated` flag.
- **Universal application**: All language bindings (Rust, Node.js, Python, WASM) share the same normalization logic.

## Frequently Asked Questions

### Does LiteParse require manual configuration to handle rotated text?

No. According to the source code in [`projection.rs`](https://github.com/run-llama/liteparse/blob/main/projection.rs), rotation handling is automatic and executes within `project_to_grid`. The `LiteParseConfig` in [`config.rs`](https://github.com/run-llama/liteparse/blob/main/config.rs) exposes no user-visible flags for rotation—the pipeline detects and normalizes angles without manual intervention.

### What rotation angles does LiteParse support?

The library canonicalizes angles to **0°, 90°, 180°, and 270°** using a ±2° tolerance. The `canonical_rotation` function snaps near-cardinal angles to these exact values. Angles outside this tolerance are preserved as-is, treating them as intentional skews rather than structured rotations.

### How does LiteParse distinguish between rotated labels and body text?

The system uses **spatial clustering** with a threshold of `3 × max_item_height` to separate rotated items vertically. Additionally, the `global_overlap` detection in [`projection.rs`](https://github.com/run-llama/liteparse/blob/main/projection.rs) identifies rotated text that visually intersects with normal text flow, flattening these groups onto common baselines while marking them with the `rotated` flag to exclude them from column-anchor calculations.

### Is the rotation handling available in all language bindings?

Yes. The rotation normalization logic resides in the Rust core ([`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs)), which is wrapped by `liteparse-napi` (Node.js), `liteparse-python` (Python), and `liteparse-wasm` (WebAssembly). All bindings delegate to `handle_rotation_reading_order`, ensuring consistent behavior across Rust, JavaScript, Python, and browser environments.