# XY-Cut++ Algorithm for Determining Reading Order in Complex PDF Layouts

> Discover the XY-Cut++ algorithm for complex PDF reading order. This technique efficiently reconstructs natural content flow in challenging layouts, ensuring accurate analysis.

- Repository: [opendataloader-project/opendataloader-pdf](https://github.com/opendataloader-project/opendataloader-pdf)
- Tags: how-to-guide
- Published: 2026-03-20

---

**The XY-Cut++ algorithm is a four-phase recursive segmentation technique that reconstructs natural reading order in complex PDFs by detecting cross-layout elements, analyzing page density to select optimal cutting axes, and reintegrating spanning content into the sorted sequence.**

OpenDataLoader’s implementation in `opendataloader-project/opendataloader-pdf` extends the classic XY-cut method to handle multi-column newspapers, academic papers with full-width headers, and mixed layouts without GPU acceleration. The algorithm resides in [`XYCutPlusPlusSorter.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/XYCutPlusPlusSorter.java) and executes automatically when `Config.READING_ORDER_XYCUT` is enabled, transforming raw PDF objects into a logically ordered stream essential for RAG pipelines and document understanding tasks.

## Four-Phase Architecture of XY-Cut++

According to the class header and implementation comments in [`XYCutPlusPlusSorter.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/XYCutPlusPlusSorter.java)【L26-L44】 and the technical documentation in `reading-order.mdx`【L25-L71】, the algorithm processes each page through four distinct phases orchestrated by the public `sort` method.

### Phase 1: Cross-Layout Element Detection

First, the algorithm identifies elements spanning the full page width—titles, headers, footers, or wide tables—that would otherwise disrupt column detection. The `identifyCrossLayoutElements` method (lines 46-84) employs `hasMinimumOverlaps` to flag these regions based on the **Beta threshold** (`DEFAULT_BETA` = 2.0). Elements exceeding this width ratio are temporarily removed from the segmentation pool and stored for later reintegration, preventing column misdetection.

### Phase 2: Density-Based Axis Selection

Before cutting, `computeDensityRatio` (lines 60-80) calculates the ratio of occupied text area to total page bounding-box area. A high density ratio (> 0.9) signals dense newspaper-style layouts, triggering **horizontal-first** cuts that prioritize row-wise reading. Lower densities trigger **vertical-first** cuts suitable for standard multi-column documents. This adaptive logic prevents fragmentation of logical text blocks in compact layouts.

### Phase 3: Recursive Segmentation

The core `recursiveSegment` method (lines 53-98) projects remaining objects onto X and Y axes to locate the largest whitespace gaps. It evaluates `findBestVerticalCutWithProjection` against `findBestHorizontalCutWithProjection`, selecting the axis with the maximum gap exceeding the **Minimum gap threshold** (`MIN_GAP_THRESHOLD` = 5.0 pts). The algorithm also filters **narrow elements** (`NARROW_ELEMENT_WIDTH_RATIO` = 0.1)—such as page numbers or footnote markers—that could falsely bridge column gaps. Recursion continues until no valid gaps remain, producing a depth-first reading order.

### Phase 4: Cross-Layout Reintegration

Finally, `mergeCrossLayoutElements` (lines 106-124) reinserts the previously removed spanning elements into the sorted sequence based on their Y-coordinates. This ensures headers appear at section tops and footers at the bottom, maintaining logical document flow regardless of recursive partitioning.

## Key Configuration Parameters

The implementation exposes four tunable constants in [`XYCutPlusPlusSorter.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/XYCutPlusPlusSorter.java) that control detection sensitivity:

| Parameter | Default | Function |
|-----------|---------|----------|
| **Beta threshold** (`DEFAULT_BETA`) | 2.0 | Minimum width multiplier for cross-layout classification; lower values detect more spanning elements |
| **Density threshold** (`DEFAULT_DENSITY_THRESHOLD`) | 0.9 | Toggles between horizontal-first (>0.9) and vertical-first (<0.9) cutting strategies |
| **Minimum gap** (`MIN_GAP_THRESHOLD`) | 5.0 pts | Smallest valid whitespace gap for cutting; prevents segmentation on trivial pixel-level noise |
| **Narrow element filter** (`NARROW_ELEMENT_WIDTH_RATIO`) | 0.1 | Width ratio below which objects are ignored during gap calculation |

## Implementation Examples

### Python: Automatic Processing

XY-Cut++ runs automatically under the default configuration when processing through the Python API:

```python
import opendataloader_pdf

# Process batch PDFs; XY-Cut++ sorts content automatically

opendataloader_pdf.convert(
    input_path=["newspaper.pdf", "research_paper.pdf"],
    output_dir="extracted/",
    format="markdown,json"
)

```

The `convert` function spawns the Java engine where `XYCutPlusPlusSorter.sort` processes each page's `IObject` list before output generation.

### CLI: Disabling Layout-Aware Sorting

To extract text in raw PDF object order without recursive segmentation:

```bash
opendataloader-pdf --reading-order off document.pdf

```

This flag maps to `Config.READING_ORDER_OFF` in [`Config.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/Config.java)【L31-L34】, useful for debugging or processing single-column documents where layout analysis adds unnecessary overhead.

### Java: Direct Invocation with Custom Parameters

For advanced use cases requiring threshold adjustment:

```java
import org.opendataloader.pdf.processors.readingorder.XYCutPlusPlusSorter;
import org.verapdf.wcag.algorithms.entities.IObject;

List<IObject> rawObjects = extractPageObjects(page);
List<IObject> ordered = XYCutPlusPlusSorter.sort(rawObjects, 1.5, 0.85);

```

This calls the static `sort` method with a reduced Beta threshold (1.5) and lower density cutoff (0.85), making the detector more aggressive for layouts with subtle spanning headers.

## Summary

- **XY-Cut++** extends recursive XY-cut with four distinct phases: cross-layout detection, density analysis, recursive segmentation, and element reintegration.
- The algorithm automatically adapts cutting strategy based on page density, preferring horizontal cuts for dense newspaper layouts and vertical cuts for standard multi-column documents.
- Critical implementation files include [`XYCutPlusPlusSorter.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/XYCutPlusPlusSorter.java) (core logic), [`Config.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/Config.java) (feature flags), and `reading-order.mdx` (technical documentation).
- Four tunable parameters—Beta threshold, Density threshold, Minimum gap, and Narrow element filter—control detection sensitivity without requiring GPU resources.
- The public `sort` method serves as the entry point for both automatic Python pipelines and direct Java integration.

## Frequently Asked Questions

### How does XY-Cut++ differ from the original XY-cut algorithm?

The original XY-cut recursively bisects pages based solely on whitespace gaps, which fails when spanning titles or dense newspaper layouts create ambiguous cut points. XY-Cut++ adds pre-processing to remove cross-layout elements, density-aware axis selection, and post-processing to reintegrate spanning content, resulting in deterministic reading order for complex layouts that confuse classic implementations.

### When should I disable XY-Cut++ processing?

Disable the algorithm via `--reading-order off` or `Config.READING_ORDER_OFF` when processing single-column documents, scanned images with uniform text blocks, or when preserving the raw PDF content stream order is required for legal compliance. The recursive segmentation overhead provides no benefit for layouts without columns or spanning elements.

### Can I tune XY-Cut++ for right-to-left or vertical reading orders?

Yes. While defaults optimize for left-to-right, top-to-bottom reading, adjusting `DEFAULT_BETA` and `MIN_GAP_THRESHOLD` accommodates other scripts. For vertical CJK layouts, increasing the density threshold above 0.9 forces horizontal-first cuts that respect row-wise progression before columnar segmentation.

### What performance characteristics should I expect?

XY-Cut++ performs CPU-bound geometric calculations without GPU requirements. The algorithm processes pages in O(n log n) time relative to text object count, with recursion depth limited by `MIN_GAP_THRESHOLD`. Complex newspapers with hundreds of text blocks process in milliseconds per page on standard hardware, suitable for high-throughput document processing pipelines.