# How pdf-inspector Handles RTL Text and CJK Character Extraction in PDFs

> Learn how pdf-inspector expertly extracts RTL text and CJK characters, preserving logical reading order with advanced algorithms. Get accurate data extraction today.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-06

---

**pdf-inspector treats Right-to-Left (RTL) and CJK scripts as first-class citizens through character-level detection, line direction analysis, and script-aware layout algorithms that preserve logical reading order.**

The [firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector) Rust library specializes in converting PDF documents to clean Markdown. Unlike generic PDF extractors, it implements dedicated handling for **RTL text extraction** (Arabic, Hebrew, Syriac) and **CJK character extraction** (Chinese, Japanese, Korean) through a multi-layered pipeline in [`src/text_utils.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_utils.rs).

## Character-Level Classification for RTL and CJK Scripts

The foundation of pdf-inspector's script handling lies in two fast Unicode range checks:

- **`is_rtl_char`** (lines 97-112): Detects Hebrew (`U+0590–U+05FF`), Arabic (`U+0600–U+06FF`), Syriac, Thaana, NKo, Samaritan, Mandaic, and other RTL scripts
- **`is_cjk_char`** (lines 80-95): Matches Hangul Jamo, Hiragana/Katakana, CJK Unified Ideographs, Compatibility Ideographs, and Half-width/Full-width forms

These functions are invoked throughout the extraction pipeline whenever script-sensitive decisions are required—gap calculations, space insertion, and sorting logic all depend on these classifications.

## Detecting Line Direction with is_rtl_text

Before processing any line, pdf-inspector determines its dominant direction using `is_rtl_text` (lines 120-136):

```rust
let rtl = is_rtl_text(items.iter().map(|i| &i.text));

```

This function iterates every character across all text items in a line, counting RTL code-points against LTR alphabetic characters (excluding CJK). When RTL characters outnumber LTR, the entire line is flagged as right-to-left. This directional flag then drives all subsequent layout decisions.

## Reordering RTL Lines for Logical Reading Order

When a line is identified as RTL, `sort_line_items` (lines 138-145) reverses the X-coordinate ordering:

```rust
pub(crate) fn sort_line_items(items: &mut [TextItem]) {
    let rtl = is_rtl_text(items.iter().map(|i| &i.text));
    if rtl {
        // Reverse X-coordinates so the line reads right-to-left.
        items.sort_by(|a, b| b.x.total_cmp(&a.x));
    } else {
        items.sort_by(|a, b| a.x.total_cmp(&b.x));
    }
}

```

This sorting occurs in three critical locations:
- [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) (line 2788) — after line collection
- [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs) (line 2258) — after column detection
- [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) (line 3313) — final conversion to Markdown

The result: RTL text appears in logical reading order in the output, not the visual left-to-right storage order common in PDF internals.

## CJK-Aware Join Logic: Spaces vs. Geometry

CJK scripts present a unique challenge—they don't use word separators. The `should_join_items` function handles this with script-specific logic:

```rust
// CJK case – always join adjacent items
let is_cjk = prev_last.is_some_and(is_cjk_char) || curr_first.is_some_and(is_cjk_char);
if is_cjk {
    return gap < char_width * 0.8;
}

```

When either adjacent item contains CJK characters, pdf-inspector joins them unless the geometric gap exceeds 80% of a character width. This geometric heuristic replaces space-based word detection entirely.

For RTL lines, gap calculation itself is direction-aware (lines 440-452):

```rust
let gap = if prev_item.x <= curr_item.x {
    // LTR
    curr_item.x - (prev_item.x + prev_item.width)
} else {
    // RTL
    prev_item.x - (curr_item.x + curr_item.width)
};

```

The same threshold constants work correctly regardless of text direction.

## Arabic Presentation Forms and Visual-Order Reversal

PDFs often store Arabic using **presentation forms** (glyph variants like `U+FEE1` for isolated/medial/final shapes). These are stored in visual order—opposite to logical reading order. The `expand_ligatures` function (lines 184-250) resolves this:

1. **Detection**: `is_arabic_presentation_form` identifies visual-form code-points
2. **Normalization**: NFKC conversion transforms presentation forms to base Arabic characters
3. **Reversal**: `reverse_visual_arabic` reorders purely RTL runs while preserving embedded LTR sequences (numbers, Latin text)

This two-step pipeline ensures Arabic text emerges in correct logical order, not the reversed visual order from the PDF content stream.

## Practical Example: RTL Line Processing

```rust
use pdf_inspector::text_utils::sort_line_items;
use pdf_inspector::types::TextItem;

// Arabic text stored left-to-right in PDF (visual order)
let mut items = vec![
    TextItem::new("العالم", 20.0, 0.0, 5.0, "Font", 12.0),  // "world" at left
    TextItem::new("مرحبا", 30.0, 0.0, 5.0, "Font", 12.0),   // "hello" at right
];

sort_line_items(&mut items);
// Now items[0] = "مرحبا", items[1] = "العالم" — correct reading order

```

## Key Implementation Files

| File | Role in RTL/CJK Handling |
|------|--------------------------|
| [`src/text_utils.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_utils.rs) | Core detection functions, sorting, gap logic, Arabic normalization |
| [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) | High-level extraction; invokes `sort_line_items` after line construction |
| [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs) | Column detection; sorting for multi-column RTL documents |
| [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) | Final Markdown conversion; last-chance RTL sorting |
| [`src/types.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs) | `TextItem` struct with x, y, width fields used for geometric calculations |

## Summary

- **Character detection**: `is_rtl_char` and `is_cjk_char` provide fast Unicode-based script identification
- **Line direction**: `is_rtl_text` counts characters to determine dominant direction
- **Reading order**: `sort_line_items` reverses X-coordinates for RTL lines to preserve logical sequence
- **CJK spacing**: Geometric gap thresholds replace space-based word detection; CJK items always join unless gaps are excessive
- **Arabic normalization**: `expand_ligatures` converts presentation forms and reverses visual order to logical order

## Frequently Asked Questions

### How does pdf-inspector determine if a PDF line is RTL or LTR?

The `is_rtl_text` function in [`src/text_utils.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_utils.rs) iterates all characters in a line's text items, counting RTL code-points against LTR alphabetic characters. RTL wins when its count exceeds LTR. CJK characters are excluded from this tally to avoid skewing direction detection in mixed-script documents.

### Why doesn't CJK text need spaces between words in the output?

CJK scripts (Chinese, Japanese, Korean) don't use word-separating spaces. pdf-inspector detects CJK characters via `is_cjk_char` and applies geometric joining logic in `should_join_items`—items join when their gap falls below 80% of character width, regardless of space presence in the PDF.

### What problem does Arabic presentation form handling solve?

PDFs store Arabic in visual order using presentation forms (glyph variants for letter positions). `expand_ligatures` applies NFKC normalization to convert these to standard Arabic, then `reverse_visual_arabic` reorders the characters to logical reading order while preserving embedded numbers or Latin text.

### Does pdf-inspector support bidirectional text mixing RTL and LTR?

Yes. The RTL detection and sorting operate at the line level, so lines with mixed content are classified by their dominant direction. Within Arabic processing, `reverse_visual_arabic` specifically preserves embedded LTR runs (like phone numbers or URLs) without reversing them, maintaining correct bidirectional presentation.