# How pdf-inspector Handles Right-to-Left (RTL) Text and CJK Characters

> Discover how pdf-inspector accurately processes Right-to-Left text and CJK characters using Unicode checks and advanced sorting for precise PDF to Markdown conversion. Ensure correct reading order.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: deep-dive
- Published: 2026-08-08

---

**pdf-inspector detects RTL scripts and CJK characters using Unicode range checks in [`src/text_utils.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_utils.rs), then applies direction-aware sorting and tokenization to preserve logical reading order during PDF-to-Markdown extraction.**

The `firecrawl/pdf-inspector` library is a Rust-based tool that converts PDF documents into compact, semantically-rich Markdown. Because PDF stores glyphs in visual order rather than logical reading order, the extractor must explicitly identify **Right-to-Left (RTL)** text (such as Arabic and Hebrew) and **CJK** characters (Chinese, Japanese, Korean) to reconstruct proper word boundaries and line ordering.

## Character Classification Utilities

All low-level script detection lives in **[`src/text_utils.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_utils.rs)**, which exports three primary classification functions used throughout the extraction pipeline.

### Detecting CJK Characters

The `is_cjk_char` function returns `true` for Unicode code points belonging to CJK blocks, including Hangul Jamo, Hiragana, Katakana, and CJK Unified Ideographs. According to the source code at [`src/text_utils.rs#L90-L102`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_utils.rs#L90-L102), this check suppresses space-based word-boundary heuristics because CJK text does not use explicit word separators.

### Detecting RTL Scripts

For RTL detection, `is_rtl_char` identifies characters in the Unicode ranges for Hebrew, Arabic, Syriac, Thaana, N'Ko, Samaritan, and Mandaic scripts as implemented at [`src/text_utils.rs#L104-L118`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_utils.rs#L104-L118). This mirrors the script blocks defined by the Unicode Standard.

### Line-Level Directionality Detection

The `is_rtl_text` function at [`src/text_utils.rs#L127-L143`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_utils.rs#L127-L143) accepts an iterator of strings and implements a majority-vote algorithm. It counts RTL versus left-to-right (LTR) alphabetic characters (excluding CJK), returning `true` only when `rtl > 0 && rtl > ltr`. This determines whether an entire line or paragraph should be treated as RTL.

## How Directionality Affects the Extraction Pipeline

The detection utilities integrate into four critical stages of the extraction process to ensure semantic correctness.

### Sorting Line Items by Reading Order

When building a line from a collection of `TextItem` structs, the extractor calls `sort_line_items` (also in **[`src/text_utils.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_utils.rs)**). This function first queries `is_rtl_text` to determine directionality:

- If **RTL**: Items are sorted by **decreasing** `x` coordinate (`b.x.total_cmp(&a.x)`)
- If **LTR**: Items are sorted by **increasing** `x` coordinate

```rust
// src/text_utils.rs
pub(crate) fn sort_line_items(items: &mut [TextItem]) {
    let rtl = is_rtl_text(items.iter().map(|i| &i.text));
    if rtl {
        items.sort_by(|a, b| b.x.total_cmp(&a.x));
    } else {
        items.sort_by(|a, b| a.x.total_cmp(&b.x));
    }
}

```

This ensures that visual order matches logical reading order regardless of script direction.

### Word-Boundary Heuristics for CJK

Functions that split lines into words—located in **[`src/markdown/preprocess.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/preprocess.rs)** and **[`src/markdown/analysis.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/analysis.rs)**—first invoke `is_cjk_char`. When CJK characters are present, the code **skips space-based splitting**, treating the entire run as a single token. This prevents false word breaks inside Chinese or Japanese text that lacks inter-word spaces.

### Reading-Order Detection

The **[`src/extractor/reading_order.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/reading_order.rs)** module combines `is_cjk_char` with `is_rtl_text` to determine whether a block should be read left-to-right, right-to-left, or top-to-bottom (for vertical CJK). The module filters character streams using these utilities to establish the correct extraction sequence. The main extraction flow in **[`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs)** integrates these checks and contains assertions verifying detection correctness.

```rust
// extractor/reading_order.rs (excerpt)
use crate::text_utils::{effective_width, is_cjk_char, is_rtl_text};
// … later …
.filter(|character| is_cjk_char(*character))

```

### Post-Processing Adjustments

After the Markdown tree is constructed, **[`src/markdown/postprocess.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/postprocess.rs)** references RTL flags to correctly attach punctuation and avoid re-ordering already-corrected lines. This final pass ensures that extracted Markdown maintains proper bi-directional formatting.

## Practical Code Examples

You can leverage these utilities directly or rely on the high-level API that automatically applies them.

### Example 1: Manual RTL Detection

```rust
use pdf_inspector::text_utils::is_rtl_text;

/// Detect if a paragraph is RTL.
fn paragraph_is_rtl(paragraph: &str) -> bool {
    // `is_rtl_text` expects an iterator of strings; a single paragraph is wrapped in a vec.
    is_rtl_text(std::iter::once(paragraph))
}

fn main() {
    let txt = "مرحبا بالعالم"; // Arabic
    println!("RTL? {}", paragraph_is_rtl(txt)); // → true
}

```

### Example 2: High-Level API Usage

```rust
use pdf_inspector::process_pdf_with_options;
use pdf_inspector::process_mode::ProcessMode;

fn main() -> Result<(), Box<dyn std::error::Error>> {
    // Options enable default detection of RTL and CJK.
    let opts = pdf_inspector::Options {
        mode: ProcessMode::Extract,
        ..Default::default()
    };

    // `process_pdf_with_options` returns a structured Markdown string.
    let markdown = process_pdf_with_options("sample_rtl.pdf", opts)?;
    println!("{}", markdown);
}

```

In this example, RTL lines in `sample_rtl.pdf` are automatically sorted in reverse `x` order, while CJK runs remain unsplit.

### Example 3: Sorting TextItem Vectors

```rust
use pdf_inspector::text_utils::sort_line_items;
use pdf_inspector::types::TextItem;

fn main() {
    let mut items = vec![
        TextItem { text: "العالم".into(), x: 100.0, ..Default::default() },
        TextItem { text: "مرحبا".into(), x: 200.0, ..Default::default() },
    ];

    // The helper detects RTL and sorts accordingly.
    sort_line_items(&mut items);

    // Items are now ordered from right-most to left-most.
    for i in items {
        println!("{} @ {}", i.text, i.x);
    }
}

```

## Summary

- **pdf-inspector** detects CJK and RTL scripts using Unicode range checks in [`src/text_utils.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_utils.rs) via `is_cjk_char`, `is_rtl_char`, and `is_rtl_text`.
- **RTL lines** are sorted by decreasing x-coordinate in `sort_line_items` to preserve logical reading order.
- **CJK text** bypasses space-based tokenization to prevent artificial word breaks in scripts without explicit separators.
- The extraction pipeline integrates these checks in [`extractor/reading_order.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/extractor/reading_order.rs), [`extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/extractor/mod.rs), and both pre/post-processing Markdown modules.
- Both low-level utilities and high-level APIs automatically apply these directionality corrections.

## Frequently Asked Questions

### How does pdf-inspector determine if a line of text is RTL?

The library uses the `is_rtl_text` function in `src/text_utils.rs#L127-L143` to count RTL versus LTR alphabetic characters. If the RTL count exceeds zero and is greater than the LTR count, the line is classified as RTL. This majority-vote approach handles mixed-script content gracefully.

### Why does CJK text require special handling during extraction?

CJK scripts (Chinese, Japanese, Korean) do not use spaces to separate words. Without detection, standard word-boundary heuristics would incorrectly split characters or insert artificial spaces. The `is_cjk_char` function identifies these scripts so the extractor treats text runs as single tokens, preserving semantic boundaries for downstream LLM or indexing tasks.

### Can I manually check if a string contains RTL characters using the library?

Yes. Import `pdf_inspector::text_utils::is_rtl_text` and pass an iterator of strings to detect RTL content at the paragraph or line level. For character-level detection, use `is_rtl_char` which checks individual Unicode code points against RTL script blocks including Arabic, Hebrew, and Syriac.

### How does the sorting algorithm differ for RTL versus LTR text?

In `sort_line_items` ([`src/text_utils.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_utils.rs)), LTR text sorts by increasing x-coordinate (left-to-right), while RTL text sorts by decreasing x-coordinate (right-to-left). This coordinate reversal ensures that the extracted Markdown reflects logical reading order rather than the visual glyph order stored in the PDF.