How pdf-inspector Handles RTL Text and CJK Character Extraction in PDFs
pdf-inspector treats Right-to-Left (RTL) and CJK scripts as first-class citizens through character-level detection, line direction analysis, and script-aware layout algorithms that preserve logical reading order.
The firecrawl/pdf-inspector Rust library specializes in converting PDF documents to clean Markdown. Unlike generic PDF extractors, it implements dedicated handling for RTL text extraction (Arabic, Hebrew, Syriac) and CJK character extraction (Chinese, Japanese, Korean) through a multi-layered pipeline in src/text_utils.rs.
Character-Level Classification for RTL and CJK Scripts
The foundation of pdf-inspector's script handling lies in two fast Unicode range checks:
is_rtl_char(lines 97-112): Detects Hebrew (U+0590–U+05FF), Arabic (U+0600–U+06FF), Syriac, Thaana, NKo, Samaritan, Mandaic, and other RTL scriptsis_cjk_char(lines 80-95): Matches Hangul Jamo, Hiragana/Katakana, CJK Unified Ideographs, Compatibility Ideographs, and Half-width/Full-width forms
These functions are invoked throughout the extraction pipeline whenever script-sensitive decisions are required—gap calculations, space insertion, and sorting logic all depend on these classifications.
Detecting Line Direction with is_rtl_text
Before processing any line, pdf-inspector determines its dominant direction using is_rtl_text (lines 120-136):
let rtl = is_rtl_text(items.iter().map(|i| &i.text));
This function iterates every character across all text items in a line, counting RTL code-points against LTR alphabetic characters (excluding CJK). When RTL characters outnumber LTR, the entire line is flagged as right-to-left. This directional flag then drives all subsequent layout decisions.
Reordering RTL Lines for Logical Reading Order
When a line is identified as RTL, sort_line_items (lines 138-145) reverses the X-coordinate ordering:
pub(crate) fn sort_line_items(items: &mut [TextItem]) {
let rtl = is_rtl_text(items.iter().map(|i| &i.text));
if rtl {
// Reverse X-coordinates so the line reads right-to-left.
items.sort_by(|a, b| b.x.total_cmp(&a.x));
} else {
items.sort_by(|a, b| a.x.total_cmp(&b.x));
}
}
This sorting occurs in three critical locations:
src/extractor/mod.rs(line 2788) — after line collectionsrc/extractor/layout.rs(line 2258) — after column detectionsrc/lib.rs(line 3313) — final conversion to Markdown
The result: RTL text appears in logical reading order in the output, not the visual left-to-right storage order common in PDF internals.
CJK-Aware Join Logic: Spaces vs. Geometry
CJK scripts present a unique challenge—they don't use word separators. The should_join_items function handles this with script-specific logic:
// CJK case – always join adjacent items
let is_cjk = prev_last.is_some_and(is_cjk_char) || curr_first.is_some_and(is_cjk_char);
if is_cjk {
return gap < char_width * 0.8;
}
When either adjacent item contains CJK characters, pdf-inspector joins them unless the geometric gap exceeds 80% of a character width. This geometric heuristic replaces space-based word detection entirely.
For RTL lines, gap calculation itself is direction-aware (lines 440-452):
let gap = if prev_item.x <= curr_item.x {
// LTR
curr_item.x - (prev_item.x + prev_item.width)
} else {
// RTL
prev_item.x - (curr_item.x + curr_item.width)
};
The same threshold constants work correctly regardless of text direction.
Arabic Presentation Forms and Visual-Order Reversal
PDFs often store Arabic using presentation forms (glyph variants like U+FEE1 for isolated/medial/final shapes). These are stored in visual order—opposite to logical reading order. The expand_ligatures function (lines 184-250) resolves this:
- Detection:
is_arabic_presentation_formidentifies visual-form code-points - Normalization: NFKC conversion transforms presentation forms to base Arabic characters
- Reversal:
reverse_visual_arabicreorders purely RTL runs while preserving embedded LTR sequences (numbers, Latin text)
This two-step pipeline ensures Arabic text emerges in correct logical order, not the reversed visual order from the PDF content stream.
Practical Example: RTL Line Processing
use pdf_inspector::text_utils::sort_line_items;
use pdf_inspector::types::TextItem;
// Arabic text stored left-to-right in PDF (visual order)
let mut items = vec![
TextItem::new("العالم", 20.0, 0.0, 5.0, "Font", 12.0), // "world" at left
TextItem::new("مرحبا", 30.0, 0.0, 5.0, "Font", 12.0), // "hello" at right
];
sort_line_items(&mut items);
// Now items[0] = "مرحبا", items[1] = "العالم" — correct reading order
Key Implementation Files
| File | Role in RTL/CJK Handling |
|---|---|
src/text_utils.rs |
Core detection functions, sorting, gap logic, Arabic normalization |
src/extractor/mod.rs |
High-level extraction; invokes sort_line_items after line construction |
src/extractor/layout.rs |
Column detection; sorting for multi-column RTL documents |
src/lib.rs |
Final Markdown conversion; last-chance RTL sorting |
src/types.rs |
TextItem struct with x, y, width fields used for geometric calculations |
Summary
- Character detection:
is_rtl_charandis_cjk_charprovide fast Unicode-based script identification - Line direction:
is_rtl_textcounts characters to determine dominant direction - Reading order:
sort_line_itemsreverses X-coordinates for RTL lines to preserve logical sequence - CJK spacing: Geometric gap thresholds replace space-based word detection; CJK items always join unless gaps are excessive
- Arabic normalization:
expand_ligaturesconverts presentation forms and reverses visual order to logical order
Frequently Asked Questions
How does pdf-inspector determine if a PDF line is RTL or LTR?
The is_rtl_text function in src/text_utils.rs iterates all characters in a line's text items, counting RTL code-points against LTR alphabetic characters. RTL wins when its count exceeds LTR. CJK characters are excluded from this tally to avoid skewing direction detection in mixed-script documents.
Why doesn't CJK text need spaces between words in the output?
CJK scripts (Chinese, Japanese, Korean) don't use word-separating spaces. pdf-inspector detects CJK characters via is_cjk_char and applies geometric joining logic in should_join_items—items join when their gap falls below 80% of character width, regardless of space presence in the PDF.
What problem does Arabic presentation form handling solve?
PDFs store Arabic in visual order using presentation forms (glyph variants for letter positions). expand_ligatures applies NFKC normalization to convert these to standard Arabic, then reverse_visual_arabic reorders the characters to logical reading order while preserving embedded numbers or Latin text.
Does pdf-inspector support bidirectional text mixing RTL and LTR?
Yes. The RTL detection and sorting operate at the line level, so lines with mixed content are classified by their dominant direction. Within Arabic processing, reverse_visual_arabic specifically preserves embedded LTR runs (like phone numbers or URLs) without reversing them, maintaining correct bidirectional presentation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →