How pdf-inspector Handles Right-to-Left (RTL) Text and CJK Characters
pdf-inspector detects RTL scripts and CJK characters using Unicode range checks in src/text_utils.rs, then applies direction-aware sorting and tokenization to preserve logical reading order during PDF-to-Markdown extraction.
The firecrawl/pdf-inspector library is a Rust-based tool that converts PDF documents into compact, semantically-rich Markdown. Because PDF stores glyphs in visual order rather than logical reading order, the extractor must explicitly identify Right-to-Left (RTL) text (such as Arabic and Hebrew) and CJK characters (Chinese, Japanese, Korean) to reconstruct proper word boundaries and line ordering.
Character Classification Utilities
All low-level script detection lives in src/text_utils.rs, which exports three primary classification functions used throughout the extraction pipeline.
Detecting CJK Characters
The is_cjk_char function returns true for Unicode code points belonging to CJK blocks, including Hangul Jamo, Hiragana, Katakana, and CJK Unified Ideographs. According to the source code at src/text_utils.rs#L90-L102, this check suppresses space-based word-boundary heuristics because CJK text does not use explicit word separators.
Detecting RTL Scripts
For RTL detection, is_rtl_char identifies characters in the Unicode ranges for Hebrew, Arabic, Syriac, Thaana, N'Ko, Samaritan, and Mandaic scripts as implemented at src/text_utils.rs#L104-L118. This mirrors the script blocks defined by the Unicode Standard.
Line-Level Directionality Detection
The is_rtl_text function at src/text_utils.rs#L127-L143 accepts an iterator of strings and implements a majority-vote algorithm. It counts RTL versus left-to-right (LTR) alphabetic characters (excluding CJK), returning true only when rtl > 0 && rtl > ltr. This determines whether an entire line or paragraph should be treated as RTL.
How Directionality Affects the Extraction Pipeline
The detection utilities integrate into four critical stages of the extraction process to ensure semantic correctness.
Sorting Line Items by Reading Order
When building a line from a collection of TextItem structs, the extractor calls sort_line_items (also in src/text_utils.rs). This function first queries is_rtl_text to determine directionality:
- If RTL: Items are sorted by decreasing
xcoordinate (b.x.total_cmp(&a.x)) - If LTR: Items are sorted by increasing
xcoordinate
// src/text_utils.rs
pub(crate) fn sort_line_items(items: &mut [TextItem]) {
let rtl = is_rtl_text(items.iter().map(|i| &i.text));
if rtl {
items.sort_by(|a, b| b.x.total_cmp(&a.x));
} else {
items.sort_by(|a, b| a.x.total_cmp(&b.x));
}
}
This ensures that visual order matches logical reading order regardless of script direction.
Word-Boundary Heuristics for CJK
Functions that split lines into words—located in src/markdown/preprocess.rs and src/markdown/analysis.rs—first invoke is_cjk_char. When CJK characters are present, the code skips space-based splitting, treating the entire run as a single token. This prevents false word breaks inside Chinese or Japanese text that lacks inter-word spaces.
Reading-Order Detection
The src/extractor/reading_order.rs module combines is_cjk_char with is_rtl_text to determine whether a block should be read left-to-right, right-to-left, or top-to-bottom (for vertical CJK). The module filters character streams using these utilities to establish the correct extraction sequence. The main extraction flow in src/extractor/mod.rs integrates these checks and contains assertions verifying detection correctness.
// extractor/reading_order.rs (excerpt)
use crate::text_utils::{effective_width, is_cjk_char, is_rtl_text};
// … later …
.filter(|character| is_cjk_char(*character))
Post-Processing Adjustments
After the Markdown tree is constructed, src/markdown/postprocess.rs references RTL flags to correctly attach punctuation and avoid re-ordering already-corrected lines. This final pass ensures that extracted Markdown maintains proper bi-directional formatting.
Practical Code Examples
You can leverage these utilities directly or rely on the high-level API that automatically applies them.
Example 1: Manual RTL Detection
use pdf_inspector::text_utils::is_rtl_text;
/// Detect if a paragraph is RTL.
fn paragraph_is_rtl(paragraph: &str) -> bool {
// `is_rtl_text` expects an iterator of strings; a single paragraph is wrapped in a vec.
is_rtl_text(std::iter::once(paragraph))
}
fn main() {
let txt = "مرحبا بالعالم"; // Arabic
println!("RTL? {}", paragraph_is_rtl(txt)); // → true
}
Example 2: High-Level API Usage
use pdf_inspector::process_pdf_with_options;
use pdf_inspector::process_mode::ProcessMode;
fn main() -> Result<(), Box<dyn std::error::Error>> {
// Options enable default detection of RTL and CJK.
let opts = pdf_inspector::Options {
mode: ProcessMode::Extract,
..Default::default()
};
// `process_pdf_with_options` returns a structured Markdown string.
let markdown = process_pdf_with_options("sample_rtl.pdf", opts)?;
println!("{}", markdown);
}
In this example, RTL lines in sample_rtl.pdf are automatically sorted in reverse x order, while CJK runs remain unsplit.
Example 3: Sorting TextItem Vectors
use pdf_inspector::text_utils::sort_line_items;
use pdf_inspector::types::TextItem;
fn main() {
let mut items = vec![
TextItem { text: "العالم".into(), x: 100.0, ..Default::default() },
TextItem { text: "مرحبا".into(), x: 200.0, ..Default::default() },
];
// The helper detects RTL and sorts accordingly.
sort_line_items(&mut items);
// Items are now ordered from right-most to left-most.
for i in items {
println!("{} @ {}", i.text, i.x);
}
}
Summary
- pdf-inspector detects CJK and RTL scripts using Unicode range checks in
src/text_utils.rsviais_cjk_char,is_rtl_char, andis_rtl_text. - RTL lines are sorted by decreasing x-coordinate in
sort_line_itemsto preserve logical reading order. - CJK text bypasses space-based tokenization to prevent artificial word breaks in scripts without explicit separators.
- The extraction pipeline integrates these checks in
extractor/reading_order.rs,extractor/mod.rs, and both pre/post-processing Markdown modules. - Both low-level utilities and high-level APIs automatically apply these directionality corrections.
Frequently Asked Questions
How does pdf-inspector determine if a line of text is RTL?
The library uses the is_rtl_text function in src/text_utils.rs#L127-L143 to count RTL versus LTR alphabetic characters. If the RTL count exceeds zero and is greater than the LTR count, the line is classified as RTL. This majority-vote approach handles mixed-script content gracefully.
Why does CJK text require special handling during extraction?
CJK scripts (Chinese, Japanese, Korean) do not use spaces to separate words. Without detection, standard word-boundary heuristics would incorrectly split characters or insert artificial spaces. The is_cjk_char function identifies these scripts so the extractor treats text runs as single tokens, preserving semantic boundaries for downstream LLM or indexing tasks.
Can I manually check if a string contains RTL characters using the library?
Yes. Import pdf_inspector::text_utils::is_rtl_text and pass an iterator of strings to detect RTL content at the paragraph or line level. For character-level detection, use is_rtl_char which checks individual Unicode code points against RTL script blocks including Arabic, Hebrew, and Syriac.
How does the sorting algorithm differ for RTL versus LTR text?
In sort_line_items (src/text_utils.rs), LTR text sorts by increasing x-coordinate (left-to-right), while RTL text sorts by decreasing x-coordinate (right-to-left). This coordinate reversal ensures that the extracted Markdown reflects logical reading order rather than the visual glyph order stored in the PDF.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →