How pdf-inspector Performs Text Quality Analysis and Detects Garbage Text/Encoding Issues in Rust
pdf-inspector uses a dual-layer detection system—span-level TextItem analysis and page-level Markdown screening—to identify broken font encodings, malformed CID-to-Unicode mappings, and garbled PDF text before it reaches downstream consumers.
The pdf-inspector library, developed by Firecrawl, implements a comprehensive text quality analysis pipeline that automatically flags unreliable extracted text and routes problematic pages to OCR. This article explains how the Rust codebase detects encoding failures, garbage text, and substitution-cipher corruption using production-ready heuristics.
Overview of the Two-Layer Detection Architecture
pdf-inspector evaluates text reliability through complementary approaches:
- Span-level analysis—examines individual
TextItemobjects for character-level corruption - Markdown-level analysis—screens rendered page text for encoding artifacts and structural garbage
Both layers feed into a final decision in analyze_text_quality (src/text_quality.rs, line 50) that produces a TextQualityReport. This report drives per-page OCR routing and sets a document-wide has_encoding_issues flag.
Document-Wide Processing Flow
The entry points process_pdf and extract_pages_markdown_mem (in src/lib.rs, lines 84-102) orchestrate quality detection:
// Extract all TextItems, then run quality analysis
let quality_report = analyze_text_quality(&all_items);
// Later, during per-page markdown generation...
for page in pages {
if detect_encoding_issues(&page.markdown)
|| is_cid_garbage(&page.markdown)
|| is_garbage_text(&page.markdown)
{
page.needs_ocr = true;
page.ocr_reason = Some(OCR_REASON_SUSPECTED_GARBLED_TEXT);
}
}
The three Markdown-level checks operate independently—any single failure triggers OCR fallback with a human-readable reason.
Span-Level TextItem Analysis
The analyze_text_quality function iterates over every TextItem (defined in src/types.rs) and classifies evidence into two severity tiers.
Strong Issues (Immediate OCR Trigger)
These defects guarantee encoding failure and bypass accumulation:
- Replacement characters (U+FFFD)—detected by
has_replacement_text_run(src/text_quality.rs, lines 73-76) - Private-Use Area / C1 control runs—detected by
has_private_use_text_run(lines 78-84) - Dollar-as-space pattern—detected by
has_dollar_as_space_pattern(lines 57-70), where$$sequences indicate broken PDF-to-text conversion
Replacement Issues (Accumulated Evidence)
Suspect character runs are tallied per-page, then evaluated by:
- Cipher-garble statistics—collected in
CipherGarbleStats(lines 87-124) - Final judgment—
looks_garbled(lines 185-225) applies statistical thresholds
The helper page_replacement_evidence_needs_ocr (lines 53-71) makes the final page-level decision based on replacement density, maximum run length, and cipher-garble detection.
Markdown-Level Quality Helpers
Three specialized functions screen rendered page text before output delivery.
detect_encoding_issues (Lines 31-55)
Combines three signals into a boolean verdict:
- Replacement character presence
- Dollar-as-space pattern detection
- Substitution-cipher signature matching
Returns true if any indicator fires.
is_garbage_text (Lines 31-71)
Implements a statistical garbage detector:
- Ignores markdown syntax characters (
#,*,`, etc.) - Counts alphanumeric vs. non-alphanumeric characters
- Requires minimum 50 characters to avoid false positives on short strings
- Flags text as garbage when non-alphanumeric proportion exceeds 50%
is_cid_garbage (Lines 73-120)
Specialized for CID-keyed font corruption:
- First delegates to
is_garbage_textfor general screening - If that passes, checks for C1 control character excess or high-Latin-1 anomalies—patterns typical of malformed CID-to-Unicode CMap tables
Working with the Text Quality API
Full Pipeline with Quality Reporting
use pdf_inspector::{process_pdf, OCR_REASON_SUSPECTED_GARBLED_TEXT};
fn main() -> Result<(), pdf_inspector::PdfError> {
let result = process_pdf("sample.pdf")?;
println!("Pages: {}", result.page_count);
println!("OCR required: {:?}", result.pages_needing_ocr);
if result.has_encoding_issues {
eprintln!("Document has encoding issues—using OCR fallback");
}
Ok(())
}
Per-Page Extraction with Explicit Flags
use pdf_inspector::extract_pages_markdown_mem;
let bytes = std::fs::read("sample.pdf")?;
let pages = extract_pages_markdown_mem(&bytes, None)?;
for page in pages.pages {
match page.needs_ocr {
true => println!(
"Page {} → OCR (reason: {:?})",
page.page + 1,
page.ocr_reason
),
false => println!(
"Page {} → Clean markdown:\n{}",
page.page + 1,
page.markdown
),
}
}
Direct Low-Level Checks for Custom Pipelines
let markdown = "# Title\n$$$ 文字 $$$";
let has_encoding = pdf_inspector::detect_encoding_issues(markdown);
let is_garbage = pdf_inspector::is_garbage_text(markdown);
let is_cid_bad = pdf_inspector::is_cid_garbage(markdown);
println!(
"encoding={}, garbage={}, cid_garbage={}",
has_encoding, is_garbage, is_cid_bad
);
// Output: encoding=true, garbage=false, cid_garbage=true
Key Source Files and Responsibilities
| File | Role in Text Quality Analysis |
|---|---|
| src/text_quality.rs | Core heuristic implementations: detect_encoding_issues, is_garbage_text, is_cid_garbage, analyze_text_quality, and cipher-garble detection |
| src/lib.rs | Orchestration logic—calls analyze_text_quality, applies Markdown checks, routes to OCR |
| src/types.rs | TextItem and ItemType definitions used for span-level inspection |
| src/markdown/*.rs | Markdown generation consumed by quality screening functions |
Summary
- pdf-inspector implements dual-layer detection: span-level
TextIteminspection plus page-level Markdown screening - Strong signals (replacement chars, private-use areas, dollar-as-space) trigger immediate OCR
- Accumulated evidence (replacement density, cipher-garble statistics) triggers OCR when thresholds are exceeded
- Markdown-level heuristics catch encoding artifacts that survive span-level filtering
- The
has_encoding_issuesdocument flag andOCR_REASON_SUSPECTED_GARBLED_TEXTreason provide actionable downstream signals
Frequently Asked Questions
What is the dollar-as-space pattern in PDF text extraction?
The dollar-as-space pattern occurs when $ characters appear in runs like $$ or $$$ as a byproduct of broken font encoding or PDF-to-text conversion. According to the pdf-inspector source code in src/text_quality.rs (lines 57-70), this pattern is detected by has_dollar_as_space_pattern and treated as a strong encoding issue that immediately flags a page for OCR.
How does pdf-inspector distinguish between garbage text and legitimate non-English content?
The is_garbage_text function in src/text_quality.rs (lines 31-71) uses character class proportions, not language detection. It ignores markdown syntax and requires text to exceed 50 characters. If more than 50% of remaining characters are non-alphanumeric, the text is flagged. This heuristic catches encoding corruption while preserving legitimate Chinese, Arabic, or Cyrillic text—since these scripts still produce substantial alphanumeric output in Unicode.
When does pdf-inspector set the has_encoding_issues document-wide flag?
The has_encoding_issues flag is set in analyze_text_quality (src/text_quality.rs, line 50) when any page in the document exhibits strong encoding failures or accumulated replacement evidence exceeding OCR thresholds. This allows downstream systems to apply document-level OCR strategies rather than per-page handling.
Can I use pdf-inspector's quality checks without the full extraction pipeline?
Yes. The quality functions are exposed for direct use: detect_encoding_issues, is_garbage_text, and is_cid_garbage operate on &str inputs. Import them for custom PDF processing workflows that need encoding detection without full markdown generation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →