When to Expect Encoding Issues with pdf-inspector and When to Fall Back to OCR
Fall back to OCR when pdf-inspector detects replacement characters (U+FFFD), dollar-as-space patterns, substitution-cipher style garbling, CID-only fonts, or extremely low-quality text extraction.
The pdf-inspector library extracts text from PDFs by interpreting font-encoding information, specifically the ToUnicode CMap that maps character codes to Unicode. When this mapping fails, the library automatically flags affected pages for OCR fallback. Understanding these detection mechanisms helps you predict when the extraction pipeline will need OCR and when you can trust the direct text output.
What Triggers Encoding Issues in pdf-inspector
PDFs encode text through complex font systems. Most modern PDFs include reliable Unicode mappings, but legacy files, scanned documents, and malformed exports often break this pipeline. The pdf-inspector source code identifies five specific failure modes that trigger automatic OCR fallback.
Replacement Characters (U+FFFD)
The most immediate sign of encoding failure appears when the decoder cannot map a byte sequence to a valid Unicode code point. This produces the Unicode replacement character U+FFFD (�), which pdf-inspector treats as a hard failure.
The detect_encoding_issues() function in src/text_quality.rs short-circuits on the first occurrence of '\u{FFFD}':
// src/text_quality.rs lines 31-42
pub fn detect_encoding_issues(text: &str) -> bool {
text.contains('\u{FFFD}')
}
This check runs early in the pipeline. If any page contains replacement characters, it is immediately flagged for OCR without further analysis.
Dollar-as-Space Pattern
Broken CMaps frequently misinterpret the dollar sign ($) as a word separator. This produces strings like Word$Word$Word where spaces should appear.
The has_dollar_as_space_pattern() function detects this through statistical analysis:
// src/text_quality.rs lines 57-73
pub fn has_dollar_as_space_pattern(text: &str) -> bool {
let total_dollars = text.matches('$').count();
let letter_dollar_letter = text
.matches(|c: char| c.is_ascii_alphabetic())
.filter(/* ... */)
.count();
letter_dollar_letter > 20 || letter_dollar_letter * 2 > total_dollars
}
A page triggers OCR when either:
- More than 20 letter-dollar-letter sequences appear
- Letter-dollar-letter sequences exceed half of all dollar characters
Substitution-Cipher Style Garbling
Some PDFs apply systematic character shifts that preserve letter shapes but scramble meaning. For example, "Certificate" becomes "8VceZWZTReV".
The CipherGarbleStats struct in src/text_quality.rs detects this through histogram comparison:
// src/text_quality.rs lines 78-110
impl CipherGarbleStats {
pub fn looks_garbled(&self) -> bool {
let english_cosine = self.english_cosine();
let shape_cosine = self.english_shape_cosine();
self.sample_size > MIN_SAMPLE_SIZE
&& english_cosine < ENGLISH_COSINE_THRESHOLD
&& shape_cosine < SHAPE_COSINE_THRESHOLD
}
}
The implementation builds ASCII letter histograms and measures cosine similarity against expected English letter frequency distributions. When both similarity scores fall below empirically tuned thresholds, the page is considered corrupted.
CID-Only Fonts
Some fonts expose only Character IDs (CIDs)—glyph indices without Unicode mappings. These fonts typically set their Encoding to Identity-H or Identity-V.
The extractor in src/extractor/fonts.rs detects these cases, and is_cid_garbage() reports the problem. Pages containing CID-only fonts are automatically routed to OCR.
Empty or Extremely Low-Quality Text
The final quality gate combines multiple heuristics. The has_text_quality_issue function integrates:
is_garbage_text()results- Empty-text detection
- All encoding heuristics above
If any flag is true, OCR is forced regardless of other quality metrics.
How pdf-inspector Decides to Fall Back to OCR
The decision pipeline lives in src/lib.rs (lines 820-830). When processing completes, the library evaluates each page against the detection criteria:
// Conceptual flow based on src/lib.rs implementation
for (page_idx, page_text) in extracted_pages.iter().enumerate() {
let mut ocr_reason = None;
if detect_encoding_issues(page_text) {
ocr_reason = Some(EncodingIssue::ReplacementCharacter);
} else if has_dollar_as_space_pattern(page_text) {
ocr_reason = Some(EncodingIssue::DollarAsSpace);
} else if CipherGarbleStats::analyze(page_text).looks_garbled() {
ocr_reason = Some(EncodingIssue::SubstitutionCipher);
} else if is_cid_garbage(&page_fonts) {
ocr_reason = Some(EncodingIssue::CidOnlyFont);
} else if has_text_quality_issue(page_text) {
ocr_reason = Some(EncodingIssue::LowQuality);
}
if let Some(reason) = ocr_reason {
ocr_queue.push(page_idx);
ocr_reasons.push(reason);
}
}
Flagged pages are added to an OCR queue and processed by the downstream OCR engine (typically Tesseract). Unflagged pages bypass OCR entirely, preserving processing speed.
Using pdf-inspector's OCR Fallback in Practice
Rust API: Checking OCR Decisions Programmatically
use pdf_inspector::process_pdf_with_options;
use pdf_inspector::options::ProcessOptions;
let opts = ProcessOptions {
force_ocr_all: false, // Let the library decide based on encoding issues
..Default::default()
};
let result = process_pdf_with_options("sample.pdf", opts);
for (i, page) in result.pages.iter().enumerate() {
if page.needs_ocr {
println!("Page {} triggered OCR fallback", i + 1);
// Inspect the specific reason
if let Some(reason) = &page.ocr_reason {
println!(" Reason: {:?}", reason);
}
}
}
CLI: Automatic OCR with JSON Reporting
The pdf2md binary (source: src/bin/pdf2md.rs) handles OCR decisions automatically:
# Process with automatic OCR fallback
$ pdf2md sample.pdf --json > output.json
# Review which pages required OCR
$ jq '.pages[] | select(.needs_ocr) | {page: .page_number, reason: .ocr_reason}' output.json
Each page in the JSON output includes a needs_ocr boolean and ocr_reason field when applicable.
Python API: Accessing Encoding Issue Flags
from pdf_inspector import PdfInspector
inspector = PdfInspector()
result = inspector.process("sample.pdf")
for i, page in enumerate(result.pages, 1):
if page.needs_ocr:
print(f"Page {i} required OCR due to: {page.ocr_reason}")
elif page.has_encoding_issues:
# Encoding issues were detected but possibly resolved
print(f"Page {i} had encoding warnings")
The Python bindings (defined in src/python.rs) expose has_encoding_issues and needs_ocr attributes on PdfResult objects.
Forcing OCR When Automatic Detection Is Insufficient
You can bypass the automatic detection entirely through the force_ocr_all flag:
let opts = ProcessOptions {
force_ocr_all: true, // OCR every page regardless of quality
..Default::default()
};
This is useful when:
- You know the PDF contains complex layouts that fool the heuristics
- You need consistent processing regardless of per-page quality
- The automatic detection misses subtle encoding issues in your document type
Key Source Files for Understanding Encoding Detection
| File | Responsibility |
|---|---|
src/text_quality.rs |
Core heuristics: detect_encoding_issues, has_dollar_as_space_pattern, CipherGarbleStats |
src/lib.rs |
Orchestration: builds OCR queue based on quality signals (lines 820-830) |
src/detector.rs |
PDF type classification; coordinates with text-quality module |
src/extractor/fonts.rs |
Font encoding extraction; CID-only font detection |
src/bin/pdf2md.rs |
CLI interface; reports OCR decisions in JSON |
src/python.rs |
Python API bindings for encoding issue flags |
Summary
- Expect encoding issues when PDFs lack proper ToUnicode CMaps, use legacy encodings, contain malformed CMaps, or rely on CID-only fonts.
- Automatic OCR fallback triggers on: replacement characters (U+FFFD), dollar-as-space patterns, substitution-cipher statistics, CID-only fonts, or low overall text quality.
- Detection is fast because
pdf-inspectorruns heuristics before invoking expensive OCR operations. - Override is available through
force_ocr_allwhen you need guaranteed OCR coverage. - Results are inspectable via
needs_ocrandocr_reasonfields in all API bindings.
Frequently Asked Questions
How accurate is pdf-inspector's automatic OCR detection?
The heuristics are tuned for common PDF failure modes based on empirical analysis of production documents. According to the pdf-inspector source code, the substitution-cipher detection uses cosine similarity thresholds that minimize false positives while catching systematic garbling. For critical applications, inspect the ocr_reason field to verify detection rationale.
Can I disable automatic OCR and handle encoding issues manually?
Yes. Set force_ocr_all: false (the default) and check has_encoding_issues or needs_ocr in the result. The raw extracted text remains available even when encoding issues are detected—you decide whether to use it, run custom processing, or invoke OCR separately.
Why does pdf-inspector sometimes miss subtle encoding problems?
The detection prioritizes speed and avoids false positives. The CipherGarbleStats implementation requires minimum sample sizes before evaluating English similarity, and single-character substitutions may evade detection. For documents where any encoding error is unacceptable, use force_ocr_all or implement additional validation.
What OCR engine does pdf-inspector use?
The library itself does not embed an OCR engine. When pages are flagged for OCR, the binary invokes a downstream processor (typically Tesseract through an external call). The OCR step occurs after the pdf-inspector quality analysis completes, as orchestrated in src/lib.rs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →