# When to Expect Encoding Issues with pdf-inspector and When to Fall Back to OCR

> Learn when pdf-inspector encounters encoding issues like replacement characters or CID-only fonts. Discover when to switch to OCR for reliable text extraction.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: troubleshooting
- Published: 2026-08-06

---

**Fall back to OCR when pdf-inspector detects replacement characters (U+FFFD), dollar-as-space patterns, substitution-cipher style garbling, CID-only fonts, or extremely low-quality text extraction.**

The `pdf-inspector` library extracts text from PDFs by interpreting font-encoding information, specifically the *ToUnicode* CMap that maps character codes to Unicode. When this mapping fails, the library automatically flags affected pages for OCR fallback. Understanding these detection mechanisms helps you predict when the extraction pipeline will need OCR and when you can trust the direct text output.

## What Triggers Encoding Issues in pdf-inspector

PDFs encode text through complex font systems. Most modern PDFs include reliable Unicode mappings, but legacy files, scanned documents, and malformed exports often break this pipeline. The `pdf-inspector` source code identifies five specific failure modes that trigger automatic OCR fallback.

### Replacement Characters (U+FFFD)

The most immediate sign of encoding failure appears when the decoder cannot map a byte sequence to a valid Unicode code point. This produces the Unicode replacement character `U+FFFD` (�), which `pdf-inspector` treats as a hard failure.

The `detect_encoding_issues()` function in [`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs) short-circuits on the first occurrence of `'\u{FFFD}'`:

```rust
// src/text_quality.rs lines 31-42
pub fn detect_encoding_issues(text: &str) -> bool {
    text.contains('\u{FFFD}')
}

```

This check runs early in the pipeline. If any page contains replacement characters, it is immediately flagged for OCR without further analysis.

### Dollar-as-Space Pattern

Broken CMaps frequently misinterpret the dollar sign (`$`) as a word separator. This produces strings like `Word$Word$Word` where spaces should appear.

The `has_dollar_as_space_pattern()` function detects this through statistical analysis:

```rust
// src/text_quality.rs lines 57-73
pub fn has_dollar_as_space_pattern(text: &str) -> bool {
    let total_dollars = text.matches('$').count();
    let letter_dollar_letter = text
        .matches(|c: char| c.is_ascii_alphabetic())
        .filter(/* ... */)
        .count();
    
    letter_dollar_letter > 20 || letter_dollar_letter * 2 > total_dollars
}

```

A page triggers OCR when either:
- More than 20 letter-dollar-letter sequences appear
- Letter-dollar-letter sequences exceed half of all dollar characters

### Substitution-Cipher Style Garbling

Some PDFs apply systematic character shifts that preserve letter shapes but scramble meaning. For example, "Certificate" becomes "8VceZWZTReV".

The `CipherGarbleStats` struct in [`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs) detects this through histogram comparison:

```rust
// src/text_quality.rs lines 78-110
impl CipherGarbleStats {
    pub fn looks_garbled(&self) -> bool {
        let english_cosine = self.english_cosine();
        let shape_cosine = self.english_shape_cosine();
        
        self.sample_size > MIN_SAMPLE_SIZE
            && english_cosine < ENGLISH_COSINE_THRESHOLD
            && shape_cosine < SHAPE_COSINE_THRESHOLD
    }
}

```

The implementation builds ASCII letter histograms and measures cosine similarity against expected English letter frequency distributions. When both similarity scores fall below empirically tuned thresholds, the page is considered corrupted.

### CID-Only Fonts

Some fonts expose only **Character IDs (CIDs)**—glyph indices without Unicode mappings. These fonts typically set their *Encoding* to `Identity-H` or `Identity-V`.

The extractor in [`src/extractor/fonts.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/fonts.rs) detects these cases, and `is_cid_garbage()` reports the problem. Pages containing CID-only fonts are automatically routed to OCR.

### Empty or Extremely Low-Quality Text

The final quality gate combines multiple heuristics. The `has_text_quality_issue` function integrates:
- `is_garbage_text()` results
- Empty-text detection
- All encoding heuristics above

If any flag is true, OCR is forced regardless of other quality metrics.

## How pdf-inspector Decides to Fall Back to OCR

The decision pipeline lives in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) (lines 820-830). When processing completes, the library evaluates each page against the detection criteria:

```rust
// Conceptual flow based on src/lib.rs implementation
for (page_idx, page_text) in extracted_pages.iter().enumerate() {
    let mut ocr_reason = None;
    
    if detect_encoding_issues(page_text) {
        ocr_reason = Some(EncodingIssue::ReplacementCharacter);
    } else if has_dollar_as_space_pattern(page_text) {
        ocr_reason = Some(EncodingIssue::DollarAsSpace);
    } else if CipherGarbleStats::analyze(page_text).looks_garbled() {
        ocr_reason = Some(EncodingIssue::SubstitutionCipher);
    } else if is_cid_garbage(&page_fonts) {
        ocr_reason = Some(EncodingIssue::CidOnlyFont);
    } else if has_text_quality_issue(page_text) {
        ocr_reason = Some(EncodingIssue::LowQuality);
    }
    
    if let Some(reason) = ocr_reason {
        ocr_queue.push(page_idx);
        ocr_reasons.push(reason);
    }
}

```

Flagged pages are added to an OCR queue and processed by the downstream OCR engine (typically Tesseract). Unflagged pages bypass OCR entirely, preserving processing speed.

## Using pdf-inspector's OCR Fallback in Practice

### Rust API: Checking OCR Decisions Programmatically

```rust
use pdf_inspector::process_pdf_with_options;
use pdf_inspector::options::ProcessOptions;

let opts = ProcessOptions {
    force_ocr_all: false,  // Let the library decide based on encoding issues
    ..Default::default()
};

let result = process_pdf_with_options("sample.pdf", opts);

for (i, page) in result.pages.iter().enumerate() {
    if page.needs_ocr {
        println!("Page {} triggered OCR fallback", i + 1);
        // Inspect the specific reason
        if let Some(reason) = &page.ocr_reason {
            println!("  Reason: {:?}", reason);
        }
    }
}

```

### CLI: Automatic OCR with JSON Reporting

The `pdf2md` binary (source: [`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs)) handles OCR decisions automatically:

```bash

# Process with automatic OCR fallback

$ pdf2md sample.pdf --json > output.json

# Review which pages required OCR

$ jq '.pages[] | select(.needs_ocr) | {page: .page_number, reason: .ocr_reason}' output.json

```

Each page in the JSON output includes a `needs_ocr` boolean and `ocr_reason` field when applicable.

### Python API: Accessing Encoding Issue Flags

```python
from pdf_inspector import PdfInspector

inspector = PdfInspector()
result = inspector.process("sample.pdf")

for i, page in enumerate(result.pages, 1):
    if page.needs_ocr:
        print(f"Page {i} required OCR due to: {page.ocr_reason}")
    elif page.has_encoding_issues:
        # Encoding issues were detected but possibly resolved

        print(f"Page {i} had encoding warnings")

```

The Python bindings (defined in [`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs)) expose `has_encoding_issues` and `needs_ocr` attributes on `PdfResult` objects.

## Forcing OCR When Automatic Detection Is Insufficient

You can bypass the automatic detection entirely through the `force_ocr_all` flag:

```rust
let opts = ProcessOptions {
    force_ocr_all: true,  // OCR every page regardless of quality
    ..Default::default()
};

```

This is useful when:
- You know the PDF contains complex layouts that fool the heuristics
- You need consistent processing regardless of per-page quality
- The automatic detection misses subtle encoding issues in your document type

## Key Source Files for Understanding Encoding Detection

| File | Responsibility |
|------|--------------|
| [`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs) | Core heuristics: `detect_encoding_issues`, `has_dollar_as_space_pattern`, `CipherGarbleStats` |
| [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) | Orchestration: builds OCR queue based on quality signals (lines 820-830) |
| [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) | PDF type classification; coordinates with text-quality module |
| [`src/extractor/fonts.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/fonts.rs) | Font encoding extraction; CID-only font detection |
| [`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) | CLI interface; reports OCR decisions in JSON |
| [`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs) | Python API bindings for encoding issue flags |

## Summary

- **Expect encoding issues** when PDFs lack proper *ToUnicode* CMaps, use legacy encodings, contain malformed CMaps, or rely on CID-only fonts.
- **Automatic OCR fallback** triggers on: replacement characters (U+FFFD), dollar-as-space patterns, substitution-cipher statistics, CID-only fonts, or low overall text quality.
- **Detection is fast** because `pdf-inspector` runs heuristics before invoking expensive OCR operations.
- **Override is available** through `force_ocr_all` when you need guaranteed OCR coverage.
- **Results are inspectable** via `needs_ocr` and `ocr_reason` fields in all API bindings.

## Frequently Asked Questions

### How accurate is pdf-inspector's automatic OCR detection?

The heuristics are tuned for common PDF failure modes based on empirical analysis of production documents. According to the `pdf-inspector` source code, the substitution-cipher detection uses cosine similarity thresholds that minimize false positives while catching systematic garbling. For critical applications, inspect the `ocr_reason` field to verify detection rationale.

### Can I disable automatic OCR and handle encoding issues manually?

Yes. Set `force_ocr_all: false` (the default) and check `has_encoding_issues` or `needs_ocr` in the result. The raw extracted text remains available even when encoding issues are detected—you decide whether to use it, run custom processing, or invoke OCR separately.

### Why does pdf-inspector sometimes miss subtle encoding problems?

The detection prioritizes speed and avoids false positives. The `CipherGarbleStats` implementation requires minimum sample sizes before evaluating English similarity, and single-character substitutions may evade detection. For documents where any encoding error is unacceptable, use `force_ocr_all` or implement additional validation.

### What OCR engine does pdf-inspector use?

The library itself does not embed an OCR engine. When pages are flagged for OCR, the binary invokes a downstream processor (typically Tesseract through an external call). The OCR step occurs after the `pdf-inspector` quality analysis completes, as orchestrated in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs).