# How pdf-inspector Detects PDF Encoding Issues: 3 Critical Heuristics Explained

> Discover how pdf-inspector detects PDF encoding issues like Unicode replacement, dollar signs, and cipher garbling using 3 critical heuristics. Ensure accurate text extraction.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: deep-dive
- Published: 2026-08-08

---

**pdf-inspector detects PDF encoding issues by scanning extracted Markdown text for Unicode replacement characters, dollar-sign substitution patterns, and substitution-cipher garbling using statistical cosine-similarity analysis, automatically flagging pages for OCR processing when the `detect_encoding_issues` function in [`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs) identifies corruption.**

The `firecrawl/pdf-inspector` library validates text extraction quality by identifying encoding issues that occur when PDF font-to-Unicode mappings fail. It analyzes converted Markdown strings in [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) to catch corrupted character encodings before they propagate to downstream applications, ensuring reliable text extraction even from damaged or poorly encoded PDF documents.

## The Three Core Encoding Issue Heuristics

The `detect_encoding_issues` function in [`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs) (lines 31‑55) evaluates the final Markdown output against three specific heuristics. If any test returns `true`, the page is marked as having encoding issues and routed to OCR downstream.

### Unicode Replacement Characters (U+FFFD)

The detector scans for the Unicode replacement character `\u{FFFD}` () in the extracted text. According to the source code in [`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs), this character appears when the PDF-to-Unicode mapping completely fails, indicating the extractor could not resolve specific byte sequences to valid Unicode characters. This signals a fundamental encoding breakdown where the font's ToUnicode CMap lacks valid entries for certain glyphs.

### Dollar-Sign Space Substitution Patterns

This heuristic identifies when the dollar symbol `$` appears between letters in patterns like `Word$Word$Word`. The `has_dollar_as_space_pattern` function counts these occurrences and flags the page when the count exceeds **20 instances** or when they constitute **more than 50%** of all dollar symbols in the text. This pattern emerges when corrupted ToUnicode CMaps incorrectly map regular letters or spaces to `$`, causing the extractor to use `$` as a surrogate for missing whitespace or alphabet characters.

### Substitution Cipher Detection with CipherGarbleStats

The `CipherGarbleStats` analyzer detects substitution-cipher style garbling through statistical analysis of ASCII letter frequencies. It builds a histogram of observed letters and calculates two cosine-similarity scores:

- **`english_cosine` < 0.60** — indicates low similarity to normal English letter distribution
- **`english_shape_cosine` ≥ 0.90** — indicates the distribution shape still resembles English

This two-stage test identifies cases where broken CMaps apply constant offsets to characters (for example, transforming "Certificate" into "8VceZWZTReV"), creating perfect alphabet permutations that mimic substitution ciphers while preserving English letter frequency shapes.

## Span-Level Encoding Validation

While `detect_encoding_issues` operates on full Markdown pages, additional helpers in [`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs) catch encoding problems at the individual text span level through `text_span_decoding_issue_kind`:

- **`has_private_use_text_run`** — detects long runs of Private-Use Area characters, typically caused by CID-to-Unicode mapping failures in [`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs)
- **`is_cid_garbage`** — identifies C1-control characters (U+0080–U+009F) or high-Latin-1 characters resulting from misinterpreting CID values as Latin-1
- **`has_cid_control_token`** — flags tokens containing unusually high proportions of C1 control bytes

These functions determine if individual `TextItem` spans are "strongly" garbled, complementing the page-wide heuristics to ensure comprehensive detection of PDF encoding issues.

## Implementation and Usage Examples

To check for encoding issues programmatically, use the `process_pdf_with_options` function exposed in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs):

```rust
use pdf_inspector::process_pdf_with_options;
use pdf_inspector::options::ProcessOptions;

/// Run the extractor and print whether any page had encoding problems.
fn main() {
    let opts = ProcessOptions::default(); // default opts enable quality checks
    let result = process_pdf_with_options("sample.pdf", opts).unwrap();

    // `has_encoding_issues` is set if any of the heuristics above triggered.
    if result.has_encoding_issues {
        eprintln!("⚠️  Encoding issues detected – OCR may be required.");
    } else {
        println!("✅  No encoding problems found.");
    }
}

```

For command-line validation, the bundled binary exposes encoding detection via JSON output:

```bash
pdf2md sample.pdf --json | jq '.has_encoding_issues'

```

The boolean `has_encoding_issues` field directly reflects the outcome of `detect_encoding_issues`, while [`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs) handles the low-level ToUnicode CMap processing that can generate the broken encodings these heuristics catch.

## Summary

- pdf-inspector detects encoding issues in [`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs) using three heuristics: Unicode replacement characters, dollar-as-space patterns, and substitution-cipher statistical analysis via `CipherGarbleStats`
- The **U+FFFD** check catches complete mapping failures, while the **CipherGarbleStats** cosine-similarity test identifies character offset permutations that preserve English distribution shapes
- Span-level functions including `has_private_use_text_run` and `is_cid_garbage` detect CID-to-Unicode corruption at the individual text item level
- When any heuristic triggers, the library automatically sets `has_encoding_issues` and routes affected pages to OCR fallback processing to maintain extraction quality

## Frequently Asked Questions

### What triggers the encoding issue flag in pdf-inspector?

The `detect_encoding_issues` function returns `true` when any of three conditions are met: the text contains Unicode replacement characters (U+FFFD), exhibits dollar-sign substitution patterns exceeding threshold counts (more than 20 occurrences or 50% of `$` symbols), or fails the `CipherGarbleStats` statistical similarity test indicating substitution-cipher style corruption with `english_cosine` below 0.60 and `english_shape_cosine` above 0.90.

### How does the dollar-as-space heuristic work?

The `has_dollar_as_space_pattern` function scans for `$` symbols appearing between letters in patterns like `Word$Word`. According to the implementation in [`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs), it flags the page when such patterns exceed 20 occurrences or represent more than 50% of all dollar symbols, indicating corrupted ToUnicode CMaps that mistakenly map spaces or letters to the dollar character.

### What is CipherGarbleStats and how does it detect garbled text?

`CipherGarbleStats` is a statistical analyzer that compares letter frequency histograms against English language distributions using cosine similarity. It detects substitution-cipher garbling when `english_cosine` falls below 0.60 while `english_shape_cosine` remains above 0.90, indicating a constant character offset has been applied to the text—such as when "Certificate" becomes "8VceZWZTReV" due to broken font encoding mappings.

### Does pdf-inspector automatically fix encoding issues?

No, pdf-inspector does not repair corrupted encodings. Instead, it detects encoding issues in [`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs) and sets the `has_encoding_issues` boolean flag on the result object, routing affected pages to OCR processing or alerting downstream systems that the extracted Markdown requires fallback handling according to the Firecrawl PDF extraction pipeline.