How pdf-inspector Detects PDF Encoding Issues: 3 Critical Heuristics Explained

pdf-inspector detects PDF encoding issues by scanning extracted Markdown text for Unicode replacement characters, dollar-sign substitution patterns, and substitution-cipher garbling using statistical cosine-similarity analysis, automatically flagging pages for OCR processing when the detect_encoding_issues function in src/text_quality.rs identifies corruption.

The firecrawl/pdf-inspector library validates text extraction quality by identifying encoding issues that occur when PDF font-to-Unicode mappings fail. It analyzes converted Markdown strings in src/markdown/convert.rs to catch corrupted character encodings before they propagate to downstream applications, ensuring reliable text extraction even from damaged or poorly encoded PDF documents.

The Three Core Encoding Issue Heuristics

The detect_encoding_issues function in src/text_quality.rs (lines 31‑55) evaluates the final Markdown output against three specific heuristics. If any test returns true, the page is marked as having encoding issues and routed to OCR downstream.

Unicode Replacement Characters (U+FFFD)

The detector scans for the Unicode replacement character \u{FFFD} () in the extracted text. According to the source code in src/text_quality.rs, this character appears when the PDF-to-Unicode mapping completely fails, indicating the extractor could not resolve specific byte sequences to valid Unicode characters. This signals a fundamental encoding breakdown where the font's ToUnicode CMap lacks valid entries for certain glyphs.

Dollar-Sign Space Substitution Patterns

This heuristic identifies when the dollar symbol $ appears between letters in patterns like Word$Word$Word. The has_dollar_as_space_pattern function counts these occurrences and flags the page when the count exceeds 20 instances or when they constitute more than 50% of all dollar symbols in the text. This pattern emerges when corrupted ToUnicode CMaps incorrectly map regular letters or spaces to $, causing the extractor to use $ as a surrogate for missing whitespace or alphabet characters.

Substitution Cipher Detection with CipherGarbleStats

The CipherGarbleStats analyzer detects substitution-cipher style garbling through statistical analysis of ASCII letter frequencies. It builds a histogram of observed letters and calculates two cosine-similarity scores:

  • english_cosine < 0.60 — indicates low similarity to normal English letter distribution
  • english_shape_cosine ≥ 0.90 — indicates the distribution shape still resembles English

This two-stage test identifies cases where broken CMaps apply constant offsets to characters (for example, transforming "Certificate" into "8VceZWZTReV"), creating perfect alphabet permutations that mimic substitution ciphers while preserving English letter frequency shapes.

Span-Level Encoding Validation

While detect_encoding_issues operates on full Markdown pages, additional helpers in src/text_quality.rs catch encoding problems at the individual text span level through text_span_decoding_issue_kind:

  • has_private_use_text_run — detects long runs of Private-Use Area characters, typically caused by CID-to-Unicode mapping failures in src/tounicode.rs
  • is_cid_garbage — identifies C1-control characters (U+0080–U+009F) or high-Latin-1 characters resulting from misinterpreting CID values as Latin-1
  • has_cid_control_token — flags tokens containing unusually high proportions of C1 control bytes

These functions determine if individual TextItem spans are "strongly" garbled, complementing the page-wide heuristics to ensure comprehensive detection of PDF encoding issues.

Implementation and Usage Examples

To check for encoding issues programmatically, use the process_pdf_with_options function exposed in src/lib.rs:

use pdf_inspector::process_pdf_with_options;
use pdf_inspector::options::ProcessOptions;

/// Run the extractor and print whether any page had encoding problems.
fn main() {
    let opts = ProcessOptions::default(); // default opts enable quality checks
    let result = process_pdf_with_options("sample.pdf", opts).unwrap();

    // `has_encoding_issues` is set if any of the heuristics above triggered.
    if result.has_encoding_issues {
        eprintln!("⚠️  Encoding issues detected – OCR may be required.");
    } else {
        println!("✅  No encoding problems found.");
    }
}

For command-line validation, the bundled binary exposes encoding detection via JSON output:

pdf2md sample.pdf --json | jq '.has_encoding_issues'

The boolean has_encoding_issues field directly reflects the outcome of detect_encoding_issues, while src/tounicode.rs handles the low-level ToUnicode CMap processing that can generate the broken encodings these heuristics catch.

Summary

  • pdf-inspector detects encoding issues in src/text_quality.rs using three heuristics: Unicode replacement characters, dollar-as-space patterns, and substitution-cipher statistical analysis via CipherGarbleStats
  • The U+FFFD check catches complete mapping failures, while the CipherGarbleStats cosine-similarity test identifies character offset permutations that preserve English distribution shapes
  • Span-level functions including has_private_use_text_run and is_cid_garbage detect CID-to-Unicode corruption at the individual text item level
  • When any heuristic triggers, the library automatically sets has_encoding_issues and routes affected pages to OCR fallback processing to maintain extraction quality

Frequently Asked Questions

What triggers the encoding issue flag in pdf-inspector?

The detect_encoding_issues function returns true when any of three conditions are met: the text contains Unicode replacement characters (U+FFFD), exhibits dollar-sign substitution patterns exceeding threshold counts (more than 20 occurrences or 50% of $ symbols), or fails the CipherGarbleStats statistical similarity test indicating substitution-cipher style corruption with english_cosine below 0.60 and english_shape_cosine above 0.90.

How does the dollar-as-space heuristic work?

The has_dollar_as_space_pattern function scans for $ symbols appearing between letters in patterns like Word$Word. According to the implementation in src/text_quality.rs, it flags the page when such patterns exceed 20 occurrences or represent more than 50% of all dollar symbols, indicating corrupted ToUnicode CMaps that mistakenly map spaces or letters to the dollar character.

What is CipherGarbleStats and how does it detect garbled text?

CipherGarbleStats is a statistical analyzer that compares letter frequency histograms against English language distributions using cosine similarity. It detects substitution-cipher garbling when english_cosine falls below 0.60 while english_shape_cosine remains above 0.90, indicating a constant character offset has been applied to the text—such as when "Certificate" becomes "8VceZWZTReV" due to broken font encoding mappings.

Does pdf-inspector automatically fix encoding issues?

No, pdf-inspector does not repair corrupted encodings. Instead, it detects encoding issues in src/text_quality.rs and sets the has_encoding_issues boolean flag on the result object, routing affected pages to OCR processing or alerting downstream systems that the extracted Markdown requires fallback handling according to the Firecrawl PDF extraction pipeline.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →