# How to Troubleshoot Common Issues with pdf-inspector: Complete Diagnostics Guide

> Troubleshoot pdf-inspector issues with this diagnostics guide. Learn to identify text extraction failures and garbled output using PdfProcessResult flags like has_encoding_issues and pages_needing_ocr.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-04

---

**pdf-inspector exposes diagnostic flags in `PdfProcessResult`—including `has_encoding_issues`, `pages_needing_ocr`, and `ocr_reasons_by_page`—that pinpoint exactly why text extraction fails or produces garbled output.**

pdf-inspector is a Rust library that detects PDF types, extracts text, and converts documents to Markdown while handling tables, multi-column layouts, and OCR fallbacks. When you encounter garbage text, missing tables, or unexpected OCR triggers, systematic troubleshooting requires inspecting specific fields in the processing result. This guide maps common symptoms to their root causes in the source code and provides concrete verification steps.

## Diagnosing Text Encoding and Garbage Output

Garbage text appears as replacement characters (U+FFFD) or "dollar-as-space" artifacts when fonts lack proper ToUnicode CMaps. The library detects this in `detect_encoding_issues` and `region_items_have_decoding_issue` within **[src/text_quality.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs)**.

When encoding issues are detected, `extract_pages_markdown_mem` (lines ≈ 73‑78) and `extract_text_in_regions_mem` (lines ≈ 84‑89) automatically flag the page for OCR fallback. To verify:

1. Run `process_pdf` and check `has_encoding_issues` in the returned `PdfProcessResult`.
2. Inspect `PageMarkdown.markdown` for `` sequences or odd spacing.
3. To suppress false positives, ensure the PDF embeds fonts with valid ToUnicode CMaps or fallback fonts.

## Fixing GID-Encoded Font Issues (CID / Glyph ID)

Fonts exposing only glyph IDs (GIDs) without Unicode mappings produce empty or garbled text. The detection happens in `extract_pages_markdown_mem` via the `has_gid` flag and in `extract_text_in_regions_mem` via the `gid_pages` set.

The final OCR decision logic (around line ≈ 84) evaluates:

```rust
needs_ocr = ocr_reason.is_some() 
    || md.trim().is_empty() 
    || has_gid 
    || is_garbage_text(&md)

```

To troubleshoot:

- Check `pages_needing_ocr` for entries flagged with `OCR_REASON_VECTOR_TEXT`.
- Verify font objects contain valid ToUnicode maps using `pdfinfo` or `mutool`.
- If CMaps are missing, allow the OCR pipeline to run or pre-process the PDF to embed proper mappings.

## Handling Encrypted PDFs

Encrypted PDFs throw `InvalidFileHeader`-like errors when `lopdf` cannot decrypt them. The library attempts decryption in `load_document_from_path_with_password`, called from `process_pdf_with_options` around line ≈ 3490 in **[src/lib.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)**.

Resolve encryption issues by passing the password through `PdfOptions`:

```rust
let opts = PdfOptions::new().password("secret");
let result = pdf_inspector::process_pdf_with_options("encrypted.pdf", opts)?;

```

If decryption still fails, the PDF likely uses non-standard encryption that `lopdf` does not support. Pre-decrypt the file using `qpdf` or similar tools before processing.

## Resolving Missing Pages and Index Errors

Requesting a page index beyond the document bounds yields empty markdown with an OCR flag. The guard logic in `extract_pages_markdown_mem` (lines ≈ 14‑24) pushes an empty `PageMarkdown` with `needs_ocr: true` for out-of-range requests.

Always validate indices against the `page_count` field from `PdfProcessResult`. Remember that pdf-inspector uses **0-based indexing** for internal operations, though `PdfOptions` accepts 1-based page numbers for user convenience.

## Troubleshooting Table Detection Failures

Table detection runs through three stages: rectangle-based (`detect_tables_from_rects`), line-based (`detect_tables_from_lines`), and heuristic (`detect_tables_with_page_width`). The selection logic resides in `extract_text_in_regions_mem` (lines ≈ 730‑770) and `extract_tables_in_regions_mem` (lines ≈ 840‑900).

If tables are missing from output:

1. Inspect `ocr_reasons_by_page` for `OCR_REASON_SUSPECTED_GARBLED_TEXT`, which causes early abort.
2. Verify the region contains sufficient text items using `region_items_have_decoding_issue`.
3. If the heuristic path triggers, adjust the region size or `base_font_size` (computed around line ≈ 1045 in **[src/markdown/analysis.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/analysis.rs)**).

## Optimizing Performance and Slow Processing

Long processing times usually stem from letter-spacing correction algorithms or OCR fallback cascades. The `ProcessingTimer` wrapper (lines ≈ 71‑97 in **[src/lib.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)**) records elapsed time in `processing_time_ms`.

Heavy computation occurs in:

- `extract_page_text_items` – called for every page.
- `text_utils::fix_letterspaced_items` – triggers when character spacing exceeds 0.10 thresholds.

To improve performance:

- Profile using `processing_time_ms` from `PdfProcessResult`.
- Restrict processing to specific pages using `PdfOptions::pages()` to avoid OCR on unnecessary sections.
- Disable OCR fallback entirely by setting `ProcessMode::Fast` if text quality is acceptable.

## Step-by-Step Debugging Workflow

Follow this systematic approach to isolate issues:

1. **Run quick detection** – Call `detect_pdf("my.pdf")` and inspect `pdf_type`, `page_count`, and `has_encoding_issues`.
2. **Full extraction with diagnostics** – Use `process_pdf_with_options` and examine `pages_needing_ocr`, `ocr_reasons_by_page`, and `layout`.
3. **Per-page inspection** – Call `extract_pages_markdown` and iterate over `PagesExtractionResult.pages` to identify empty outputs.
4. **Region-level troubleshooting** – For missing tables, call `extract_text_in_regions_mem` with exact bounding-box coordinates to check `RegionText.needs_ocr`.
5. **Enable debug logging** – Set `RUST_LOG=pdf_inspector::extractor=debug` to view granular messages about font-cmap loading and GID detection.

## Complete Troubleshooting Example

This Rust example demonstrates detection, selective processing, and region-level extraction:

```rust
use pdf_inspector::{process_pdf, PdfOptions, ProcessMode};

fn main() -> Result<(), pdf_inspector::PdfError> {
    // 1️⃣ Quick detection only
    let detect = pdf_inspector::detect_pdf("sample.pdf")?;
    println!("Detected type: {:?}, pages: {}", detect.pdf_type, detect.page_count);

    // 2️⃣ Full extraction with OCR diagnostics
    let result = pdf_inspector::process_pdf("sample.pdf")?;
    println!("Encoding issues? {}", result.has_encoding_issues);
    println!("Pages needing OCR: {:?}", result.pages_needing_ocr);

    // 3️⃣ Restrict to specific pages (1-based indexing in options)
    let opts = PdfOptions::new()
        .mode(ProcessMode::Full)
        .pages([1, 3, 5]);
    let limited = pdf_inspector::process_pdf_with_options("sample.pdf", opts)?;
    println!("Limited markdown length: {}", limited.markdown.unwrap_or_default().len());

    // 4️⃣ Region-level extraction for suspected tables
    let regions = vec![
        (0, vec![[72.0, 720.0, 540.0, 600.0]]), // page 0, bbox [x1, y1, x2, y2]
    ];
    let table_res = pdf_inspector::extract_text_in_regions_mem(
        &std::fs::read("sample.pdf")?,
        &regions,
    )?;
    for page_res in table_res {
        for region in page_res.regions {
            println!("Region needs OCR? {}", region.needs_ocr);
            println!("Text preview: {:.100}...", region.text);
        }
    }
    Ok(())
}

```

## Summary

- **Encoding issues** trigger OCR fallback when `has_encoding_issues` returns true; inspect [`text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/text_quality.rs) for detection logic.
- **GID-encoded fonts** force OCR via `OCR_REASON_VECTOR_TEXT` when ToUnicode CMaps are missing.
- **Encrypted PDFs** require password options in `PdfOptions` or pre-decryption with external tools.
- **Out-of-range pages** produce empty markdown; validate against `page_count` using 0-based indexing.
- **Table detection** fails when regions contain garbled text or insufficient font metadata; check `ocr_reasons_by_page` for abort signals.
- **Performance bottlenecks** appear in `processing_time_ms`; limit pages or disable OCR to reduce latency.

## Frequently Asked Questions

### Why am I seeing U+FFFD replacement characters in extracted text?

U+FFFD characters indicate that pdf-inspector detected encoding issues in [`text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/text_quality.rs) and the source font lacks a proper ToUnicode CMap. The library flags these pages in `has_encoding_issues` and may trigger OCR fallback. To fix, regenerate the PDF with embedded fonts containing valid Unicode mappings, or allow OCR to process the affected pages.

### How do I handle password-protected PDFs in pdf-inspector?

Pass the decryption password through `PdfOptions::new().password("your_password")` when calling `process_pdf_with_options`. If the PDF uses non-standard encryption that the underlying `lopdf` library cannot handle, pre-decrypt the file using `qpdf --password=secret --decrypt input.pdf output.pdf` before processing.

### Why are tables not detected even though they are visible in the PDF?

Table detection aborts early if `OCR_REASON_SUSPECTED_GARBLED_TEXT` appears in `ocr_reasons_by_page`, or if the region fails the `region_items_have_decoding_issue` check in [`text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/text_quality.rs). Ensure the table text is not flagged as garbage, and verify that `base_font_size` calculations in [`analysis.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/analysis.rs) correctly identify column boundaries. For debugging, use `extract_text_in_regions_mem` with explicit bounding boxes to bypass automatic region detection.

### How can I disable OCR fallback to improve processing speed?

Set `ProcessMode::Fast` in your `PdfOptions` to skip OCR pipelines entirely, or restrict processing to specific pages using `.pages([1, 2, 3])` to avoid OCR overhead on scanned sections. Monitor `processing_time_ms` in the results to identify which pages trigger expensive `fix_letterspaced_items` operations or OCR cascades.