How to Troubleshoot Common Issues with pdf-inspector: Complete Diagnostics Guide

pdf-inspector exposes diagnostic flags in PdfProcessResult—including has_encoding_issues, pages_needing_ocr, and ocr_reasons_by_page—that pinpoint exactly why text extraction fails or produces garbled output.

pdf-inspector is a Rust library that detects PDF types, extracts text, and converts documents to Markdown while handling tables, multi-column layouts, and OCR fallbacks. When you encounter garbage text, missing tables, or unexpected OCR triggers, systematic troubleshooting requires inspecting specific fields in the processing result. This guide maps common symptoms to their root causes in the source code and provides concrete verification steps.

Diagnosing Text Encoding and Garbage Output

Garbage text appears as replacement characters (U+FFFD) or "dollar-as-space" artifacts when fonts lack proper ToUnicode CMaps. The library detects this in detect_encoding_issues and region_items_have_decoding_issue within src/text_quality.rs.

When encoding issues are detected, extract_pages_markdown_mem (lines ≈ 73‑78) and extract_text_in_regions_mem (lines ≈ 84‑89) automatically flag the page for OCR fallback. To verify:

  1. Run process_pdf and check has_encoding_issues in the returned PdfProcessResult.
  2. Inspect PageMarkdown.markdown for `` sequences or odd spacing.
  3. To suppress false positives, ensure the PDF embeds fonts with valid ToUnicode CMaps or fallback fonts.

Fixing GID-Encoded Font Issues (CID / Glyph ID)

Fonts exposing only glyph IDs (GIDs) without Unicode mappings produce empty or garbled text. The detection happens in extract_pages_markdown_mem via the has_gid flag and in extract_text_in_regions_mem via the gid_pages set.

The final OCR decision logic (around line ≈ 84) evaluates:

needs_ocr = ocr_reason.is_some() 
    || md.trim().is_empty() 
    || has_gid 
    || is_garbage_text(&md)

To troubleshoot:

  • Check pages_needing_ocr for entries flagged with OCR_REASON_VECTOR_TEXT.
  • Verify font objects contain valid ToUnicode maps using pdfinfo or mutool.
  • If CMaps are missing, allow the OCR pipeline to run or pre-process the PDF to embed proper mappings.

Handling Encrypted PDFs

Encrypted PDFs throw InvalidFileHeader-like errors when lopdf cannot decrypt them. The library attempts decryption in load_document_from_path_with_password, called from process_pdf_with_options around line ≈ 3490 in src/lib.rs.

Resolve encryption issues by passing the password through PdfOptions:

let opts = PdfOptions::new().password("secret");
let result = pdf_inspector::process_pdf_with_options("encrypted.pdf", opts)?;

If decryption still fails, the PDF likely uses non-standard encryption that lopdf does not support. Pre-decrypt the file using qpdf or similar tools before processing.

Resolving Missing Pages and Index Errors

Requesting a page index beyond the document bounds yields empty markdown with an OCR flag. The guard logic in extract_pages_markdown_mem (lines ≈ 14‑24) pushes an empty PageMarkdown with needs_ocr: true for out-of-range requests.

Always validate indices against the page_count field from PdfProcessResult. Remember that pdf-inspector uses 0-based indexing for internal operations, though PdfOptions accepts 1-based page numbers for user convenience.

Troubleshooting Table Detection Failures

Table detection runs through three stages: rectangle-based (detect_tables_from_rects), line-based (detect_tables_from_lines), and heuristic (detect_tables_with_page_width). The selection logic resides in extract_text_in_regions_mem (lines ≈ 730‑770) and extract_tables_in_regions_mem (lines ≈ 840‑900).

If tables are missing from output:

  1. Inspect ocr_reasons_by_page for OCR_REASON_SUSPECTED_GARBLED_TEXT, which causes early abort.
  2. Verify the region contains sufficient text items using region_items_have_decoding_issue.
  3. If the heuristic path triggers, adjust the region size or base_font_size (computed around line ≈ 1045 in src/markdown/analysis.rs).

Optimizing Performance and Slow Processing

Long processing times usually stem from letter-spacing correction algorithms or OCR fallback cascades. The ProcessingTimer wrapper (lines ≈ 71‑97 in src/lib.rs) records elapsed time in processing_time_ms.

Heavy computation occurs in:

  • extract_page_text_items – called for every page.
  • text_utils::fix_letterspaced_items – triggers when character spacing exceeds 0.10 thresholds.

To improve performance:

  • Profile using processing_time_ms from PdfProcessResult.
  • Restrict processing to specific pages using PdfOptions::pages() to avoid OCR on unnecessary sections.
  • Disable OCR fallback entirely by setting ProcessMode::Fast if text quality is acceptable.

Step-by-Step Debugging Workflow

Follow this systematic approach to isolate issues:

  1. Run quick detection – Call detect_pdf("my.pdf") and inspect pdf_type, page_count, and has_encoding_issues.
  2. Full extraction with diagnostics – Use process_pdf_with_options and examine pages_needing_ocr, ocr_reasons_by_page, and layout.
  3. Per-page inspection – Call extract_pages_markdown and iterate over PagesExtractionResult.pages to identify empty outputs.
  4. Region-level troubleshooting – For missing tables, call extract_text_in_regions_mem with exact bounding-box coordinates to check RegionText.needs_ocr.
  5. Enable debug logging – Set RUST_LOG=pdf_inspector::extractor=debug to view granular messages about font-cmap loading and GID detection.

Complete Troubleshooting Example

This Rust example demonstrates detection, selective processing, and region-level extraction:

use pdf_inspector::{process_pdf, PdfOptions, ProcessMode};

fn main() -> Result<(), pdf_inspector::PdfError> {
    // 1️⃣ Quick detection only
    let detect = pdf_inspector::detect_pdf("sample.pdf")?;
    println!("Detected type: {:?}, pages: {}", detect.pdf_type, detect.page_count);

    // 2️⃣ Full extraction with OCR diagnostics
    let result = pdf_inspector::process_pdf("sample.pdf")?;
    println!("Encoding issues? {}", result.has_encoding_issues);
    println!("Pages needing OCR: {:?}", result.pages_needing_ocr);

    // 3️⃣ Restrict to specific pages (1-based indexing in options)
    let opts = PdfOptions::new()
        .mode(ProcessMode::Full)
        .pages([1, 3, 5]);
    let limited = pdf_inspector::process_pdf_with_options("sample.pdf", opts)?;
    println!("Limited markdown length: {}", limited.markdown.unwrap_or_default().len());

    // 4️⃣ Region-level extraction for suspected tables
    let regions = vec![
        (0, vec![[72.0, 720.0, 540.0, 600.0]]), // page 0, bbox [x1, y1, x2, y2]
    ];
    let table_res = pdf_inspector::extract_text_in_regions_mem(
        &std::fs::read("sample.pdf")?,
        &regions,
    )?;
    for page_res in table_res {
        for region in page_res.regions {
            println!("Region needs OCR? {}", region.needs_ocr);
            println!("Text preview: {:.100}...", region.text);
        }
    }
    Ok(())
}

Summary

  • Encoding issues trigger OCR fallback when has_encoding_issues returns true; inspect text_quality.rs for detection logic.
  • GID-encoded fonts force OCR via OCR_REASON_VECTOR_TEXT when ToUnicode CMaps are missing.
  • Encrypted PDFs require password options in PdfOptions or pre-decryption with external tools.
  • Out-of-range pages produce empty markdown; validate against page_count using 0-based indexing.
  • Table detection fails when regions contain garbled text or insufficient font metadata; check ocr_reasons_by_page for abort signals.
  • Performance bottlenecks appear in processing_time_ms; limit pages or disable OCR to reduce latency.

Frequently Asked Questions

Why am I seeing U+FFFD replacement characters in extracted text?

U+FFFD characters indicate that pdf-inspector detected encoding issues in text_quality.rs and the source font lacks a proper ToUnicode CMap. The library flags these pages in has_encoding_issues and may trigger OCR fallback. To fix, regenerate the PDF with embedded fonts containing valid Unicode mappings, or allow OCR to process the affected pages.

How do I handle password-protected PDFs in pdf-inspector?

Pass the decryption password through PdfOptions::new().password("your_password") when calling process_pdf_with_options. If the PDF uses non-standard encryption that the underlying lopdf library cannot handle, pre-decrypt the file using qpdf --password=secret --decrypt input.pdf output.pdf before processing.

Why are tables not detected even though they are visible in the PDF?

Table detection aborts early if OCR_REASON_SUSPECTED_GARBLED_TEXT appears in ocr_reasons_by_page, or if the region fails the region_items_have_decoding_issue check in text_quality.rs. Ensure the table text is not flagged as garbage, and verify that base_font_size calculations in analysis.rs correctly identify column boundaries. For debugging, use extract_text_in_regions_mem with explicit bounding boxes to bypass automatic region detection.

How can I disable OCR fallback to improve processing speed?

Set ProcessMode::Fast in your PdfOptions to skip OCR pipelines entirely, or restrict processing to specific pages using .pages([1, 2, 3]) to avoid OCR overhead on scanned sections. Monitor processing_time_ms in the results to identify which pages trigger expensive fix_letterspaced_items operations or OCR cascades.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →