How to Troubleshoot Common Issues with pdf-inspector: Complete Diagnostics Guide
pdf-inspector exposes diagnostic flags in PdfProcessResult—including has_encoding_issues, pages_needing_ocr, and ocr_reasons_by_page—that pinpoint exactly why text extraction fails or produces garbled output.
pdf-inspector is a Rust library that detects PDF types, extracts text, and converts documents to Markdown while handling tables, multi-column layouts, and OCR fallbacks. When you encounter garbage text, missing tables, or unexpected OCR triggers, systematic troubleshooting requires inspecting specific fields in the processing result. This guide maps common symptoms to their root causes in the source code and provides concrete verification steps.
Diagnosing Text Encoding and Garbage Output
Garbage text appears as replacement characters (U+FFFD) or "dollar-as-space" artifacts when fonts lack proper ToUnicode CMaps. The library detects this in detect_encoding_issues and region_items_have_decoding_issue within src/text_quality.rs.
When encoding issues are detected, extract_pages_markdown_mem (lines ≈ 73‑78) and extract_text_in_regions_mem (lines ≈ 84‑89) automatically flag the page for OCR fallback. To verify:
- Run
process_pdfand checkhas_encoding_issuesin the returnedPdfProcessResult. - Inspect
PageMarkdown.markdownfor `` sequences or odd spacing. - To suppress false positives, ensure the PDF embeds fonts with valid ToUnicode CMaps or fallback fonts.
Fixing GID-Encoded Font Issues (CID / Glyph ID)
Fonts exposing only glyph IDs (GIDs) without Unicode mappings produce empty or garbled text. The detection happens in extract_pages_markdown_mem via the has_gid flag and in extract_text_in_regions_mem via the gid_pages set.
The final OCR decision logic (around line ≈ 84) evaluates:
needs_ocr = ocr_reason.is_some()
|| md.trim().is_empty()
|| has_gid
|| is_garbage_text(&md)
To troubleshoot:
- Check
pages_needing_ocrfor entries flagged withOCR_REASON_VECTOR_TEXT. - Verify font objects contain valid ToUnicode maps using
pdfinfoormutool. - If CMaps are missing, allow the OCR pipeline to run or pre-process the PDF to embed proper mappings.
Handling Encrypted PDFs
Encrypted PDFs throw InvalidFileHeader-like errors when lopdf cannot decrypt them. The library attempts decryption in load_document_from_path_with_password, called from process_pdf_with_options around line ≈ 3490 in src/lib.rs.
Resolve encryption issues by passing the password through PdfOptions:
let opts = PdfOptions::new().password("secret");
let result = pdf_inspector::process_pdf_with_options("encrypted.pdf", opts)?;
If decryption still fails, the PDF likely uses non-standard encryption that lopdf does not support. Pre-decrypt the file using qpdf or similar tools before processing.
Resolving Missing Pages and Index Errors
Requesting a page index beyond the document bounds yields empty markdown with an OCR flag. The guard logic in extract_pages_markdown_mem (lines ≈ 14‑24) pushes an empty PageMarkdown with needs_ocr: true for out-of-range requests.
Always validate indices against the page_count field from PdfProcessResult. Remember that pdf-inspector uses 0-based indexing for internal operations, though PdfOptions accepts 1-based page numbers for user convenience.
Troubleshooting Table Detection Failures
Table detection runs through three stages: rectangle-based (detect_tables_from_rects), line-based (detect_tables_from_lines), and heuristic (detect_tables_with_page_width). The selection logic resides in extract_text_in_regions_mem (lines ≈ 730‑770) and extract_tables_in_regions_mem (lines ≈ 840‑900).
If tables are missing from output:
- Inspect
ocr_reasons_by_pageforOCR_REASON_SUSPECTED_GARBLED_TEXT, which causes early abort. - Verify the region contains sufficient text items using
region_items_have_decoding_issue. - If the heuristic path triggers, adjust the region size or
base_font_size(computed around line ≈ 1045 in src/markdown/analysis.rs).
Optimizing Performance and Slow Processing
Long processing times usually stem from letter-spacing correction algorithms or OCR fallback cascades. The ProcessingTimer wrapper (lines ≈ 71‑97 in src/lib.rs) records elapsed time in processing_time_ms.
Heavy computation occurs in:
extract_page_text_items– called for every page.text_utils::fix_letterspaced_items– triggers when character spacing exceeds 0.10 thresholds.
To improve performance:
- Profile using
processing_time_msfromPdfProcessResult. - Restrict processing to specific pages using
PdfOptions::pages()to avoid OCR on unnecessary sections. - Disable OCR fallback entirely by setting
ProcessMode::Fastif text quality is acceptable.
Step-by-Step Debugging Workflow
Follow this systematic approach to isolate issues:
- Run quick detection – Call
detect_pdf("my.pdf")and inspectpdf_type,page_count, andhas_encoding_issues. - Full extraction with diagnostics – Use
process_pdf_with_optionsand examinepages_needing_ocr,ocr_reasons_by_page, andlayout. - Per-page inspection – Call
extract_pages_markdownand iterate overPagesExtractionResult.pagesto identify empty outputs. - Region-level troubleshooting – For missing tables, call
extract_text_in_regions_memwith exact bounding-box coordinates to checkRegionText.needs_ocr. - Enable debug logging – Set
RUST_LOG=pdf_inspector::extractor=debugto view granular messages about font-cmap loading and GID detection.
Complete Troubleshooting Example
This Rust example demonstrates detection, selective processing, and region-level extraction:
use pdf_inspector::{process_pdf, PdfOptions, ProcessMode};
fn main() -> Result<(), pdf_inspector::PdfError> {
// 1️⃣ Quick detection only
let detect = pdf_inspector::detect_pdf("sample.pdf")?;
println!("Detected type: {:?}, pages: {}", detect.pdf_type, detect.page_count);
// 2️⃣ Full extraction with OCR diagnostics
let result = pdf_inspector::process_pdf("sample.pdf")?;
println!("Encoding issues? {}", result.has_encoding_issues);
println!("Pages needing OCR: {:?}", result.pages_needing_ocr);
// 3️⃣ Restrict to specific pages (1-based indexing in options)
let opts = PdfOptions::new()
.mode(ProcessMode::Full)
.pages([1, 3, 5]);
let limited = pdf_inspector::process_pdf_with_options("sample.pdf", opts)?;
println!("Limited markdown length: {}", limited.markdown.unwrap_or_default().len());
// 4️⃣ Region-level extraction for suspected tables
let regions = vec![
(0, vec![[72.0, 720.0, 540.0, 600.0]]), // page 0, bbox [x1, y1, x2, y2]
];
let table_res = pdf_inspector::extract_text_in_regions_mem(
&std::fs::read("sample.pdf")?,
®ions,
)?;
for page_res in table_res {
for region in page_res.regions {
println!("Region needs OCR? {}", region.needs_ocr);
println!("Text preview: {:.100}...", region.text);
}
}
Ok(())
}
Summary
- Encoding issues trigger OCR fallback when
has_encoding_issuesreturns true; inspecttext_quality.rsfor detection logic. - GID-encoded fonts force OCR via
OCR_REASON_VECTOR_TEXTwhen ToUnicode CMaps are missing. - Encrypted PDFs require password options in
PdfOptionsor pre-decryption with external tools. - Out-of-range pages produce empty markdown; validate against
page_countusing 0-based indexing. - Table detection fails when regions contain garbled text or insufficient font metadata; check
ocr_reasons_by_pagefor abort signals. - Performance bottlenecks appear in
processing_time_ms; limit pages or disable OCR to reduce latency.
Frequently Asked Questions
Why am I seeing U+FFFD replacement characters in extracted text?
U+FFFD characters indicate that pdf-inspector detected encoding issues in text_quality.rs and the source font lacks a proper ToUnicode CMap. The library flags these pages in has_encoding_issues and may trigger OCR fallback. To fix, regenerate the PDF with embedded fonts containing valid Unicode mappings, or allow OCR to process the affected pages.
How do I handle password-protected PDFs in pdf-inspector?
Pass the decryption password through PdfOptions::new().password("your_password") when calling process_pdf_with_options. If the PDF uses non-standard encryption that the underlying lopdf library cannot handle, pre-decrypt the file using qpdf --password=secret --decrypt input.pdf output.pdf before processing.
Why are tables not detected even though they are visible in the PDF?
Table detection aborts early if OCR_REASON_SUSPECTED_GARBLED_TEXT appears in ocr_reasons_by_page, or if the region fails the region_items_have_decoding_issue check in text_quality.rs. Ensure the table text is not flagged as garbage, and verify that base_font_size calculations in analysis.rs correctly identify column boundaries. For debugging, use extract_text_in_regions_mem with explicit bounding boxes to bypass automatic region detection.
How can I disable OCR fallback to improve processing speed?
Set ProcessMode::Fast in your PdfOptions to skip OCR pipelines entirely, or restrict processing to specific pages using .pages([1, 2, 3]) to avoid OCR overhead on scanned sections. Monitor processing_time_ms in the results to identify which pages trigger expensive fix_letterspaced_items operations or OCR cascades.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →