How to Troubleshoot Errors with Firecrawl pdf-inspector: A Complete Debugging Guide

Enable structured logging with RUST_LOG environment variables and trace errors to their source module—detector, extractor, layout, or markdown—to quickly diagnose PDF parsing failures.

Firecrawl pdf-inspector is a Rust-based PDF processing library (with Python, Node.js, and WebAssembly bindings) that classifies documents, extracts positioned text, and converts content to clean Markdown. Because it operates at a low level—parsing PDF objects, fonts, content-stream operators, and layout structures—errors can originate from multiple pipeline stages. This guide walks you through systematic troubleshooting using the actual source code from the firecrawl/pdf-inspector repository.

Understanding the Processing Pipeline

Before troubleshooting, you need to know which stage failed. The pdf-inspector pipeline flows through distinct modules, each with dedicated source files:


PDF bytes
   ├─► detector            → PdfType (TextBased / Scanned / ImageBased / Mixed)
   └─► extractor
        ├─ fonts            → font widths, encodings
        ├─ content_stream   → walk PDF operators → TextItems + PdfRects
        ├─ xobjects          → Form XObject text, image placeholders
        ├─ links             → hyperlinks, AcroForm fields
        └─ layout            → column detection → line grouping → reading order
              │
              ├─► tables
              │     ├─ detect_rects      → rectangle‑based tables (union‑find)
              │     ├─ detect_heuristic  → alignment‑based tables
              │     ├─ grid              → column/row assignment → cells
              │     └─ format            → cells → Markdown table
              │
              └─► markdown
                    ├─ analysis     → font stats, heading tiers
                    ├─ preprocess   → merge headings, drop caps
                    ├─ convert      → line loop + table/image insertion
                    ├─ classify     → captions, lists, code
                    └─ postprocess  → cleanup → final Markdown

Key source files to bookmark:

Component Source File Purpose
Public API [src/lib.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) Entry point for all language bindings
PDF type detection [src/detector.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) Fast classification of document type
Extraction core [src/extractor/mod.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) Orchestrates text and layout extraction
Font & encoding [src/tounicode.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs) Decodes CID fonts and CMaps
Layout engine [src/extractor/layout.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs) Column detection and reading order
Table detection [src/tables/detect_rects.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs) Rectangle-based table identification
Markdown output [src/markdown/convert.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) Final Markdown generation

Enabling Structured Debug Logging

pdf-inspector uses the tracing crate for structured logging. Set the RUST_LOG environment variable to target specific modules when you troubleshoot errors with firecrawl pdf-inspector.

Basic Logging Examples


# Show content-stream operator walking (verbose)

RUST_LOG=pdf_inspector::extractor::content_stream=trace cargo run --bin pdf2md -- file.pdf > /dev/null

# Debug font encoding issues

RUST_LOG=pdf_inspector::tounicode=debug cargo run --bin pdf2md -- file.pdf

# Full debug output for all modules

RUST_LOG=pdf_inspector=debug cargo run --bin pdf2md -- file.pdf

Module-Specific Log Targets

Use Case RUST_LOG Value What You'll See
Font metadata, CID encoding, ligatures pdf_inspector::extractor::fonts=debug Font name, encoding tables, fallback paths
ToUnicode CMap parsing pdf_inspector::tounicode=debug CMap parsing success/failure, character mappings
Column detection & reading order pdf_inspector::extractor::layout=debug Histogram values, valley detection, column splits
Font-size stats, heading tiers pdf_inspector::markdown::analysis=debug Font clusters, calculated heading levels
Table detection pdf_inspector::tables=debug Detected rectangles, heuristic scores, grid assignments

The complete debug target reference is documented in [docs/debugging.md](https://github.com/firecrawl/pdf-inspector/blob/main/docs/debugging.md).

Common Error Patterns and Diagnostic Steps

ParseError: InvalidFileHeader

Symptom: Immediate failure with InvalidFileHeader message.

Likely cause: Corrupt PDF header or unsupported PDF version.

Diagnostic command:

RUST_LOG=pdf_inspector::detector=debug cargo run --bin pdf2md -- corrupted.pdf

What to check: Verify the file starts with %PDF- signature. The detector logs the header bytes it found.

Missing or Garbled Text

Symptom: Output contains blank spaces or mojibake instead of expected characters.

Likely cause: CID font encoding without proper ToUnicode CMap.

Diagnostic commands:


# Check font encoding decisions

RUST_LOG=pdf_inspector::extractor::fonts=debug,pdf_inspector::tounicode=debug cargo run --bin pdf2md -- document.pdf

What to look for: Messages about "fallback to UTF-16BE" or "fallback to UTF-8" indicate missing CMaps. In [src/tounicode.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs), the parse_cmap function handles these mappings—missing CMaps trigger encoding errors logged here.

Incorrect Column Order

Symptom: Multi-column text appears scrambled or in wrong reading sequence.

Likely cause: Layout detection misidentified column structure.

Diagnostic command:

RUST_LOG=pdf_inspector::extractor::layout=debug cargo run --bin pdf2md -- newspaper.pdf

What to check: The layout module in [src/extractor/layout.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs) logs histogram values and valley detection thresholds. Look for early_exit strategy decisions that may short-circuit column detection.

Tables Not Detected

Symptom: Tabular data appears as plain paragraphs instead of Markdown tables.

Likely cause: Rectangle-based detection failed (no drawing operators) or heuristic thresholds too strict.

Diagnostic command:

RUST_LOG=pdf_inspector::tables=debug cargo run --bin pdf2md -- report.pdf

What to inspect: The rectangle detector in [src/tables/detect_rects.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs) uses union-find clustering. If you see "no rects found," the PDF lacks explicit cell borders. Check heuristic detection fallback in [src/tables/detect_heuristic.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs).

Tuning option: Adjust TABLE_MAX_COLUMNS (default 25) for unusually wide tables.

Unexpected Headings or List Markers

Symptom: Wrong heading levels, missed bullet points, or false list items.

Likely cause: Font-size clustering failed or bullet character regex didn't match.

Diagnostic command:

RUST_LOG=pdf_inspector::markdown::analysis=debug,pdf_inspector::markdown::classify=debug cargo run --bin pdf2md -- document.pdf

What to verify: Font-size ratios in clustering output, and list-prefix regex matches for characters like •, -, 1..

Crashes and Panics

Symptom: Process terminates unexpectedly without clean error.

Likely cause: Unhandled PDF construct, such as malformed XObject.

Diagnostic command:

RUST_LOG=pdf_inspector=trace cargo run --bin pdf2md -- problematic.pdf 2>&1 | tee crash.log

What to find: Stack trace with offending PDF object ID, usually logged near the panic location.

Handling Errors in Each Language Binding

Rust API

All public functions return Result<T, PdfError>. Match on error variants for targeted handling:

use pdf_inspector::{process_pdf, PdfError};

match process_pdf("file.pdf") {
    Ok(result) => println!("Pages: {}", result.pages.len()),
    Err(PdfError::ParseError(msg)) => {
        eprintln!("PDF parse error: {msg}");
        // Suggest re-exporting from source application
    }
    Err(PdfError::EncodingError(msg)) => {
        eprintln!("Font encoding error: {msg}");
        // Enable tounicode debug logging for diagnosis
    }
    Err(e) => eprintln!("Unexpected pdf-inspector error: {e}"),
}

Python Bindings

Errors surface as PdfError exceptions:

import pdf_inspector
import logging
import os

# Enable Rust logging via environment before import

os.environ["RUST_LOG"] = "pdf_inspector::tounicode=debug"

logging.basicConfig(level=logging.INFO)

try:
    result = pdf_inspector.process_pdf("document.pdf")
    print(result.markdown)
except pdf_inspector.PdfError as err:
    logging.error("pdf-inspector processing failed: %s", err)
    # Log error and re-run with more verbose target

Node.js Bindings

Errors throw as standard JavaScript Error objects:

import { readFileSync } from 'fs';
import { processPdf } from '@firecrawl/pdf-inspector';

const pdfBuffer = readFileSync('document.pdf');

try {
  const result = processPdf(pdfBuffer);
  console.log(result.markdown);
} catch (e) {
  console.error('pdf-inspector error:', e.message);
  // Inspect e.stack for Rust backtrace when compiled with debug symbols
}

Enable Rust logging from shell before Node execution:

RUST_LOG=pdf_inspector::extractor::layout=debug node process-pdf.js

WebAssembly Limitations

The WASM build does not directly expose RUST_LOG. To troubleshoot errors with firecrawl pdf-inspector in browser environments:

  1. Reproduce the issue with the CLI using equivalent RUST_LOG settings
  2. Inspect the debug output to identify the failing module
  3. Apply fixes or workarounds in your WASM integration code

Practical Debugging Workflows

Workflow 1: Diagnose a Missing Table


# Step 1: Check if rectangles were detected

RUST_LOG=pdf_inspector::tables::detect_rects=debug cargo run --bin pdf2md -- invoice.pdf

# Step 2: If no rects, check heuristic fallback

RUST_LOG=pdf_inspector::tables::detect_heuristic=debug cargo run --bin pdf2md -- invoice.pdf

# Step 3: Examine the source files to understand thresholds

cat src/tables/detect_rects.rs | grep -A5 "union_find"
cat src/tables/detect_heuristic.rs | grep -A5 "threshold"

Workflow 2: Fix Garbled Japanese Text


# Isolate font encoding decisions

RUST_LOG=pdf_inspector::extractor::fonts=debug,pdf_inspector::tounicode=debug \
  cargo run --bin pdf2md -- japanese.pdf 2>&1 | grep -E "(CMap|encoding|fallback)"

# Check if specific font is missing ToUnicode data

grep -r "ToUnicode" src/tounicode.rs

Workflow 3: Debug Layout Reading Order


# Capture full layout analysis

RUST_LOG=pdf_inspector::extractor::layout=trace cargo run --bin pdf2md -- magazine.pdf > layout.log 2>&1

# Search for column decisions

grep -E "(column|valley|histogram|early_exit)" layout.log

Summary

To effectively troubleshoot errors with firecrawl pdf-inspector:

  • Map errors to pipeline stages using the module architecture: detector → extractor (fonts, content_stream, layout) → tables → markdown
  • Use RUST_LOG environment variables to enable targeted debug output for specific modules
  • Cross-reference log messages with source files—each major component has a dedicated file in src/
  • Handle PdfError variants explicitly in your language binding to provide user-friendly failure messages
  • Validate PDF input independently when parse errors occur—corrupted files fail early in the detector

Frequently Asked Questions

How do I enable debug logging in pdf-inspector?

Set the RUST_LOG environment variable to target specific modules using the tracing crate syntax. For example: RUST_LOG=pdf_inspector::extractor::layout=debug for column detection issues, or pdf_inspector=debug for all modules. Run this before invoking the CLI, Python, or Node.js bindings.

Why is my PDF text garbled or missing characters?

This typically indicates a ToUnicode CMap parsing failure in [src/tounicode.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs). Enable pdf_inspector::tounicode=debug logging to see fallback encoding decisions. CID fonts without embedded CMaps force UTF-16BE or UTF-8 fallback, which often produces incorrect character mappings.

Why aren't my tables being detected?

pdf-inspector tries rectangle-based detection first in [src/tables/detect_rects.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs), which requires PDF drawing operators defining cell borders. If no rectangles are found, it falls back to heuristic detection in [src/tables/detect_heuristic.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs). Check pdf_inspector::tables=debug logs to see which path was taken and why detection failed.

What's the difference between ParseError and EncodingError?

ParseError indicates structural PDF problems—invalid headers, malformed objects, or unexpected end-of-file. These originate in the detector and early extraction phases. EncodingError signals font-specific issues, particularly missing or corrupt ToUnicode CMaps that prevent proper character decoding. Match on these variants in Rust to implement appropriate recovery strategies.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →