# How to Troubleshoot Errors with Firecrawl pdf-inspector: A Complete Debugging Guide

> Troubleshoot firecrawl pdf-inspector errors by enabling structured logging and tracing issues to their source module. Quickly diagnose PDF parsing failures with this debugging guide.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-07

---

**Enable structured logging with `RUST_LOG` environment variables and trace errors to their source module—detector, extractor, layout, or markdown—to quickly diagnose PDF parsing failures.**

Firecrawl **pdf-inspector** is a Rust-based PDF processing library (with Python, Node.js, and WebAssembly bindings) that classifies documents, extracts positioned text, and converts content to clean Markdown. Because it operates at a low level—parsing PDF objects, fonts, content-stream operators, and layout structures—errors can originate from multiple pipeline stages. This guide walks you through systematic troubleshooting using the actual source code from the [firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector) repository.

## Understanding the Processing Pipeline

Before troubleshooting, you need to know which stage failed. The pdf-inspector pipeline flows through distinct modules, each with dedicated source files:

```

PDF bytes
   ├─► detector            → PdfType (TextBased / Scanned / ImageBased / Mixed)
   └─► extractor
        ├─ fonts            → font widths, encodings
        ├─ content_stream   → walk PDF operators → TextItems + PdfRects
        ├─ xobjects          → Form XObject text, image placeholders
        ├─ links             → hyperlinks, AcroForm fields
        └─ layout            → column detection → line grouping → reading order
              │
              ├─► tables
              │     ├─ detect_rects      → rectangle‑based tables (union‑find)
              │     ├─ detect_heuristic  → alignment‑based tables
              │     ├─ grid              → column/row assignment → cells
              │     └─ format            → cells → Markdown table
              │
              └─► markdown
                    ├─ analysis     → font stats, heading tiers
                    ├─ preprocess   → merge headings, drop caps
                    ├─ convert      → line loop + table/image insertion
                    ├─ classify     → captions, lists, code
                    └─ postprocess  → cleanup → final Markdown

```

Key source files to bookmark:

| Component | Source File | Purpose |
|-----------|-------------|---------|
| Public API | [[`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) | Entry point for all language bindings |
| PDF type detection | [[`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) | Fast classification of document type |
| Extraction core | [[`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) | Orchestrates text and layout extraction |
| Font & encoding | [[`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs) | Decodes CID fonts and CMaps |
| Layout engine | [[`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs) | Column detection and reading order |
| Table detection | [[`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs) | Rectangle-based table identification |
| Markdown output | [[`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) | Final Markdown generation |

## Enabling Structured Debug Logging

pdf-inspector uses the `tracing` crate for structured logging. Set the `RUST_LOG` environment variable to target specific modules when you troubleshoot errors with firecrawl pdf-inspector.

### Basic Logging Examples

```bash

# Show content-stream operator walking (verbose)

RUST_LOG=pdf_inspector::extractor::content_stream=trace cargo run --bin pdf2md -- file.pdf > /dev/null

# Debug font encoding issues

RUST_LOG=pdf_inspector::tounicode=debug cargo run --bin pdf2md -- file.pdf

# Full debug output for all modules

RUST_LOG=pdf_inspector=debug cargo run --bin pdf2md -- file.pdf

```

### Module-Specific Log Targets

| Use Case | `RUST_LOG` Value | What You'll See |
|----------|------------------|-----------------|
| Font metadata, CID encoding, ligatures | `pdf_inspector::extractor::fonts=debug` | Font name, encoding tables, fallback paths |
| ToUnicode CMap parsing | `pdf_inspector::tounicode=debug` | CMap parsing success/failure, character mappings |
| Column detection & reading order | `pdf_inspector::extractor::layout=debug` | Histogram values, valley detection, column splits |
| Font-size stats, heading tiers | `pdf_inspector::markdown::analysis=debug` | Font clusters, calculated heading levels |
| Table detection | `pdf_inspector::tables=debug` | Detected rectangles, heuristic scores, grid assignments |

The complete debug target reference is documented in [[`docs/debugging.md`](https://github.com/firecrawl/pdf-inspector/blob/main/docs/debugging.md)](https://github.com/firecrawl/pdf-inspector/blob/main/docs/debugging.md).

## Common Error Patterns and Diagnostic Steps

### ParseError: InvalidFileHeader

**Symptom:** Immediate failure with `InvalidFileHeader` message.

**Likely cause:** Corrupt PDF header or unsupported PDF version.

**Diagnostic command:**

```bash
RUST_LOG=pdf_inspector::detector=debug cargo run --bin pdf2md -- corrupted.pdf

```

**What to check:** Verify the file starts with `%PDF-` signature. The detector logs the header bytes it found.

### Missing or Garbled Text

**Symptom:** Output contains blank spaces or mojibake instead of expected characters.

**Likely cause:** CID font encoding without proper ToUnicode CMap.

**Diagnostic commands:**

```bash

# Check font encoding decisions

RUST_LOG=pdf_inspector::extractor::fonts=debug,pdf_inspector::tounicode=debug cargo run --bin pdf2md -- document.pdf

```

**What to look for:** Messages about "fallback to UTF-16BE" or "fallback to UTF-8" indicate missing CMaps. In [[`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs), the `parse_cmap` function handles these mappings—missing CMaps trigger encoding errors logged here.

### Incorrect Column Order

**Symptom:** Multi-column text appears scrambled or in wrong reading sequence.

**Likely cause:** Layout detection misidentified column structure.

**Diagnostic command:**

```bash
RUST_LOG=pdf_inspector::extractor::layout=debug cargo run --bin pdf2md -- newspaper.pdf

```

**What to check:** The layout module in [[`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs) logs histogram values and valley detection thresholds. Look for `early_exit` strategy decisions that may short-circuit column detection.

### Tables Not Detected

**Symptom:** Tabular data appears as plain paragraphs instead of Markdown tables.

**Likely cause:** Rectangle-based detection failed (no drawing operators) or heuristic thresholds too strict.

**Diagnostic command:**

```bash
RUST_LOG=pdf_inspector::tables=debug cargo run --bin pdf2md -- report.pdf

```

**What to inspect:** The rectangle detector in [[`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs) uses union-find clustering. If you see "no rects found," the PDF lacks explicit cell borders. Check heuristic detection fallback in [[`src/tables/detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs).

**Tuning option:** Adjust `TABLE_MAX_COLUMNS` (default 25) for unusually wide tables.

### Unexpected Headings or List Markers

**Symptom:** Wrong heading levels, missed bullet points, or false list items.

**Likely cause:** Font-size clustering failed or bullet character regex didn't match.

**Diagnostic command:**

```bash
RUST_LOG=pdf_inspector::markdown::analysis=debug,pdf_inspector::markdown::classify=debug cargo run --bin pdf2md -- document.pdf

```

**What to verify:** Font-size ratios in clustering output, and list-prefix regex matches for characters like `•`, `-`, `1.`.

### Crashes and Panics

**Symptom:** Process terminates unexpectedly without clean error.

**Likely cause:** Unhandled PDF construct, such as malformed XObject.

**Diagnostic command:**

```bash
RUST_LOG=pdf_inspector=trace cargo run --bin pdf2md -- problematic.pdf 2>&1 | tee crash.log

```

**What to find:** Stack trace with offending PDF object ID, usually logged near the panic location.

## Handling Errors in Each Language Binding

### Rust API

All public functions return `Result<T, PdfError>`. Match on error variants for targeted handling:

```rust
use pdf_inspector::{process_pdf, PdfError};

match process_pdf("file.pdf") {
    Ok(result) => println!("Pages: {}", result.pages.len()),
    Err(PdfError::ParseError(msg)) => {
        eprintln!("PDF parse error: {msg}");
        // Suggest re-exporting from source application
    }
    Err(PdfError::EncodingError(msg)) => {
        eprintln!("Font encoding error: {msg}");
        // Enable tounicode debug logging for diagnosis
    }
    Err(e) => eprintln!("Unexpected pdf-inspector error: {e}"),
}

```

### Python Bindings

Errors surface as `PdfError` exceptions:

```python
import pdf_inspector
import logging
import os

# Enable Rust logging via environment before import

os.environ["RUST_LOG"] = "pdf_inspector::tounicode=debug"

logging.basicConfig(level=logging.INFO)

try:
    result = pdf_inspector.process_pdf("document.pdf")
    print(result.markdown)
except pdf_inspector.PdfError as err:
    logging.error("pdf-inspector processing failed: %s", err)
    # Log error and re-run with more verbose target

```

### Node.js Bindings

Errors throw as standard JavaScript `Error` objects:

```javascript
import { readFileSync } from 'fs';
import { processPdf } from '@firecrawl/pdf-inspector';

const pdfBuffer = readFileSync('document.pdf');

try {
  const result = processPdf(pdfBuffer);
  console.log(result.markdown);
} catch (e) {
  console.error('pdf-inspector error:', e.message);
  // Inspect e.stack for Rust backtrace when compiled with debug symbols
}

```

Enable Rust logging from shell before Node execution:

```bash
RUST_LOG=pdf_inspector::extractor::layout=debug node process-pdf.js

```

### WebAssembly Limitations

The WASM build does not directly expose `RUST_LOG`. To troubleshoot errors with firecrawl pdf-inspector in browser environments:

1. Reproduce the issue with the CLI using equivalent `RUST_LOG` settings
2. Inspect the debug output to identify the failing module
3. Apply fixes or workarounds in your WASM integration code

## Practical Debugging Workflows

### Workflow 1: Diagnose a Missing Table

```bash

# Step 1: Check if rectangles were detected

RUST_LOG=pdf_inspector::tables::detect_rects=debug cargo run --bin pdf2md -- invoice.pdf

# Step 2: If no rects, check heuristic fallback

RUST_LOG=pdf_inspector::tables::detect_heuristic=debug cargo run --bin pdf2md -- invoice.pdf

# Step 3: Examine the source files to understand thresholds

cat src/tables/detect_rects.rs | grep -A5 "union_find"
cat src/tables/detect_heuristic.rs | grep -A5 "threshold"

```

### Workflow 2: Fix Garbled Japanese Text

```bash

# Isolate font encoding decisions

RUST_LOG=pdf_inspector::extractor::fonts=debug,pdf_inspector::tounicode=debug \
  cargo run --bin pdf2md -- japanese.pdf 2>&1 | grep -E "(CMap|encoding|fallback)"

# Check if specific font is missing ToUnicode data

grep -r "ToUnicode" src/tounicode.rs

```

### Workflow 3: Debug Layout Reading Order

```bash

# Capture full layout analysis

RUST_LOG=pdf_inspector::extractor::layout=trace cargo run --bin pdf2md -- magazine.pdf > layout.log 2>&1

# Search for column decisions

grep -E "(column|valley|histogram|early_exit)" layout.log

```

## Summary

To effectively troubleshoot errors with firecrawl pdf-inspector:

- **Map errors to pipeline stages** using the module architecture: detector → extractor (fonts, content_stream, layout) → tables → markdown
- **Use `RUST_LOG` environment variables** to enable targeted debug output for specific modules
- **Cross-reference log messages with source files**—each major component has a dedicated file in `src/`
- **Handle `PdfError` variants explicitly** in your language binding to provide user-friendly failure messages
- **Validate PDF input** independently when parse errors occur—corrupted files fail early in the detector

## Frequently Asked Questions

### How do I enable debug logging in pdf-inspector?

Set the `RUST_LOG` environment variable to target specific modules using the `tracing` crate syntax. For example: `RUST_LOG=pdf_inspector::extractor::layout=debug` for column detection issues, or `pdf_inspector=debug` for all modules. Run this before invoking the CLI, Python, or Node.js bindings.

### Why is my PDF text garbled or missing characters?

This typically indicates a **ToUnicode CMap** parsing failure in [[`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs). Enable `pdf_inspector::tounicode=debug` logging to see fallback encoding decisions. CID fonts without embedded CMaps force UTF-16BE or UTF-8 fallback, which often produces incorrect character mappings.

### Why aren't my tables being detected?

pdf-inspector tries **rectangle-based detection first** in [[`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs), which requires PDF drawing operators defining cell borders. If no rectangles are found, it falls back to **heuristic detection** in [[`src/tables/detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs). Check `pdf_inspector::tables=debug` logs to see which path was taken and why detection failed.

### What's the difference between `ParseError` and `EncodingError`?

**`ParseError`** indicates structural PDF problems—invalid headers, malformed objects, or unexpected end-of-file. These originate in the detector and early extraction phases. **`EncodingError`** signals font-specific issues, particularly missing or corrupt ToUnicode CMaps that prevent proper character decoding. Match on these variants in Rust to implement appropriate recovery strategies.