How to Debug PDF Extraction Issues in pdf-inspector Using `RUST_LOG` Environment Variables

Set RUST_LOG to module-specific debug levels before running pdf2md or detect-pdf to trace exactly where the extraction pipeline drops or misclassifies content.

The pdf-inspector repository from Firecrawl provides powerful PDF-to-Markdown extraction through a Rust-based pipeline. When extraction fails—whether a table disappears, columns merge incorrectly, or text quality degrades—you need visibility into the internal decision-making. This guide shows you how to debug PDF extraction issues in pdf-inspector using RUST_LOG environment variables to activate targeted logging across the multi-layer pipeline.

Understanding the Logging Architecture

Both command-line binaries—pdf2md and detect-pdf—initialize logging through env_logger::init(). In src/bin/pdf2md.rs at line 221 and src/bin/detect-pdf.rs at line 67, the logger reads the RUST_LOG environment variable once at startup. This means you can control verbosity per module without recompiling the codebase.

The env_logger crate recognizes patterns like module=level separated by commas. Valid levels are error, warn, info, debug, and trace (from least to most verbose).

Extraction Pipeline Layers and Their Debug Modules

The pipeline spans several specialized layers. Each emits debug! statements that reveal intermediate decisions:

Layer Key Source Files What Debug Logs Reveal
PDF type detection src/detector.rs Template image detection (has_template_image), vector text signals (has_vector_text)
Content-stream parsing src/extractor/content_stream.rs PDF operators (Tj, TJ, Td, Tm), glyph decoding, GID-font detection
Layout analysis src/extractor/layout.rs, src/tables/*.rs Column histogram valleys, table-grid construction, rejection reasons
Text quality src/text_quality.rs Encoding issues, CID-garbage detection
Markdown conversion src/markdown/*.rs Line classification, heading merging, post-processing

For example, src/tables/detect_rects.rs contains statements like:

debug!("page {}: {} clusters …", page, clusters.len());

These logs expose exactly why a table candidate was discarded or why a page triggered OCR fallback.

Step-by-Step Debugging Workflow

1. Identify the Suspected Layer

Start with the highest-level module matching your symptom:

  • Missing or broken tables → pdf_inspector::tables
  • Incorrect column detection → pdf_inspector::extractor::layout
  • Garbled text or encoding issues → pdf_inspector::text_quality or pdf_inspector::extractor::content_stream

2. Set RUST_LOG and Run

Use a comma-separated list of module=level pairs. Enable only what you need to reduce noise.

RUST_LOG=pdf_inspector::extractor::layout=debug,pdf_inspector::tables=debug \
cargo run --bin pdf2md -- my.pdf

This keeps unrelated modules at their default (info) level while surfacing layout and table decisions.

3. Analyze Output

Expect structured messages like:

[2026-08-06T12:34:56Z] pdf_inspector::extractor::layout DEBUG: page 3: 2 columns detected, valley at x=112.4
[2026-08-06T12:34:57Z] pdf_inspector::tables DEBUG: page 3: 5 clusters with >= 6 rects
[2026-08-06T12:34:57Z] pdf_inspector::tables DEBUG: rejected: fill ratio 0.24 < 0.30

The fill ratio rejection (0.24 below 0.30 threshold) explains why your table disappeared.

4. Narrow to Finer Modules

Once you locate the problematic step, drill deeper:

RUST_LOG=pdf_inspector::extractor::content_stream=debug,pdf_inspector::text_utils=debug \
cargo run --bin pdf2md -- my.pdf

5. Iterate Until Root Cause Is Clear

Adjust filters to illuminate the full execution path. Since env_logger reads RUST_LOG only at startup, you must set the variable before each invocation.

Practical Command Examples

Run these in your shell before executing binaries:


# Layout-only debugging for column detection problems

RUST_LOG=pdf_inspector::extractor::layout=debug cargo run --release --bin pdf2md -- my.pdf

# Full extraction pipeline visibility

RUST_LOG=pdf_inspector=debug cargo run --release --bin pdf2md -- my.pdf

# Persist across multiple commands (bash/zsh)

export RUST_LOG=pdf_inspector::tables=debug,pdf_inspector::extractor::content_stream=debug
cargo run --release --bin detect-pdf -- sample.pdf

Critical Implementation Details

  • RUST_LOG is read once at startup — changing it after the process starts has no effect.
  • Module paths use crate names — pdf_inspector:: prefix is required, not src/.
  • Release builds benefit too — --release preserves debug logging; no separate debug build needed.
  • Log format includes timestamps and module paths — grep friendly for automated analysis.

Key Source Files for Reference

When tracing log output back to code logic, consult these locations:

Summary

  • RUST_LOG controls per-module verbosity without recompilation
  • Both binaries (pdf2md, detect-pdf) initialize via env_logger::init()
  • Target specific layers: extractor::layout for columns, tables for grids, content_stream for glyph parsing
  • Set before invocation — startup-only read, no runtime changes
  • Use release builds — debug logging works in optimized code

Frequently Asked Questions

Why don't I see any debug output when I set RUST_LOG=debug?

You likely set the variable after starting the binary or misspelled the module path. Verify with echo $RUST_LOG and ensure the binary reads it at launch. Also confirm you're using pdf_inspector:: prefixes, not file paths like src/.

Can I debug OCR fallback decisions?

Yes. The pdf_inspector::detector module logs has_template_image and has_vector_text results. Enable RUST_LOG=pdf_inspector::detector=debug to see why a page was classified as needing OCR versus native extraction.

How do I silence noisy modules while keeping others verbose?

Use multiple module=level pairs. For example: RUST_LOG=pdf_inspector::tables=debug,hyper=error suppresses HTTP client noise while surfacing table logic. Order does not matter; most specific match wins.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →