How to Debug PDF Extraction Issues Using RUST_LOG Environment Variables

Set the RUST_LOG environment variable to <target>=<level> when running pdf-inspector binaries to activate module-specific debug logs without recompiling.

The firecrawl/pdf-inspector repository uses the tracing crate to instrument its Rust codebase with structured, hierarchical logging. To debug extraction issues using RUST_LOG environment variables, you enable specific logging targets that correspond to distinct pipeline stages—such as content stream parsing, layout analysis, or table detection—allowing you to isolate whether a failure stems from font decoding, column detection heuristics, or PDF operator interpretation.

How RUST_LOG Controls the Tracing Subsystem

The project implements logging through the tracing crate, where each major component declares its own target path. When you set RUST_LOG, the subscriber filters which debug!, info!, or trace! statements actually emit to stderr based on the target hierarchy.

The variable accepts a comma-separated list of <target>=<level> pairs:

  • Target: The Rust module path (e.g., pdf_inspector::extractor::layout)
  • Level: One of trace, debug, info, warn, or error

Because the binary reads this variable at runtime, you can toggle verbose diagnostics for any module without modifying source code or recompiling.

Essential Debugging Targets by Component

Content Stream Operators (pdf_inspector::extractor::content_stream)

The file src/extractor/content_stream.rs parses low-level PDF text operators such as Tj, TJ, and Td. Enabling trace-level logging here reveals the raw text positioning commands and the glyph strings extracted from the page content stream.

RUST_LOG=pdf_inspector::extractor::content_stream=trace \
cargo run --bin pdf2md -- file.pdf > /dev/null

Font Metrics and CMap Resolution (pdf_inspector::extractor::fonts)

When character widths appear incorrect or glyphs map to wrong Unicode values, target the fonts module. This emits font width tables, CMap parsing decisions, and fallback handling logic from the font extraction layer.

RUST_LOG=pdf_inspector::extractor::fonts=debug \
cargo run --bin pdf2md -- file.pdf > /dev/null

CID-to-Unicode Mapping (pdf_inspector::tounicode)

For PDFs with embedded CID fonts, extraction failures often occur in the ToUnicode mapping phase. This target exposes CID-to-Unicode translation attempts and decoding failures.

RUST_LOG=pdf_inspector::tounicode=debug \
cargo run --bin pdf2md -- file.pdf > /dev/null

Layout Detection and Column Analysis (pdf_inspector::extractor::layout)

The file src/extractor/layout.rs handles column detection, newspaper versus tabular classification, and pre-masking heuristics. Debug logs here show "column histogram" calculations and "valley detection" thresholds that determine how text blocks are grouped.

RUST_LOG=pdf_inspector::extractor::layout=debug \
cargo run --bin pdf2md -- file.pdf > /dev/null

Table Extraction Strategies (pdf_inspector::tables)

Defined in src/tables/mod.rs, this module orchestrates three detection strategies: rectangle-based clustering, line-based analysis, and heuristic pattern matching. Enabling this target reveals why a table was missed or misidentified, including "table rect" and "union-find clustering" messages.

RUST_LOG=pdf_inspector::tables=debug \
cargo run --bin pdf2md -- file.pdf > /dev/null

Markdown Structure Analysis (pdf_inspector::markdown::analysis)

The file src/markdown/analysis.rs performs font-size statistics and heading tier detection. Use this target to debug why certain lines become headers versus paragraphs based on font size thresholds.

RUST_LOG=pdf_inspector::markdown::analysis=debug \
cargo run --bin pdf2md -- file.pdf > /dev/null

PDF Type Classification (pdf_inspector::detector)

The file src/detector.rs classifies documents as TextBased, Scanned, or Mixed, and runs tiled-scan detection. Unlike the pdf2md binary, this component typically runs via the detect-pdf binary.

RUST_LOG=pdf_inspector::detector=debug \
cargo run --release --bin detect-pdf -- file.pdf

Combining Multiple Log Targets

You can debug extraction issues across multiple subsystems simultaneously by separating targets with commas. This is useful when tracing interactions between layout detection and table recognition.

RUST_LOG=pdf_inspector::extractor::layout=debug,pdf_inspector::tables=debug \
cargo run --bin pdf2md -- file.pdf > /dev/null

For comprehensive diagnostics when you are unsure which module is failing, enable debug logging for the entire crate:

RUST_LOG=pdf_inspector=debug \
cargo run --bin pdf2md -- file.pdf > /dev/null

Interpreting Log Output Format

Structured log lines follow the pattern <timestamp> <level> <target>: <message>. When you debug extraction issues using RUST_LOG, look for these specific log signatures:

  • "column histogram" or "valley detection" → Indicates the layout engine is calculating text column boundaries in src/extractor/layout.rs.
  • "table rect" or "union-find clustering" → Shows table detection progressing through geometric analysis in src/tables/mod.rs.
  • "font width" or "cmap fallback" → Signals font metric resolution issues in the font extraction pipeline.
  • "CID" or "ToUnicode" → Highlights Unicode mapping failures for embedded CID fonts.

By matching these messages to their emission points in the source code, you can pinpoint whether an extraction failure originates from operator parsing, font decoding, or layout heuristics.

Summary

  • The firecrawl/pdf-inspector uses tracing targets that map directly to source files like src/extractor/layout.rs and src/tables/mod.rs.
  • Set RUST_LOG=<target>=<level> to enable debug output without recompilation.
  • Use pdf_inspector::extractor::content_stream=trace for low-level PDF operator inspection.
  • Use pdf_inspector::extractor::layout=debug to diagnose column detection and text flow issues.
  • Combine multiple targets with commas to trace cross-module interactions.
  • Redirect stdout to /dev/null when running pdf2md to isolate stderr logs for easier reading.

Frequently Asked Questions

What is the difference between debug and trace levels in RUST_LOG?

The trace level emits the most verbose output, including every low-level operation such as individual PDF text operators (Tj, TJ) in src/extractor/content_stream.rs. The debug level provides higher-level state summaries, such as column detection results or font mapping decisions, without the per-operator noise. Use trace when you need granular step-through data, and debug for understanding algorithmic decisions.

Why am I not seeing any log output after setting RUST_LOG?

Ensure you are viewing stderr rather than stdout. The tracing crate writes logs to stderr by default, while pdf2md writes extracted content to stdout. Run your command with 2>&1 to merge streams, or redirect stdout to /dev/null (as shown in examples) to display only logs. Also verify that your target path matches exactly, such as pdf_inspector::tables not pdf_inspector::table.

Can I redirect the debug logs to a file instead of the terminal?

Yes. Because RUST_LOG controls the filter and the tracing subscriber writes to stderr, you can redirect stderr to a file using standard shell redirection. For example: RUST_LOG=pdf_inspector=debug cargo run --bin pdf2md -- file.pdf 2> debug.log. This captures all structured logs to debug.log while keeping the extracted markdown output separate.

Which target should I use if text extraction returns garbled characters?

Start with pdf_inspector::tounicode=debug to check CID-to-Unicode mapping failures, then enable pdf_inspector::extractor::fonts=debug to inspect font width tables and CMap decisions. If the issue involves missing text rather than incorrect glyphs, use pdf_inspector::extractor::content_stream=trace to verify that the PDF operators are being parsed correctly from the content stream.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →