How to Debug PDF Extraction Issues Using RUST_LOG Environment Variables
Set the RUST_LOG environment variable to <target>=<level> when running pdf-inspector binaries to activate module-specific debug logs without recompiling.
The firecrawl/pdf-inspector repository uses the tracing crate to instrument its Rust codebase with structured, hierarchical logging. To debug extraction issues using RUST_LOG environment variables, you enable specific logging targets that correspond to distinct pipeline stages—such as content stream parsing, layout analysis, or table detection—allowing you to isolate whether a failure stems from font decoding, column detection heuristics, or PDF operator interpretation.
How RUST_LOG Controls the Tracing Subsystem
The project implements logging through the tracing crate, where each major component declares its own target path. When you set RUST_LOG, the subscriber filters which debug!, info!, or trace! statements actually emit to stderr based on the target hierarchy.
The variable accepts a comma-separated list of <target>=<level> pairs:
- Target: The Rust module path (e.g.,
pdf_inspector::extractor::layout) - Level: One of
trace,debug,info,warn, orerror
Because the binary reads this variable at runtime, you can toggle verbose diagnostics for any module without modifying source code or recompiling.
Essential Debugging Targets by Component
Content Stream Operators (pdf_inspector::extractor::content_stream)
The file src/extractor/content_stream.rs parses low-level PDF text operators such as Tj, TJ, and Td. Enabling trace-level logging here reveals the raw text positioning commands and the glyph strings extracted from the page content stream.
RUST_LOG=pdf_inspector::extractor::content_stream=trace \
cargo run --bin pdf2md -- file.pdf > /dev/null
Font Metrics and CMap Resolution (pdf_inspector::extractor::fonts)
When character widths appear incorrect or glyphs map to wrong Unicode values, target the fonts module. This emits font width tables, CMap parsing decisions, and fallback handling logic from the font extraction layer.
RUST_LOG=pdf_inspector::extractor::fonts=debug \
cargo run --bin pdf2md -- file.pdf > /dev/null
CID-to-Unicode Mapping (pdf_inspector::tounicode)
For PDFs with embedded CID fonts, extraction failures often occur in the ToUnicode mapping phase. This target exposes CID-to-Unicode translation attempts and decoding failures.
RUST_LOG=pdf_inspector::tounicode=debug \
cargo run --bin pdf2md -- file.pdf > /dev/null
Layout Detection and Column Analysis (pdf_inspector::extractor::layout)
The file src/extractor/layout.rs handles column detection, newspaper versus tabular classification, and pre-masking heuristics. Debug logs here show "column histogram" calculations and "valley detection" thresholds that determine how text blocks are grouped.
RUST_LOG=pdf_inspector::extractor::layout=debug \
cargo run --bin pdf2md -- file.pdf > /dev/null
Table Extraction Strategies (pdf_inspector::tables)
Defined in src/tables/mod.rs, this module orchestrates three detection strategies: rectangle-based clustering, line-based analysis, and heuristic pattern matching. Enabling this target reveals why a table was missed or misidentified, including "table rect" and "union-find clustering" messages.
RUST_LOG=pdf_inspector::tables=debug \
cargo run --bin pdf2md -- file.pdf > /dev/null
Markdown Structure Analysis (pdf_inspector::markdown::analysis)
The file src/markdown/analysis.rs performs font-size statistics and heading tier detection. Use this target to debug why certain lines become headers versus paragraphs based on font size thresholds.
RUST_LOG=pdf_inspector::markdown::analysis=debug \
cargo run --bin pdf2md -- file.pdf > /dev/null
PDF Type Classification (pdf_inspector::detector)
The file src/detector.rs classifies documents as TextBased, Scanned, or Mixed, and runs tiled-scan detection. Unlike the pdf2md binary, this component typically runs via the detect-pdf binary.
RUST_LOG=pdf_inspector::detector=debug \
cargo run --release --bin detect-pdf -- file.pdf
Combining Multiple Log Targets
You can debug extraction issues across multiple subsystems simultaneously by separating targets with commas. This is useful when tracing interactions between layout detection and table recognition.
RUST_LOG=pdf_inspector::extractor::layout=debug,pdf_inspector::tables=debug \
cargo run --bin pdf2md -- file.pdf > /dev/null
For comprehensive diagnostics when you are unsure which module is failing, enable debug logging for the entire crate:
RUST_LOG=pdf_inspector=debug \
cargo run --bin pdf2md -- file.pdf > /dev/null
Interpreting Log Output Format
Structured log lines follow the pattern <timestamp> <level> <target>: <message>. When you debug extraction issues using RUST_LOG, look for these specific log signatures:
- "column histogram" or "valley detection" → Indicates the layout engine is calculating text column boundaries in
src/extractor/layout.rs. - "table rect" or "union-find clustering" → Shows table detection progressing through geometric analysis in
src/tables/mod.rs. - "font width" or "cmap fallback" → Signals font metric resolution issues in the font extraction pipeline.
- "CID" or "ToUnicode" → Highlights Unicode mapping failures for embedded CID fonts.
By matching these messages to their emission points in the source code, you can pinpoint whether an extraction failure originates from operator parsing, font decoding, or layout heuristics.
Summary
- The
firecrawl/pdf-inspectorusestracingtargets that map directly to source files likesrc/extractor/layout.rsandsrc/tables/mod.rs. - Set
RUST_LOG=<target>=<level>to enable debug output without recompilation. - Use
pdf_inspector::extractor::content_stream=tracefor low-level PDF operator inspection. - Use
pdf_inspector::extractor::layout=debugto diagnose column detection and text flow issues. - Combine multiple targets with commas to trace cross-module interactions.
- Redirect stdout to
/dev/nullwhen runningpdf2mdto isolate stderr logs for easier reading.
Frequently Asked Questions
What is the difference between debug and trace levels in RUST_LOG?
The trace level emits the most verbose output, including every low-level operation such as individual PDF text operators (Tj, TJ) in src/extractor/content_stream.rs. The debug level provides higher-level state summaries, such as column detection results or font mapping decisions, without the per-operator noise. Use trace when you need granular step-through data, and debug for understanding algorithmic decisions.
Why am I not seeing any log output after setting RUST_LOG?
Ensure you are viewing stderr rather than stdout. The tracing crate writes logs to stderr by default, while pdf2md writes extracted content to stdout. Run your command with 2>&1 to merge streams, or redirect stdout to /dev/null (as shown in examples) to display only logs. Also verify that your target path matches exactly, such as pdf_inspector::tables not pdf_inspector::table.
Can I redirect the debug logs to a file instead of the terminal?
Yes. Because RUST_LOG controls the filter and the tracing subscriber writes to stderr, you can redirect stderr to a file using standard shell redirection. For example: RUST_LOG=pdf_inspector=debug cargo run --bin pdf2md -- file.pdf 2> debug.log. This captures all structured logs to debug.log while keeping the extracted markdown output separate.
Which target should I use if text extraction returns garbled characters?
Start with pdf_inspector::tounicode=debug to check CID-to-Unicode mapping failures, then enable pdf_inspector::extractor::fonts=debug to inspect font width tables and CMap decisions. If the issue involves missing text rather than incorrect glyphs, use pdf_inspector::extractor::content_stream=trace to verify that the PDF operators are being parsed correctly from the content stream.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →