# How to Debug PDF Extraction Issues Using RUST_LOG Environment Variables

> Debug PDF extraction issues with RUST_LOG environment variables in firecrawl pdf-inspector. Activate module specific logs without recompiling. Get granular insights for faster troubleshooting.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-08

---

**Set the `RUST_LOG` environment variable to `<target>=<level>` when running pdf-inspector binaries to activate module-specific debug logs without recompiling.**

The `firecrawl/pdf-inspector` repository uses the `tracing` crate to instrument its Rust codebase with structured, hierarchical logging. To debug extraction issues using `RUST_LOG` environment variables, you enable specific logging targets that correspond to distinct pipeline stages—such as content stream parsing, layout analysis, or table detection—allowing you to isolate whether a failure stems from font decoding, column detection heuristics, or PDF operator interpretation.

## How RUST_LOG Controls the Tracing Subsystem

The project implements logging through the `tracing` crate, where each major component declares its own target path. When you set `RUST_LOG`, the subscriber filters which `debug!`, `info!`, or `trace!` statements actually emit to stderr based on the target hierarchy.

The variable accepts a comma-separated list of `<target>=<level>` pairs:

- **Target**: The Rust module path (e.g., `pdf_inspector::extractor::layout`)
- **Level**: One of `trace`, `debug`, `info`, `warn`, or `error`

Because the binary reads this variable at runtime, you can toggle verbose diagnostics for any module without modifying source code or recompiling.

## Essential Debugging Targets by Component

### Content Stream Operators (`pdf_inspector::extractor::content_stream`)

The file [`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs) parses low-level PDF text operators such as `Tj`, `TJ`, and `Td`. Enabling trace-level logging here reveals the raw text positioning commands and the glyph strings extracted from the page content stream.

```bash
RUST_LOG=pdf_inspector::extractor::content_stream=trace \
cargo run --bin pdf2md -- file.pdf > /dev/null

```

### Font Metrics and CMap Resolution (`pdf_inspector::extractor::fonts`)

When character widths appear incorrect or glyphs map to wrong Unicode values, target the fonts module. This emits font width tables, CMap parsing decisions, and fallback handling logic from the font extraction layer.

```bash
RUST_LOG=pdf_inspector::extractor::fonts=debug \
cargo run --bin pdf2md -- file.pdf > /dev/null

```

### CID-to-Unicode Mapping (`pdf_inspector::tounicode`)

For PDFs with embedded CID fonts, extraction failures often occur in the ToUnicode mapping phase. This target exposes CID-to-Unicode translation attempts and decoding failures.

```bash
RUST_LOG=pdf_inspector::tounicode=debug \
cargo run --bin pdf2md -- file.pdf > /dev/null

```

### Layout Detection and Column Analysis (`pdf_inspector::extractor::layout`)

The file [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs) handles column detection, newspaper versus tabular classification, and pre-masking heuristics. Debug logs here show "column histogram" calculations and "valley detection" thresholds that determine how text blocks are grouped.

```bash
RUST_LOG=pdf_inspector::extractor::layout=debug \
cargo run --bin pdf2md -- file.pdf > /dev/null

```

### Table Extraction Strategies (`pdf_inspector::tables`)

Defined in [`src/tables/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/mod.rs), this module orchestrates three detection strategies: rectangle-based clustering, line-based analysis, and heuristic pattern matching. Enabling this target reveals why a table was missed or misidentified, including "table rect" and "union-find clustering" messages.

```bash
RUST_LOG=pdf_inspector::tables=debug \
cargo run --bin pdf2md -- file.pdf > /dev/null

```

### Markdown Structure Analysis (`pdf_inspector::markdown::analysis`)

The file [`src/markdown/analysis.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/analysis.rs) performs font-size statistics and heading tier detection. Use this target to debug why certain lines become headers versus paragraphs based on font size thresholds.

```bash
RUST_LOG=pdf_inspector::markdown::analysis=debug \
cargo run --bin pdf2md -- file.pdf > /dev/null

```

### PDF Type Classification (`pdf_inspector::detector`)

The file [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) classifies documents as TextBased, Scanned, or Mixed, and runs tiled-scan detection. Unlike the `pdf2md` binary, this component typically runs via the `detect-pdf` binary.

```bash
RUST_LOG=pdf_inspector::detector=debug \
cargo run --release --bin detect-pdf -- file.pdf

```

## Combining Multiple Log Targets

You can debug extraction issues across multiple subsystems simultaneously by separating targets with commas. This is useful when tracing interactions between layout detection and table recognition.

```bash
RUST_LOG=pdf_inspector::extractor::layout=debug,pdf_inspector::tables=debug \
cargo run --bin pdf2md -- file.pdf > /dev/null

```

For comprehensive diagnostics when you are unsure which module is failing, enable debug logging for the entire crate:

```bash
RUST_LOG=pdf_inspector=debug \
cargo run --bin pdf2md -- file.pdf > /dev/null

```

## Interpreting Log Output Format

Structured log lines follow the pattern `<timestamp> <level> <target>: <message>`. When you debug extraction issues using `RUST_LOG`, look for these specific log signatures:

- **"column histogram"** or **"valley detection"** → Indicates the layout engine is calculating text column boundaries in [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs).
- **"table rect"** or **"union-find clustering"** → Shows table detection progressing through geometric analysis in [`src/tables/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/mod.rs).
- **"font width"** or **"cmap fallback"** → Signals font metric resolution issues in the font extraction pipeline.
- **"CID"** or **"ToUnicode"** → Highlights Unicode mapping failures for embedded CID fonts.

By matching these messages to their emission points in the source code, you can pinpoint whether an extraction failure originates from operator parsing, font decoding, or layout heuristics.

## Summary

- The `firecrawl/pdf-inspector` uses `tracing` targets that map directly to source files like [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs) and [`src/tables/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/mod.rs).
- Set `RUST_LOG=<target>=<level>` to enable debug output without recompilation.
- Use `pdf_inspector::extractor::content_stream=trace` for low-level PDF operator inspection.
- Use `pdf_inspector::extractor::layout=debug` to diagnose column detection and text flow issues.
- Combine multiple targets with commas to trace cross-module interactions.
- Redirect stdout to `/dev/null` when running `pdf2md` to isolate stderr logs for easier reading.

## Frequently Asked Questions

### What is the difference between `debug` and `trace` levels in RUST_LOG?

The `trace` level emits the most verbose output, including every low-level operation such as individual PDF text operators (`Tj`, `TJ`) in [`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs). The `debug` level provides higher-level state summaries, such as column detection results or font mapping decisions, without the per-operator noise. Use `trace` when you need granular step-through data, and `debug` for understanding algorithmic decisions.

### Why am I not seeing any log output after setting RUST_LOG?

Ensure you are viewing stderr rather than stdout. The `tracing` crate writes logs to stderr by default, while `pdf2md` writes extracted content to stdout. Run your command with `2>&1` to merge streams, or redirect stdout to `/dev/null` (as shown in examples) to display only logs. Also verify that your target path matches exactly, such as `pdf_inspector::tables` not `pdf_inspector::table`.

### Can I redirect the debug logs to a file instead of the terminal?

Yes. Because `RUST_LOG` controls the filter and the tracing subscriber writes to stderr, you can redirect stderr to a file using standard shell redirection. For example: `RUST_LOG=pdf_inspector=debug cargo run --bin pdf2md -- file.pdf 2> debug.log`. This captures all structured logs to `debug.log` while keeping the extracted markdown output separate.

### Which target should I use if text extraction returns garbled characters?

Start with `pdf_inspector::tounicode=debug` to check CID-to-Unicode mapping failures, then enable `pdf_inspector::extractor::fonts=debug` to inspect font width tables and CMap decisions. If the issue involves missing text rather than incorrect glyphs, use `pdf_inspector::extractor::content_stream=trace` to verify that the PDF operators are being parsed correctly from the content stream.