# Garbage Text Upgrade Mechanism for Mixed PDFs in pdf-inspector: A Technical Deep Dive

> Discover the garbage text upgrade mechanism in firecrawl pdf-inspector. Learn how it reprocesses Mixed PDFs with invisible text layers to improve text extraction and classification.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: deep-dive
- Published: 2026-08-08

---

**The garbage text upgrade mechanism in firecrawl/pdf-inspector automatically detects unreadable text extraction results from Mixed PDFs and reprocesses the document using invisible text layers, ultimately upgrading the classification to Scanned if the content remains unreadable.**

The `firecrawl/pdf-inspector` library handles complex PDF documents by implementing a sophisticated garbage text upgrade mechanism for Mixed PDFs that bridges the gap between vector text and scanned images. When a document contains both outlined text and scanned pages, the system first attempts standard extraction before automatically escalating to more aggressive parsing strategies that uncover hidden OCR layers. This ensures reliable text extraction even from documents with corrupted text streams or invisible encoding artifacts.

## How Mixed PDFs Trigger the Upgrade Mechanism

### PDF Type Classification and Mixed Detection

`pdf-inspector` classifies PDFs into four distinct types: **TextBased**, **Scanned**, **ImageBased**, and **Mixed**. A **Mixed** PDF contains both vector-outlined text and scanned-image pages, making it particularly challenging for standard extraction pipelines. According to the source code in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs), documents receive the Mixed classification when analysis reveals heterogeneous content types across different pages or layers.

### Initial Extraction and Garbage Sampling

During processing in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) (lines 3791-3808), the system samples up to 200 text items from the initial extraction to evaluate quality before committing to the full document parse. This sampling occurs specifically when `pdf_type == PdfType::Mixed`, allowing the garbage text upgrade mechanism to activate early in the pipeline:

```rust
if pdf_type == PdfType::Mixed {
    let sample: String = items.iter()
        .filter(|item| options.page_filter.as_ref()
            .is_none_or(|f| f.contains(&item.page)))
        .take(200)
        .map(|item| item.text.as_str())
        .collect();
    
    if is_garbage_text(&sample) || sample.trim().is_empty() {
        extractor::extract_positioned_text_include_invisible_with_folio_context(...)
    } else {
        result
    }
}

```

## The Garbage Detection Algorithm

The `is_garbage_text` function in [`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs) (lines 31-71) implements a statistical heuristic to identify unreadable content. The algorithm counts alphanumeric versus non-alphanumeric characters while ignoring Markdown-added syntax and decorative leaders:

```rust
pub(crate) fn is_garbage_text(markdown: &str) -> bool {
    // ... count alphanum / non-alphanum ...
    let total = alphanum + non_alphanum;
    total >= 50 && alphanum * 2 < total
}

```

Text qualifies as garbage when **non-alphanumeric characters constitute more than 50% of the content** and the **total character count exceeds 50**. This threshold effectively filters out documents containing mostly symbols, corrupted encoding sequences, or binary artifacts masquerading as readable text.

## The Two-Phase Upgrade Process

### Phase 1: Standard Extraction with Quality Check

When processing begins on a Mixed PDF, `pdf-inspector` first invokes standard text extraction methods. If the sampled text passes the garbage test (less than 50% non-alphanumeric), processing continues normally. If the sample is empty or fails the quality check, the system immediately prepares for upgrade.

### Phase 2: Invisible Text Recovery

Upon detecting garbage, the extractor upgrades to `extract_positioned_text_include_invisible_with_folio_context()`, a more permissive extraction mode that parses normally invisible text objects. These hidden layers often contain OCR results from scanned document pages that standard extraction misses. This second pass effectively turns the document into a Scanned-type PDF for downstream handling if the invisible text reveals readable content.

### Final Classification to Scanned Type

If the Markdown output still registers as garbage after the invisible text extraction (lines 4012-4015 in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)), the PDF type upgrades from Mixed to Scanned:

```rust
if pdf_type == PdfType::Mixed && markdown.as_ref().is_some_and(|m| is_garbage_text(m)) {
    // Treat the document as Scanned for downstream consumers
    // (e.g., the CLI will report "mixed → scanned")
}

```

This upgrade signals that the document requires full OCR processing rather than vector text extraction, optimizing resource allocation for image-heavy documents.

## Practical Implementation Examples

### Command Line Interface

When running the CLI on a Mixed PDF, the garbage text upgrade mechanism operates automatically:

```bash
pdf2md --json mixed-document.pdf

```

The tool first attempts normal extraction. If the sample text qualifies as garbage, it automatically re-extracts with invisible objects enabled. If the final Markdown remains unreadable, the CLI reports the PDF as "mixed → scanned" in the JSON output.

### Programmatic Library Usage

Developers can detect upgrade events programmatically:

```rust
use pdf_inspector::{process_pdf_with_options, PdfOptions};

let opts = PdfOptions::default();
let result = process_pdf_with_options("mixed-document.pdf", opts).unwrap();

if result.pdf_type == pdf_inspector::PdfType::Scanned {
    println!("Garbage-text upgrade triggered – PDF treated as scanned.");
}

```

## Summary

- The garbage text upgrade mechanism activates exclusively for **Mixed PDFs** when initial extraction yields unreadable results
- **Detection relies** on the `is_garbage_text()` function in [`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs), which flags text containing more than 50% non-alphanumeric characters
- The system performs a **two-phase extraction**: standard parsing followed by invisible text layer recovery using `extract_positioned_text_include_invisible_with_folio_context()`
- Documents failing both phases undergo **type promotion** from Mixed to Scanned in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), ensuring downstream OCR processing
- Sampling **200 text items** during the initial pass prevents wasted computation on severely corrupted documents

## Frequently Asked Questions

### What qualifies as "garbage text" in pdf-inspector?

Text qualifies as garbage when non-alphanumeric characters constitute more than 50% of the sampled content and the total character count exceeds 50, as implemented in [`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs) lines 31-71. This statistical approach filters out corrupted encoding, binary artifacts, and symbol-heavy noise while preserving legitimate documents with heavy punctuation.

### How does the upgrade mechanism handle invisible OCR layers?

When initial extraction fails the garbage test in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), the system invokes `extract_positioned_text_include_invisible_with_folio_context()` from [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs). This function parses text objects normally hidden from standard extraction, revealing OCR results embedded behind scanned images that typical vector text parsers cannot access.

### Why does pdf-inspector upgrade Mixed PDFs to Scanned type rather than retrying indefinitely?

The upgrade to Scanned type (lines 4012-4015 in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)) represents a deterministic fallback that prevents infinite processing loops and resource exhaustion. Once classified as Scanned, downstream pipelines apply full OCR processing rather than attempting further vector text extraction, optimizing CPU usage for documents where text extraction is fundamentally unreliable.

### Can developers customize the garbage detection thresholds?

Currently, the thresholds—requiring 50% non-alphanumeric ratio and 50-character minimum—are hardcoded constants in [`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs) within the `is_garbage_text` function. Similarly, the sampling limit of 200 text items in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) is fixed, ensuring consistent cross-document behavior in the garbage text upgrade mechanism.