# How pdf-inspector Detects PDF Type: TextBased, Scanned, Mixed, and ImageBased Classification

> Discover how pdf-inspector classifies PDFs as TextBased, Scanned, Mixed, or ImageBased by analyzing content streams and page metrics. Understand your documents faster.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-08

---

**pdf-inspector detects PDF type by sampling pages and analyzing content streams for text operators, images, fonts, and vector graphics, then aggregating per-page metrics to classify documents as TextBased, Scanned, Mixed, or ImageBased.**

The open-source Rust library `firecrawl/pdf-inspector` provides robust PDF classification through a three-phase pipeline implemented in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs). This system powers both the library API exposed in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) and the `detect-pdf` CLI binary found in [`src/bin/detect_pdf.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs).

## Overview of the Detection Pipeline

The detection process follows three distinct phases: **page sampling**, **content analysis**, and **hierarchical classification**. By default, the system uses a `Sample(8)` strategy that evenly distributes up to eight pages across the document, including the first and last pages. This approach balances speed with accuracy for large documents.

The core detection logic resides in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs), where the `DetectionConfig` struct defines tunable thresholds and the `PdfTypeResult` struct returns the final classification with confidence scores and OCR recommendations.

## Phase 1 – Page Sampling Strategy

pdf-inspector selects pages for analysis according to a configurable `ScanStrategy` enum. The `distribute_pages` function ensures the first and last pages are always included, with remaining indices spaced evenly throughout the document.

Available strategies include:
- **`Sample(N)`** – Evenly distributes N pages (default: 8)
- **`Pages([...])`** – Analyzes specific page numbers only
- **`Full`** – Scans every page in the document
- **`EarlyExit`** – Stops early if classification confidence reaches threshold

For a 100-page PDF using `Sample(8)`, the sampled indices might be `[1, 13, 25, 37, 49, 61, 73, 85, 100]`.

## Phase 2 – Content Stream Analysis

For each sampled page, the `analyze_page_content` function parses the PDF content stream and populates a `PageAnalysis` struct with granular metrics.

### Text Detection via Operators

The system counts `Tj` and `TJ` text-showing operators to determine text presence. A page qualifies as text-rich when `text_operator_count` exceeds the `min_text_ops_per_page` threshold (default: 3). The analysis also tracks `unique_text_chars` and `unique_alphanum_chars` to distinguish meaningful text from decorative glyphs.

### Image and Template Detection

The detector identifies images by monitoring `Do` (Draw Object) operators. It specifically flags **template images**—single full-page background images that indicate scanned document backgrounds—through the `has_template_image` boolean.

### Vector Text Identification

When a page contains vector graphics masquerading as text, the system detects this through path operation density. If path operations exceed 1,000 while unique alphanumerics remain below 30, the `has_vector_text` flag activates, indicating text rendered as curves rather than font glyphs.

### Font Resolution and Decodability

After resolving font names to their underlying `ObjectId`s while respecting PDF resource inheritance, the analyzer checks three critical font properties:
- **`has_identity_h_no_tounicode`** – Identity-H/V fonts lacking ToUnicode CMaps
- **`has_only_type3_fonts`** – All fonts are Type3 without Unicode mappings
- **`has_decodable_text_fonts`** – At least one font supports Unicode extraction

These flags determine whether text is extractable or requires OCR.

## Phase 3 – Classification Logic

The classification logic (implemented around lines 10,030–10,335 in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)) follows a strict decision hierarchy:

1. **Template-image check** – If `has_template_images` is true and `pages_with_text > 0`, classify as `Mixed` (OCR recommended)
2. **Text ratio evaluation** – If `text_ratio >= text_page_ratio_threshold` (default: 0.6), classify as `TextBased`
3. **Image-only detection** – If no text exists but images or vector text are present, classify as `Scanned` (or `ImageBased` if vector text detected)
4. **Mixed content** – If text coexists with images or vector graphics, classify as `Mixed`
5. **Fallback** – Default to `TextBased` for edge cases

A secondary "newspaper-layout" heuristic (lines 10,050–10,074) may upgrade `ocr_recommended` for dense, multi-column PDFs even when the primary type is `TextBased`.

## Configuration and Thresholds

The `DetectionConfig` struct allows customization of detection behavior:

```rust
pub struct DetectionConfig {
    pub strategy: ScanStrategy,          // Page selection method
    pub min_text_ops_per_page: u32,     // Minimum Tj/TJ ops for text pages (default: 3)
    pub text_page_ratio_threshold: f32 // Ratio for TextBased classification (default: 0.6)
}

```

Default settings favor speed while maintaining accuracy for typical document layouts.

## Using the Detection API

### Rust Library Usage

Import the detection functions from [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) to classify PDFs programmatically:

```rust
use pdf_inspector::{detect_pdf_type, DetectionConfig, ScanStrategy};

// Basic detection with defaults (Sample 8 pages)
let result = detect_pdf_type("document.pdf")?;
println!("Type: {:?}, Confidence: {}", result.pdf_type, result.confidence);

// Custom configuration – analyze only first 3 pages with stricter thresholds
let config = DetectionConfig {
    strategy: ScanStrategy::Pages(vec![1, 2, 3]),
    min_text_ops_per_page: 5,
    text_page_ratio_threshold: 0.7,
};
let result = pdf_inspector::detect_pdf_type_with_config("document.pdf", config)?;

```

### CLI Binary Usage

The `detect-pdf` binary in [`src/bin/detect_pdf.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs) provides command-line access:

```bash

# Basic human-readable output

detect-pdf path/to/document.pdf

# JSON output for integration

detect-pdf --json path/to/document.pdf

# Custom page sampling

detect-pdf --strategy=pages=1,10,20 path/to/document.pdf

```

### Accessing OCR Recommendations

The `PdfTypeResult` struct provides actionable OCR guidance:

```rust
if result.ocr_recommended {
    for page in &result.pages_needing_ocr {
        let reasons = &result.ocr_reasons_by_page[page];
        println!("Page {} needs OCR: {:?}", page, reasons);
    }
}

```

**Scanned** and **ImageBased** PDFs return all pages in `pages_needing_ocr`, while **Mixed** PDFs return only specific pages requiring OCR due to undecodable fonts or image content.

## Summary

- pdf-inspector detects PDF type through a three-phase pipeline: page sampling, content stream analysis, and hierarchical classification.
- The system analyzes text operators (`Tj`/`TJ`), image objects, vector graphics, and font properties in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs).
- Default `Sample(8)` strategy balances performance with accuracy by examining evenly distributed pages including first and last.
- Classification depends on configurable thresholds: `min_text_ops_per_page` (default 3) and `text_page_ratio_threshold` (default 0.6).
- Results include specific OCR recommendations via `pages_needing_ocr` and `ocr_reasons_by_page` for targeted processing.

## Frequently Asked Questions

### What sampling strategy does pdf-inspector use by default?

By default, pdf-inspector uses `ScanStrategy::Sample(8)`, which selects up to eight evenly distributed pages across the document while always including the first and last pages. This strategy provides representative analysis without processing every page, making it suitable for large documents. You can override this with `Full` (all pages), `Pages([...])` (specific pages), or `EarlyExit` (stop when confident).

### How does pdf-inspector distinguish between TextBased and Scanned PDFs?

pdf-inspector distinguishes TextBased from Scanned PDFs by counting text-showing operators (`Tj`/`TJ`) and calculating the ratio of text-rich pages. If the `text_ratio` meets or exceeds the `text_page_ratio_threshold` (default 0.6), the document is TextBased. If no text operators are found but images or vector text exist, the document is classified as Scanned or ImageBased, triggering OCR recommendations for all pages.

### What triggers the Mixed PDF classification?

The Mixed classification triggers in three scenarios: when template images coexist with text content, when the text ratio falls below the threshold but some text exists alongside images, or when specific pages contain undecodable fonts or vector text masquerading as content. Mixed documents return `ocr_recommended: true` with specific `pages_needing_ocr` identified for targeted OCR processing.

### Can I customize the detection thresholds?

Yes, the `DetectionConfig` struct allows full customization of detection parameters. You can adjust `min_text_ops_per_page` to require more text operators before a page counts as text-rich, modify `text_page_ratio_threshold` to change the classification sensitivity, or specify exact pages via `ScanStrategy::Pages`. Pass your custom configuration to `detect_pdf_type_with_config` available in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs).