# What Is Tiled-Scan Detection? How It Identifies JBIG2 and Strip-Image PDFs

> Learn about tiled-scan detection a powerful algorithm that identifies JBIG2 and strip-image PDFs for accurate text extraction. Understand its role in PDF analysis.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: deep-dive
- Published: 2026-08-08

---

**Tiled-scan detection is an algorithm that identifies PDFs composed of many small image tiles rather than single raster images, specifically targeting JBIG2-compressed documents and strip-image PDFs to ensure accurate text extraction.**

The `firecrawl/pdf-inspector` repository implements tiled-scan detection to solve a critical edge case in document analysis. When PDFs are generated from JBIG2 compression or stored as strip-images, they appear as fragmented collections of tiny images rather than unified scanned pages. This classification system ensures the extraction pipeline treats these documents as true scanned images rather than misidentifying them as mixed-content files.

## How Tiled-Scan Detection Works

The detection algorithm resides in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) within the `tiled_scan_detection` function. It evaluates each page's image composition to distinguish between genuine mixed-content PDFs and tiled raster documents.

### The Fragmented Image Problem

Traditional scanned PDFs contain one large image per page. However, JBIG2 encoders and certain scanning workflows store documents as horizontal strips or 1-pixel-high scan lines. Without detection, these appear as thousands of separate image objects, causing misclassification as **Mixed** documents rather than **Scanned** according to the `PdfType` enum defined in [`src/types.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs).

### The Detection Algorithm

The algorithm follows a four-step statistical analysis:

1. **Page-wise image enumeration** – The detector traverses PDF objects to gather every XObject of type **Image** on each page.
2. **Tile size aggregation** – For every image object, it records pixel dimensions (`width × height`).
3. **Threshold-based scoring** – The system accumulates total pixel count across all images and counts how many tiles exceed a minimal size threshold (typically >256 px).
4. **Classification trigger** – If the total pixel count surpasses a configurable limit (default approximately 2 million pixels) **and** the number of large tiles remains below a secondary threshold, the page is flagged as a *tiled-scan*.

## Identifying JBIG2 and Strip-Image PDFs

JBIG2-generated PDFs typically store each scan line as an individual 1-pixel-high image. The pixel-sum check captures these documents because the aggregate area of thousands of tiny strips exceeds the 2M pixel threshold while the large-tile count stays minimal.

When detected, the document is reclassified from `PdfType::Mixed` to `PdfType::Scanned`. This triggers the raster-only extraction path in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs), avoiding costly mixed-mode processing and improving text extraction reliability for these edge-case formats.

## Using Tiled-Scan Detection

The functionality is exposed both via CLI and as a Rust library.

### Command Line Analysis

Run the `detect-pdf` binary to analyze documents:

```bash
RUST_LOG=pdf_inspector::detector=debug cargo run --release --bin detect-pdf -- \
  --analyze --json path/to/document.pdf

```

Successful detection returns JSON indicating the reclassification:

```json
{
  "type": "Scanned",
  "tiled_scan": true,
  "details": {
    "total_pixels": 2137425,
    "large_tile_count": 3
  }
}

```

### Programmatic Integration

Import the detector in Rust applications:

```rust
use pdf_inspector::process_pdf_with_options;
use pdf_inspector::options::PdfOptions;

let opts = PdfOptions::default();
let result = process_pdf_with_options("my.pdf", opts)
    .expect("PDF processing failed");

println!("PDF classification: {:?}", result.pdf_type);
if let Some(details) = result.tiled_scan_details {
    println!("Tiled-scan detected: {} tiles, {} total pixels",
             details.large_tile_count, details.total_pixels);
}

```

## Summary

- Tiled-scan detection identifies PDFs built from many small image tiles rather than single-page rasters.
- The algorithm in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) uses pixel-area thresholds (~2M pixels) and large-tile counts (>256px) to classify documents.
- It specifically catches JBIG2-compressed files and strip-image PDFs that store scan lines as separate objects.
- Detection triggers reclassification from `Mixed` to `Scanned`, optimizing the extraction strategy in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs).
- Available via both the `detect-pdf` CLI binary and the Rust library interface.

## Frequently Asked Questions

### What makes JBIG2 PDFs different from regular scanned PDFs?

JBIG2-compressed documents often fragment the page into thousands of 1-pixel-high horizontal strips stored as individual image objects. Regular scanned PDFs typically contain one large image per page. Without tiled-scan detection, these fragments appear as separate elements causing misclassification as mixed-content documents rather than scanned images.

### Where is the tiled-scan detection logic implemented?

The core logic lives in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) inside the `tiled_scan_detection` function. This module evaluates per-page image statistics and returns `PdfType::Scanned` when criteria are met. The type definitions reside in [`src/types.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs), while [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) orchestrates the extraction strategy based on these classifications.

### What are the default thresholds for detecting tiled scans?

The algorithm uses two primary thresholds: a total pixel count limit defaulting to approximately 2 million pixels, and a minimum tile size of 256 pixels to distinguish "large" tiles from tiny fragments. Documents must exceed the pixel total while maintaining a low large-tile count to trigger detection and reclassification.

### Why does detection change the extraction path from Mixed to Scanned?

Mixed-mode extraction attempts to balance text and image processing, which fails on heavily fragmented raster documents. By reclassifying tiled scans as pure scanned documents, PDF Inspector switches to a raster-only extraction path that treats the entire page as an image, yielding more reliable text extraction for JBIG2 and strip-image formats.