What Is Tiled-Scan Detection? How It Identifies JBIG2 and Strip-Image PDFs
Tiled-scan detection is an algorithm that identifies PDFs composed of many small image tiles rather than single raster images, specifically targeting JBIG2-compressed documents and strip-image PDFs to ensure accurate text extraction.
The firecrawl/pdf-inspector repository implements tiled-scan detection to solve a critical edge case in document analysis. When PDFs are generated from JBIG2 compression or stored as strip-images, they appear as fragmented collections of tiny images rather than unified scanned pages. This classification system ensures the extraction pipeline treats these documents as true scanned images rather than misidentifying them as mixed-content files.
How Tiled-Scan Detection Works
The detection algorithm resides in src/detector.rs within the tiled_scan_detection function. It evaluates each page's image composition to distinguish between genuine mixed-content PDFs and tiled raster documents.
The Fragmented Image Problem
Traditional scanned PDFs contain one large image per page. However, JBIG2 encoders and certain scanning workflows store documents as horizontal strips or 1-pixel-high scan lines. Without detection, these appear as thousands of separate image objects, causing misclassification as Mixed documents rather than Scanned according to the PdfType enum defined in src/types.rs.
The Detection Algorithm
The algorithm follows a four-step statistical analysis:
- Page-wise image enumeration – The detector traverses PDF objects to gather every XObject of type Image on each page.
- Tile size aggregation – For every image object, it records pixel dimensions (
width × height). - Threshold-based scoring – The system accumulates total pixel count across all images and counts how many tiles exceed a minimal size threshold (typically >256 px).
- Classification trigger – If the total pixel count surpasses a configurable limit (default approximately 2 million pixels) and the number of large tiles remains below a secondary threshold, the page is flagged as a tiled-scan.
Identifying JBIG2 and Strip-Image PDFs
JBIG2-generated PDFs typically store each scan line as an individual 1-pixel-high image. The pixel-sum check captures these documents because the aggregate area of thousands of tiny strips exceeds the 2M pixel threshold while the large-tile count stays minimal.
When detected, the document is reclassified from PdfType::Mixed to PdfType::Scanned. This triggers the raster-only extraction path in src/extractor/mod.rs, avoiding costly mixed-mode processing and improving text extraction reliability for these edge-case formats.
Using Tiled-Scan Detection
The functionality is exposed both via CLI and as a Rust library.
Command Line Analysis
Run the detect-pdf binary to analyze documents:
RUST_LOG=pdf_inspector::detector=debug cargo run --release --bin detect-pdf -- \
--analyze --json path/to/document.pdf
Successful detection returns JSON indicating the reclassification:
{
"type": "Scanned",
"tiled_scan": true,
"details": {
"total_pixels": 2137425,
"large_tile_count": 3
}
}
Programmatic Integration
Import the detector in Rust applications:
use pdf_inspector::process_pdf_with_options;
use pdf_inspector::options::PdfOptions;
let opts = PdfOptions::default();
let result = process_pdf_with_options("my.pdf", opts)
.expect("PDF processing failed");
println!("PDF classification: {:?}", result.pdf_type);
if let Some(details) = result.tiled_scan_details {
println!("Tiled-scan detected: {} tiles, {} total pixels",
details.large_tile_count, details.total_pixels);
}
Summary
- Tiled-scan detection identifies PDFs built from many small image tiles rather than single-page rasters.
- The algorithm in
src/detector.rsuses pixel-area thresholds (~2M pixels) and large-tile counts (>256px) to classify documents. - It specifically catches JBIG2-compressed files and strip-image PDFs that store scan lines as separate objects.
- Detection triggers reclassification from
MixedtoScanned, optimizing the extraction strategy insrc/extractor/mod.rs. - Available via both the
detect-pdfCLI binary and the Rust library interface.
Frequently Asked Questions
What makes JBIG2 PDFs different from regular scanned PDFs?
JBIG2-compressed documents often fragment the page into thousands of 1-pixel-high horizontal strips stored as individual image objects. Regular scanned PDFs typically contain one large image per page. Without tiled-scan detection, these fragments appear as separate elements causing misclassification as mixed-content documents rather than scanned images.
Where is the tiled-scan detection logic implemented?
The core logic lives in src/detector.rs inside the tiled_scan_detection function. This module evaluates per-page image statistics and returns PdfType::Scanned when criteria are met. The type definitions reside in src/types.rs, while src/extractor/mod.rs orchestrates the extraction strategy based on these classifications.
What are the default thresholds for detecting tiled scans?
The algorithm uses two primary thresholds: a total pixel count limit defaulting to approximately 2 million pixels, and a minimum tile size of 256 pixels to distinguish "large" tiles from tiny fragments. Documents must exceed the pixel total while maintaining a low large-tile count to trigger detection and reclassification.
Why does detection change the extraction path from Mixed to Scanned?
Mixed-mode extraction attempts to balance text and image processing, which fails on heavily fragmented raster documents. By reclassifying tiled scans as pure scanned documents, PDF Inspector switches to a raster-only extraction path that treats the entire page as an image, yielding more reliable text extraction for JBIG2 and strip-image formats.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →