How pdf-inspector Determines If a PDF Is Text-Based or Scanned Without Loading the Entire Document
pdf-inspector classifies PDFs as TextBased, Scanned, Mixed, or ImageBased by sampling only a tiny fraction of the file—examining metadata, text quality ratios, and image tile patterns—rather than parsing every page.
The firecrawl/pdf-inspector Rust library provides fast PDF type detection that scales to documents with hundreds of pages. Its detection pipeline operates in constant time relative to page count by strategically limiting what it reads from the PDF structure. This article breaks down the three lightweight heuristics that power this classification system.
The Three-Stage Detection Pipeline
The core logic resides in [src/detector.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs). The detector runs checks in sequence and exits early once a definitive classification is reached.
Stage 1: Metadata Scan for Structure Tree
The detector first opens the PDF with lopdf and reads only the document catalog, pages tree, and first ~10 page dictionaries. This minimal object scan identifies whether a Structure Tree is present.
- Structure Tree detected → immediate TextBased classification
- Indicates a properly tagged PDF meant for accessibility and text extraction
This check avoids any text or image rendering. It purely inspects PDF object headers and cross-reference tables.
Stage 2: Text Quality Sampling via Alphanumeric Ratio
If no Structure Tree exists, the detector samples text from early pages using the helper module [src/text_quality.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs).
The process:
- Extract raw text content from sampled pages via the extractor pipeline
- Calculate the proportion of printable alphanumeric characters (ASCII/Unicode letters and numbers)
- Compare against the 50% threshold
| Ratio Result | Classification |
|---|---|
| ≥ 50% alphanumeric | TextBased |
| < 50% alphanumeric | Scanned |
This heuristic catches PDFs with hidden OCR text overlays—common in "scanned" documents that contain garbled or sparse text strings beneath image content.
Stage 3: Tiled-Scan Image Detection
When text quality is ambiguous, the detector examines XObject resources for raster images. The detect_tiled_scan function implements this:
- Aggregates total pixel count across all image tiles on sampled pages
- Checks if summed area exceeds 2 million pixels (the "tiled-scan" threshold)
- Verifies no single tile exceeds the per-tile limit
This flags JBIG2-compressed or strip-scanned PDFs where text is visually rendered as many small images rather than as actual text objects.
| Condition | Classification |
|---|---|
| Tile area > 2M px + no dominant tile | Scanned |
| Otherwise | Mixed or ImageBased depending on other factors |
Why This Approach Avoids Full Document Loading
Traditional PDF analysis requires parsing every page content stream—a O(n) operation where n equals page count. pdf-inspector's approach achieves O(1) time complexity by:
- Reading only the PDF header and object table (constant size)
- Sampling a fixed number of early pages (typically 3–10)
- Avoiding full content stream parsing, font embedding, or image decoding
The memory footprint stays low because lopdf's lazy object loading only pulls referenced objects into memory. Large images beyond the sampled pages never get touched.
Code Examples
Rust API
use pdf_inspector::process_pdf_with_options;
use pdf_inspector::types::PdfType;
// Detect type without loading entire PDF
let path = "example.pdf";
let pdf_type = process_pdf_with_options(path, Default::default())
.map(|doc| doc.pdf_type)
.expect("Failed to open PDF");
match pdf_type {
PdfType::TextBased => println!("Extractable text document"),
PdfType::Scanned => println!("Requires OCR"),
PdfType::Mixed => println!("Hybrid content"),
PdfType::ImageBased => println!("Pure image pages"),
}
Command-Line Interface
# Quick detection with JSON output
detect-pdf --json example.pdf
# Example output
# {"type":"Scanned","pages":150,"confidence":"high"}
The CLI tool in [src/bin/detect_pdf.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs) exposes the same detection pipeline with configurable sampling depth.
Key Source Files
| File | Purpose |
|---|---|
[src/detector.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) |
Core detection orchestration; implements all three heuristics |
[src/text_quality.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs) |
Alphanumeric ratio calculation and text validation |
[src/bin/detect_pdf.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs) |
Standalone CLI for batch PDF classification |
[src/lib.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) |
Public API entry point process_pdf_with_options |
Summary
- pdf-inspector detects PDF type via three sequential heuristics: Structure Tree presence, alphanumeric text ratio, and tiled-scan image analysis
- Only a constant-size sample is read: PDF header, catalog, and first few pages—never the full document
- Classification completes in O(1) time: Performance does not degrade as page count increases
- The 50% alphanumeric threshold and 2M pixel tile threshold are the key decision boundaries between TextBased and Scanned classifications
Frequently Asked Questions
How accurate is the alphanumeric ratio threshold?
The 50% threshold in [src/text_quality.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs) effectively distinguishes true text documents from OCR overlays or image-heavy PDFs. Documents with legitimate text typically exceed 70–80% alphanumeric content, while scanned PDFs with hidden OCR garbage usually fall below 30%. The threshold provides a conservative cutoff with low false-negative rates for text detection.
Can the detection fail on hybrid documents?
Yes. PDFs containing substantial text sections interleaved with scanned images may receive Mixed classification when neither heuristic produces a definitive result. The detector in [src/detector.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) returns PdfType::Mixed specifically for edge cases where sampled pages contain conflicting signals—some with clean text, others with dominant image tiles.
Does pdf-inspector ever need to load the full PDF?
No. The detection pipeline is designed to avoid full loading. However, if downstream processing requests full text extraction or image rendering, the library will subsequently read additional objects. The initial classification itself never requires complete document parsing.
What PDF libraries does pdf-inspector depend on for this detection?
The detector uses lopdf for low-level PDF object parsing and the internal extractor pipeline for sampled text retrieval. These dependencies are lightweight compared to full rendering engines like PDFium or Poppler, enabling the minimal-read detection strategy that defines pdf-inspector's performance characteristics.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →