PDF-Inspector Detection Configuration: Tuning min_text_ops_per_page and text_page_ratio_threshold

The DetectionConfig struct in pdf-inspector exposes three key parameters—strategy, min_text_ops_per_page, and text_page_ratio_threshold—that control how PDFs are classified as TextBased, Scanned, ImageBased, or Mixed.

pdf-inspector is a Rust library that classifies PDF documents by sampling content streams for text operators (Tj/TJ). The heuristics driving this classification live in the DetectionConfig struct defined in src/detector.rs at lines 69-78. Understanding these parameters lets you balance detection speed against accuracy and adapt the tool to your specific PDF collection.

Core Detection Parameters

The DetectionConfig struct provides fine-grained control over three fields:

Parameter Type Default Purpose
strategy ScanStrategy Sample(8) Chooses which pages are inspected
min_text_ops_per_page u32 3 Minimum text operators required for a page to count as text-based
text_page_ratio_threshold f32 0.6 Ratio of pages that must meet the text threshold for TextBased classification

strategy: Controlling Scan Scope

The strategy field determines detection speed versus precision. Available variants in ScanStrategy include:

  • EarlyExit — Stop scanning once classification is certain
  • Full — Inspect every page in the document
  • Sample(u32) — Sample up to N evenly-distributed pages (default: 8)
  • Pages(Vec<u32>) — Explicit list of pages to inspect

min_text_ops_per_page: Filtering Text Noise

Set min_text_ops_per_page to ignore pages with only stray text operators like page numbers or footers. The default value of 3 means pages with fewer than three Tj/TJ occurrences are treated as non-textual.

text_page_ratio_threshold: Setting Document-Level Tolerance

The text_page_ratio_threshold default of 0.6 requires 60% of examined pages to pass the min_text_ops_per_page check before the entire document is labeled TextBased. Lower this threshold for documents with heavy mixed content; raise it to demand consistently text-rich pages.

How Detection Uses These Parameters

The detection logic consumes DetectionConfig in the detect_from_document function (lines 182-199 of src/detector.rs). This implementation:

  1. Applies the selected strategy to choose which pages to sample
  2. Counts Tj/TJ operators per page
  3. Compares counts against min_text_ops_per_page
  4. Computes the ratio of passing pages against text_page_ratio_threshold
  5. Returns PdfType::TextBased, Scanned, ImageBased, or Mixed based on cumulative evidence

Code Examples

Default Configuration

Use detect_pdf_type for standard detection without tuning:

use pdf_inspector::detector::detect_pdf_type;

let result = detect_pdf_type("sample.pdf")?;
println!("Detected type: {:?}", result.pdf_type);

Custom Thresholds

Raise both thresholds for stricter TextBased classification:

use pdf_inspector::detector::{DetectionConfig, ScanStrategy, detect_pdf_type_with_config};

let mut cfg = DetectionConfig::default();
cfg.min_text_ops_per_page = 5;
cfg.text_page_ratio_threshold = 0.8;
cfg.strategy = ScanStrategy::Full;

let result = detect_pdf_type_with_config("sample.pdf", cfg)?;
println!("Detected type: {:?}", result.pdf_type);

Fast Sampling for Large PDFs

Use ScanStrategy::Sample to limit page inspections:

use pdf_inspector::detector::{DetectionConfig, ScanStrategy, detect_pdf_type_with_config};

let cfg = DetectionConfig {
    strategy: ScanStrategy::Sample(12),
    ..Default::default()
};

let result = detect_pdf_type_with_config("large.pdf", cfg)?;
println!("Detected type: {:?}", result.pdf_type);

Key Source Files

  • src/detector.rs — Contains PdfType, ScanStrategy, DetectionConfig, default values, and the detect_from_document algorithm
  • src/lib.rs — Public API with detect_pdf_type and detect_pdf_type_with_config
  • src/bin/detect_pdf.rs — CLI wrapper demonstrating command-line configuration overrides
  • tests/integration_tests.rs — Tests using custom DetectionConfig instances

Summary

  • DetectionConfig in src/detector.rs exposes three tunable parameters for pdf-inspector classification
  • strategy controls which pages are inspected, trading speed for accuracy
  • min_text_ops_per_page filters pages with insufficient textual content (default: 3)
  • text_page_ratio_threshold sets the required proportion of text-rich pages for TextBased classification (default: 0.6)
  • Use detect_pdf_type_with_config when you need non-default detection behavior

Frequently Asked Questions

How do I make pdf-inspector more strict about TextBased classification?

Increase both min_text_ops_per_page and text_page_ratio_threshold. For example, set min_text_ops_per_page to 5 or higher and text_page_ratio_threshold to 0.8 or 0.9. This reduces false positives where documents with sparse text get labeled as TextBased.

What is the fastest detection strategy for large PDFs?

Use ScanStrategy::Sample(n) with a modest n value (8-12). This samples evenly-distributed pages without reading the entire document. For even faster results, ScanStrategy::EarlyExit stops as soon as classification confidence is reached.

Can I inspect specific pages only?

Yes. Construct DetectionConfig with strategy: ScanStrategy::Pages(vec![1, 5, 10]) to examine only pages 1, 5, and 10. This is useful when you know which pages contain representative content.

Why does my PDF with some text get classified as Scanned or Mixed?

Your PDF likely falls below the default text_page_ratio_threshold of 0.6. If most pages have fewer than 3 text operators (the min_text_ops_per_page default), they don't count as text-based pages. Lower min_text_ops_per_page to 1 or 2, or reduce text_page_ratio_threshold to accommodate documents with sparse but meaningful text.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →