PDF-Inspector Detection Configuration: Tuning min_text_ops_per_page and text_page_ratio_threshold
The DetectionConfig struct in pdf-inspector exposes three key parameters—strategy, min_text_ops_per_page, and text_page_ratio_threshold—that control how PDFs are classified as TextBased, Scanned, ImageBased, or Mixed.
pdf-inspector is a Rust library that classifies PDF documents by sampling content streams for text operators (Tj/TJ). The heuristics driving this classification live in the DetectionConfig struct defined in src/detector.rs at lines 69-78. Understanding these parameters lets you balance detection speed against accuracy and adapt the tool to your specific PDF collection.
Core Detection Parameters
The DetectionConfig struct provides fine-grained control over three fields:
| Parameter | Type | Default | Purpose |
|---|---|---|---|
strategy |
ScanStrategy |
Sample(8) |
Chooses which pages are inspected |
min_text_ops_per_page |
u32 |
3 |
Minimum text operators required for a page to count as text-based |
text_page_ratio_threshold |
f32 |
0.6 |
Ratio of pages that must meet the text threshold for TextBased classification |
strategy: Controlling Scan Scope
The strategy field determines detection speed versus precision. Available variants in ScanStrategy include:
EarlyExit— Stop scanning once classification is certainFull— Inspect every page in the documentSample(u32)— Sample up to N evenly-distributed pages (default: 8)Pages(Vec<u32>)— Explicit list of pages to inspect
min_text_ops_per_page: Filtering Text Noise
Set min_text_ops_per_page to ignore pages with only stray text operators like page numbers or footers. The default value of 3 means pages with fewer than three Tj/TJ occurrences are treated as non-textual.
text_page_ratio_threshold: Setting Document-Level Tolerance
The text_page_ratio_threshold default of 0.6 requires 60% of examined pages to pass the min_text_ops_per_page check before the entire document is labeled TextBased. Lower this threshold for documents with heavy mixed content; raise it to demand consistently text-rich pages.
How Detection Uses These Parameters
The detection logic consumes DetectionConfig in the detect_from_document function (lines 182-199 of src/detector.rs). This implementation:
- Applies the selected
strategyto choose which pages to sample - Counts
Tj/TJoperators per page - Compares counts against
min_text_ops_per_page - Computes the ratio of passing pages against
text_page_ratio_threshold - Returns
PdfType::TextBased,Scanned,ImageBased, orMixedbased on cumulative evidence
Code Examples
Default Configuration
Use detect_pdf_type for standard detection without tuning:
use pdf_inspector::detector::detect_pdf_type;
let result = detect_pdf_type("sample.pdf")?;
println!("Detected type: {:?}", result.pdf_type);
Custom Thresholds
Raise both thresholds for stricter TextBased classification:
use pdf_inspector::detector::{DetectionConfig, ScanStrategy, detect_pdf_type_with_config};
let mut cfg = DetectionConfig::default();
cfg.min_text_ops_per_page = 5;
cfg.text_page_ratio_threshold = 0.8;
cfg.strategy = ScanStrategy::Full;
let result = detect_pdf_type_with_config("sample.pdf", cfg)?;
println!("Detected type: {:?}", result.pdf_type);
Fast Sampling for Large PDFs
Use ScanStrategy::Sample to limit page inspections:
use pdf_inspector::detector::{DetectionConfig, ScanStrategy, detect_pdf_type_with_config};
let cfg = DetectionConfig {
strategy: ScanStrategy::Sample(12),
..Default::default()
};
let result = detect_pdf_type_with_config("large.pdf", cfg)?;
println!("Detected type: {:?}", result.pdf_type);
Key Source Files
src/detector.rs— ContainsPdfType,ScanStrategy,DetectionConfig, default values, and thedetect_from_documentalgorithmsrc/lib.rs— Public API withdetect_pdf_typeanddetect_pdf_type_with_configsrc/bin/detect_pdf.rs— CLI wrapper demonstrating command-line configuration overridestests/integration_tests.rs— Tests using customDetectionConfiginstances
Summary
DetectionConfiginsrc/detector.rsexposes three tunable parameters for pdf-inspector classificationstrategycontrols which pages are inspected, trading speed for accuracymin_text_ops_per_pagefilters pages with insufficient textual content (default: 3)text_page_ratio_thresholdsets the required proportion of text-rich pages for TextBased classification (default: 0.6)- Use
detect_pdf_type_with_configwhen you need non-default detection behavior
Frequently Asked Questions
How do I make pdf-inspector more strict about TextBased classification?
Increase both min_text_ops_per_page and text_page_ratio_threshold. For example, set min_text_ops_per_page to 5 or higher and text_page_ratio_threshold to 0.8 or 0.9. This reduces false positives where documents with sparse text get labeled as TextBased.
What is the fastest detection strategy for large PDFs?
Use ScanStrategy::Sample(n) with a modest n value (8-12). This samples evenly-distributed pages without reading the entire document. For even faster results, ScanStrategy::EarlyExit stops as soon as classification confidence is reached.
Can I inspect specific pages only?
Yes. Construct DetectionConfig with strategy: ScanStrategy::Pages(vec![1, 5, 10]) to examine only pages 1, 5, and 10. This is useful when you know which pages contain representative content.
Why does my PDF with some text get classified as Scanned or Mixed?
Your PDF likely falls below the default text_page_ratio_threshold of 0.6. If most pages have fewer than 3 text operators (the min_text_ops_per_page default), they don't count as text-based pages. Lower min_text_ops_per_page to 1 or 2, or reduce text_page_ratio_threshold to accommodate documents with sparse but meaningful text.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →