# PDF-Inspector Detection Configuration: Tuning min_text_ops_per_page and text_page_ratio_threshold

> Tune pdf-inspector detection configuration using min_text_ops_per_page and text_page_ratio_threshold to accurately classify PDFs. Optimize your document analysis.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-06

---

**The `DetectionConfig` struct in pdf-inspector exposes three key parameters—`strategy`, `min_text_ops_per_page`, and `text_page_ratio_threshold`—that control how PDFs are classified as TextBased, Scanned, ImageBased, or Mixed.**

pdf-inspector is a Rust library that classifies PDF documents by sampling content streams for text operators (`Tj`/`TJ`). The heuristics driving this classification live in the **`DetectionConfig`** struct defined in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) at lines 69-78. Understanding these parameters lets you balance detection speed against accuracy and adapt the tool to your specific PDF collection.

## Core Detection Parameters

The `DetectionConfig` struct provides fine-grained control over three fields:

| Parameter | Type | Default | Purpose |
|-----------|------|---------|---------|
| `strategy` | `ScanStrategy` | `Sample(8)` | Chooses which pages are inspected |
| `min_text_ops_per_page` | `u32` | `3` | Minimum text operators required for a page to count as text-based |
| `text_page_ratio_threshold` | `f32` | `0.6` | Ratio of pages that must meet the text threshold for TextBased classification |

### `strategy`: Controlling Scan Scope

The **`strategy`** field determines detection speed versus precision. Available variants in `ScanStrategy` include:

- **`EarlyExit`** — Stop scanning once classification is certain
- **`Full`** — Inspect every page in the document
- **`Sample(u32)`** — Sample up to N evenly-distributed pages (default: 8)
- **`Pages(Vec<u32>)`** — Explicit list of pages to inspect

### `min_text_ops_per_page`: Filtering Text Noise

Set **`min_text_ops_per_page`** to ignore pages with only stray text operators like page numbers or footers. The default value of `3` means pages with fewer than three `Tj`/`TJ` occurrences are treated as non-textual.

### `text_page_ratio_threshold`: Setting Document-Level Tolerance

The **`text_page_ratio_threshold`** default of `0.6` requires 60% of examined pages to pass the `min_text_ops_per_page` check before the entire document is labeled *TextBased*. Lower this threshold for documents with heavy mixed content; raise it to demand consistently text-rich pages.

## How Detection Uses These Parameters

The detection logic consumes `DetectionConfig` in the `detect_from_document` function (lines 182-199 of [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)). This implementation:

1. Applies the selected `strategy` to choose which pages to sample
2. Counts `Tj`/`TJ` operators per page
3. Compares counts against `min_text_ops_per_page`
4. Computes the ratio of passing pages against `text_page_ratio_threshold`
5. Returns `PdfType::TextBased`, `Scanned`, `ImageBased`, or `Mixed` based on cumulative evidence

## Code Examples

### Default Configuration

Use `detect_pdf_type` for standard detection without tuning:

```rust
use pdf_inspector::detector::detect_pdf_type;

let result = detect_pdf_type("sample.pdf")?;
println!("Detected type: {:?}", result.pdf_type);

```

### Custom Thresholds

Raise both thresholds for stricter TextBased classification:

```rust
use pdf_inspector::detector::{DetectionConfig, ScanStrategy, detect_pdf_type_with_config};

let mut cfg = DetectionConfig::default();
cfg.min_text_ops_per_page = 5;
cfg.text_page_ratio_threshold = 0.8;
cfg.strategy = ScanStrategy::Full;

let result = detect_pdf_type_with_config("sample.pdf", cfg)?;
println!("Detected type: {:?}", result.pdf_type);

```

### Fast Sampling for Large PDFs

Use `ScanStrategy::Sample` to limit page inspections:

```rust
use pdf_inspector::detector::{DetectionConfig, ScanStrategy, detect_pdf_type_with_config};

let cfg = DetectionConfig {
    strategy: ScanStrategy::Sample(12),
    ..Default::default()
};

let result = detect_pdf_type_with_config("large.pdf", cfg)?;
println!("Detected type: {:?}", result.pdf_type);

```

## Key Source Files

- **[`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)** — Contains `PdfType`, `ScanStrategy`, `DetectionConfig`, default values, and the `detect_from_document` algorithm
- **[`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)** — Public API with `detect_pdf_type` and `detect_pdf_type_with_config`
- **[`src/bin/detect_pdf.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs)** — CLI wrapper demonstrating command-line configuration overrides
- **[`tests/integration_tests.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/tests/integration_tests.rs)** — Tests using custom `DetectionConfig` instances

## Summary

- **`DetectionConfig`** in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) exposes three tunable parameters for pdf-inspector classification
- **`strategy`** controls which pages are inspected, trading speed for accuracy
- **`min_text_ops_per_page`** filters pages with insufficient textual content (default: 3)
- **`text_page_ratio_threshold`** sets the required proportion of text-rich pages for TextBased classification (default: 0.6)
- Use **`detect_pdf_type_with_config`** when you need non-default detection behavior

## Frequently Asked Questions

### How do I make pdf-inspector more strict about TextBased classification?

Increase both `min_text_ops_per_page` and `text_page_ratio_threshold`. For example, set `min_text_ops_per_page` to `5` or higher and `text_page_ratio_threshold` to `0.8` or `0.9`. This reduces false positives where documents with sparse text get labeled as TextBased.

### What is the fastest detection strategy for large PDFs?

Use `ScanStrategy::Sample(n)` with a modest `n` value (8-12). This samples evenly-distributed pages without reading the entire document. For even faster results, `ScanStrategy::EarlyExit` stops as soon as classification confidence is reached.

### Can I inspect specific pages only?

Yes. Construct `DetectionConfig` with `strategy: ScanStrategy::Pages(vec![1, 5, 10])` to examine only pages 1, 5, and 10. This is useful when you know which pages contain representative content.

### Why does my PDF with some text get classified as Scanned or Mixed?

Your PDF likely falls below the default `text_page_ratio_threshold` of `0.6`. If most pages have fewer than 3 text operators (the `min_text_ops_per_page` default), they don't count as text-based pages. Lower `min_text_ops_per_page` to `1` or `2`, or reduce `text_page_ratio_threshold` to accommodate documents with sparse but meaningful text.