How to Configure Firecrawl PDF‑Inspector for Specific Use Cases: Detection, Speed, and Extraction Modes

Configure Firecrawl PDF‑Inspector by adjusting DetectionConfig for page‑sampling strategies and PdfOptions for pipeline modes, OCR thresholds, page filtering, and encrypted PDF handling.

The Firecrawl PDF‑Inspector is a Rust‑based modular pipeline that separates detection, extraction, and markdown generation into configurable stages. This article shows how to tailor the library for speed‑critical batch processing, precise document classification, selective page extraction, and encrypted file handling—using the actual source structures found in firecrawl/pdf‑inspector.

Core Configuration Structures

PDF‑Inspector's flexibility centers on two main types defined in the source code:

  • DetectionConfig (src/detector.rs, lines 69‑78) — Controls how PDFs are inspected for text‑based vs. scanned content
  • PdfOptions (src/lib.rs, lines 164‑180) — Builder for end‑to‑end run settings including mode, detection config, page filters, and passwords

The pipeline's execution mode is determined by ProcessMode (src/process_mode.rs, lines 7‑15), an enum with three variants:

pub enum ProcessMode {
    DetectOnly,  // Fast detection only, no extraction
    Full,        // Detection → extraction → markdown (default)
    Analyze,     // Detection + layout analysis without markdown
}

Configuring Detection Behavior with DetectionConfig

The DetectionConfig struct defines three critical thresholds:

pub struct DetectionConfig {
    pub strategy: ScanStrategy,           // Which pages to scan
    pub min_text_ops_per_page: u32,       // Minimum Tj/TJ operators for "text" page
    pub text_page_ratio_threshold: f32,   // Text pages / total pages for TextBased
}

ScanStrategy Variants for Speed vs. Accuracy

Strategy Implementation Best For
EarlyExit Stops at first non‑text page Quickly rejecting clearly scanned documents
Full Scans every page Precise Mixed/Scanned classification (legal contracts, mixed reports)
Sample(N) Samples N evenly‑distributed pages Large PDFs where speed matters
Pages(vec) Scans only specified 1‑indexed pages Known relevant pages (cover, index, appendix)

Default values (defined in src/detector.rs, lines 80‑89):

  • ScanStrategy::Sample(8)
  • min_text_ops_per_page: 3
  • text_page_ratio_threshold: 0.6

These defaults provide balanced accuracy for typical documents without scanning every page.

Configuring Pipeline Execution with PdfOptions

PdfOptions uses a fluent builder pattern to assemble complete run configurations. The struct fields (from src/lib.rs, lines 164‑180) include:

  • mode: ProcessMode — Pipeline stage selection
  • detection: DetectionConfig — Detection parameters
  • markdown: MarkdownOptions — Output formatting (Full mode only)
  • page_filter: Option<Vec<PageNum>> — Subset of pages to process
  • password: Option<String> — Decryption password (with custom Debug to prevent logging)

Builder methods are implemented in src/lib.rs at lines 30‑59.

Use Case: Default Fast Detection

For quick classification without extraction, use the detect_pdf convenience function:

use pdf_inspector::detect_pdf;

let result = detect_pdf("report.pdf")?;
println!("Detected type: {:?}", result.pdf_type);

This internally calls process_pdf_with_options with PdfOptions::detect_only(), which sets mode = ProcessMode::DetectOnly and default detection config (source: src/lib.rs, lines 73‑78).

Use Case: Speed‑Focused Sampling for Large PDFs

Reduce detection overhead by sampling fewer pages:

use pdf_inspector::{PdfOptions, DetectionConfig, ScanStrategy};

let config = DetectionConfig {
    strategy: ScanStrategy::Sample(4), // Only 4 pages examined
    ..DetectionConfig::default()
};

let opts = PdfOptions::new().detection(config);
let result = pdf_inspector::process_pdf_with_options("huge.pdf", opts)?;

println!("Pages sampled: {}", result.pages_sampled);

Performance impact: Detection becomes O(1) regardless of total page count. The four sampled pages are distributed across the document (first, last, and intermediate pages) to improve representativeness.

Use Case: Precise Mixed/Scanned Classification

Force complete page analysis for documents requiring accurate type detection:

use pdf_inspector::{PdfOptions, DetectionConfig, ScanStrategy, ProcessMode};

let config = DetectionConfig {
    strategy: ScanStrategy::Full,        // Every page scanned
    min_text_ops_per_page: 5,            // Stricter text threshold
    text_page_ratio_threshold: 0.7,      // Higher bar for TextBased
    ..DetectionConfig::default()
};

let opts = PdfOptions::new()
    .mode(ProcessMode::Full)
    .detection(config);

let result = pdf_inspector::process_pdf_with_options("contract.pdf", opts)?;
println!("PDF type: {:?}", result.pdf_type);

When to use: Legal documents, academic papers with scanned figures, or reports mixing generated and raster content. The Full strategy prevents early‑exit misclassification when a single non‑text page appears early in the document.

Use Case: Selective Page Processing

Extract only specific pages without parsing the entire file:

use pdf_inspector::PdfOptions;

let opts = PdfOptions::new()
    .pages([1, 10, 20]); // 1‑indexed page numbers

let result = pdf_inspector::process_pdf_with_options("book.pdf", opts)?;
println!("Extracted {} pages", result.pages.len());

The pages method accepts any IntoIterator<Item = PageNum> and propagates the filter through the extraction pipeline (src/extractor/mod.rs).

Use Case: Encrypted PDF Handling

Provide decryption passwords safely:

use pdf_inspector::{PdfOptions, ProcessMode};

let opts = PdfOptions::new()
    .password("my‑pw")
    .mode(ProcessMode::Full);

let result = pdf_inspector::process_pdf_with_options("encrypted.pdf", opts)?;
println!("Decrypted and extracted {} pages", result.pages.len());

Security note: The password field uses a custom Debug implementation that masks the actual value, preventing accidental exposure in logs.

Complete Configuration Example

Combine multiple options for production workflows:

use pdf_inspector::{
    ProcessMode, PdfOptions, DetectionConfig, ScanStrategy, MarkdownOptions,
};

let detection_cfg = DetectionConfig {
    strategy: ScanStrategy::Sample(6),   // Faster for 300‑page reports
    min_text_ops_per_page: 4,
    text_page_ratio_threshold: 0.65,
    ..Default::default()
};

let markdown_cfg = MarkdownOptions::default(); // Customize in src/markdown/mod.rs

let opts = PdfOptions::new()
    .mode(ProcessMode::Full)
    .detection(detection_cfg)
    .markdown(markdown_cfg)
    .pages(1..=10);                       // First ten pages only

let result = pdf_inspector::process_pdf_with_options("annual_report.pdf", opts)?;
println!("Completed: {} markdown blocks", result.markdown.len());

Key Source Files Reference

File Purpose Lines
src/detector.rs DetectionConfig, ScanStrategy, detection algorithm 69‑89
src/lib.rs PdfOptions builder, public API entry points 30‑59, 73‑78, 164‑180
src/process_mode.rs ProcessMode enum definition 7‑15
src/markdown/mod.rs MarkdownOptions and output formatting Module root
src/extractor/mod.rs Extraction orchestration respecting page filters Module root
src/tables/mod.rs Table detection strategies Module root

Summary

  • Use ProcessMode::DetectOnly for fastest classification without content extraction
  • Adjust ScanStrategy to trade speed for accuracy: Sample(N) for large files, Full for precise classification, Pages(vec) for selective analysis
  • Tune OCR thresholds via min_text_ops_per_page and text_page_ratio_threshold to match your document types
  • Apply pages() and password() methods on PdfOptions for targeted processing and secure decryption
  • Prefer PdfOptions::new() builder over raw struct initialization for forward compatibility

Frequently Asked Questions

What is the fastest way to classify PDF type without extracting content?

Use pdf_inspector::detect_pdf("file.pdf") or explicitly configure PdfOptions::new().mode(ProcessMode::DetectOnly). This skips extraction and markdown generation entirely, running only the detection algorithm with default DetectionConfig.

How do I process a 1000‑page PDF without scanning every page?

Set DetectionConfig { strategy: ScanStrategy::Sample(4), ..Default::default() } and pass it to PdfOptions::detection(). The detector examines only four evenly‑spaced pages regardless of total page count, making runtime effectively constant.

When should I use ProcessMode::Analyze instead of Full?

Use Analyze when you need structural information—table boundaries, column positions, reading order—without the overhead of markdown serialization. This mode runs detection plus layout analysis but skips the final markdown generation step, saving CPU for downstream custom formatters.

How do I prevent passwords from appearing in logs?

The PdfOptions struct implements a custom Debug trait that masks the password field. When you call .password("secret"), the value is stored for decryption but displays as redacted in any {:?} formatting or logging output.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →