# ProcessMode Enum in pdf-inspector: How It Controls PDF Processing Depth

> Learn how the ProcessMode enum in pdf-inspector controls PDF processing depth. Choose between Full, Analyze, and DetectOnly modes for tailored extraction.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: internals
- Published: 2026-08-08

---

**The `ProcessMode` enum in firecrawl/pdf-inspector determines exactly how far the extraction pipeline runs, offering three variants—`Full`, `Analyze`, and `DetectOnly`—that control whether the system performs complete markdown conversion, structural analysis only, or simple PDF type detection.**

The `ProcessMode` enum serves as the primary execution gate for the firecrawl/pdf-inspector Rust library, allowing precise control over computational depth and output generation. Defined in [`src/process_mode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/process_mode.rs), this configuration option determines which pipeline stages execute—from lightweight classification to comprehensive text extraction and markdown formatting.

## What Is the ProcessMode Enum?

The `ProcessMode` enum is defined in [[`src/process_mode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/process_mode.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/process_mode.rs) and exposed through the public API in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) via the `PdfOptions` struct. It provides three distinct processing variants that cater to different use cases, from quick file classification to full document conversion.

Each variant maps to a specific subset of the extraction pipeline:

- **Full**: Executes the complete pipeline through markdown generation
- **Analyze**: Performs structural analysis without final markdown formatting
- **DetectOnly**: Runs only the PDF type detector with no text extraction

## How ProcessMode Affects the PDF Processing Pipeline

The enum directly controls which stages of the extraction pipeline execute. According to the implementation in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs), the system checks `options.mode` to conditionally skip specific processing steps.

### Full Mode: End-to-End Extraction

**`ProcessMode::Full`** runs the complete extraction pipeline, producing structured markdown or JSON output suitable for downstream consumption. This mode executes the full sequence:

1. PDF type detection (TextBased / Scanned / Mixed / ImageBased)
2. Text extraction
3. Layout analysis
4. Table detection
5. Markdown conversion

Use this mode when you need production-ready markdown output from your PDF documents.

### Analyze Mode: Structural Analysis Without Rendering

**`ProcessMode::Analyze`** performs all analysis steps but stops before markdown formatting. This variant is ideal for gathering statistics, debugging document structure, or extracting metadata without the overhead of final text rendering.

The pipeline executes through table detection, providing access to font statistics, column detection data, and table outlines, but omits the markdown conversion step. The CLI utility [`src/bin/detect_pdf.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs) utilizes this mode to produce detailed structural analysis reports.

### DetectOnly Mode: Minimal Classification

**`ProcessMode::DetectOnly`** executes only the PDF type detector, identifying whether a document is TextBased, Scanned, Mixed, or ImageBased without extracting any text content. This is the fastest option, useful for routing decisions or quick inventory scans where you only need to know the document type.

## Using ProcessMode in Your Code

The `ProcessMode` integrates with the `PdfOptions` struct available in the public API. You configure processing depth by setting the mode when building options:

```rust
use pdf_inspector::{PdfOptions, ProcessMode};

// Complete extraction with markdown output
let opts = PdfOptions::new().mode(ProcessMode::Full);
let result = process_pdf_with_options("report.pdf", opts)?;

// Quick type detection only
let opts = PdfOptions::new().mode(ProcessMode::DetectOnly);
let result = process_pdf_with_options("scanned_doc.pdf", opts)?;

// Structural analysis without markdown generation
let opts = PdfOptions::new().mode(ProcessMode::Analyze);
let result = process_pdf_with_options("complex_layout.pdf", opts)?;

```

### Command-Line Interface Usage

In the CLI binary [[`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs), the mode is selected via the `--mode` flag. The application parses this flag and branches execution based on the enum value:

```bash

# Full extraction (default behavior)

pdf2md --mode full input.pdf

# Analysis only

pdf2md --mode analyze input.pdf

# Type detection only

pdf2md --mode detect-only input.pdf

```

The detection utility [[`src/bin/detect_pdf.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs) specifically invokes the pipeline with `ProcessMode::Analyze` to generate detailed reports without markdown output.

## Pipeline Implementation Details

The extraction pipeline in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) orchestrates processing by checking the `mode` field on the options struct. For example, when `options.mode == ProcessMode::DetectOnly`, the system returns immediately after type classification, bypassing the text extraction and layout analysis modules entirely.

This conditional execution ensures that computational resources are allocated only to the stages required by the selected mode, making `DetectOnly` significantly faster than `Full` processing for large document batches.

## Summary

- **The `ProcessMode` enum** in [`src/process_mode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/process_mode.rs) controls pipeline depth with three variants: `Full`, `Analyze`, and `DetectOnly`
- **`ProcessMode::Full`** executes the complete pipeline from type detection through markdown conversion
- **`ProcessMode::Analyze`** provides structural analysis and table detection without final markdown formatting, useful for debugging
- **`ProcessMode::DetectOnly`** performs only PDF type classification (TextBased, Scanned, Mixed, ImageBased) with minimal overhead
- **Configuration** occurs via `PdfOptions::new().mode()` in Rust code or the `--mode` flag in the [`pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/pdf2md.rs) CLI
- **Implementation** checks in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) conditionally skip pipeline stages based on the selected mode

## Frequently Asked Questions

### How do I select a ProcessMode from the command line?

The [`pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/pdf2md.rs) binary accepts a `--mode` flag that maps directly to the enum variants. Use `--mode full` for complete extraction, `--mode analyze` for structural analysis without markdown output, or `--mode detect-only` for type classification only. The CLI parses this flag and constructs the appropriate `PdfOptions` configuration before invoking `process_pdf_with_options`.

### What is the performance difference between ProcessMode variants?

**`DetectOnly`** is the fastest option, executing only the PDF type detector without text extraction. **`Analyze`** adds text extraction, layout analysis, and table detection but skips the computationally expensive markdown formatting step. **`Full`** performs the complete pipeline including final markdown generation, making it the most resource-intensive but producing the most complete output.

### When should I use ProcessMode::Analyze instead of Full?

Use **`Analyze`** when you need structural metadata—such as font statistics, column positions, or table boundaries—without the final rendered markdown. This mode is ideal for debugging extraction issues, gathering document statistics, or building custom renderers that need raw layout data. The [`detect_pdf.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/detect_pdf.rs) utility uses this mode to produce detailed analysis reports.

### Can I change ProcessMode dynamically for different pages in the same PDF?

The current implementation in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) evaluates the `ProcessMode` once at the pipeline entry point and applies it consistently across the entire document. To process different pages with different modes, you would need to invoke `process_pdf_with_options` separately for each page range with distinct `PdfOptions` configurations, as the enum governs the global pipeline execution path.