# PDF Text Extraction Options in pdf-inspector: A Complete Guide to Rust-Powered PDF Parsing

> Explore nine PDF text extraction modes in pdf-inspector, a Rust-powered PDF parser. Get coordinate-aware, region-based, table-to-Markdown, and OCR-ready extraction via Python & WASM.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-04

---

**pdf-inspector provides nine distinct PDF text extraction modes ranging from simple string output to coordinate-aware item extraction, region-based cropping, table-to-Markdown conversion, and hybrid OCR-ready pipelines, all exposed through a Rust API with Python and WebAssembly bindings.**

The `firecrawl/pdf-inspector` library ships a flexible extraction framework that handles everything from basic text retrieval to complex layout analysis. Whether you need raw strings for indexing, positional data for UI rendering, or structured Markdown for content management, the toolkit offers granular control over the extraction process through its Rust-based engine.

## Basic Text Extraction Modes

### Full Document Plain Text

For simple use cases requiring only the readable characters in reading order, use the high-level convenience functions in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs). The `pdf_inspector::extract_text(path)` function returns a complete `String` containing all concatenated text from the document.

```rust
let txt = pdf_inspector::extract_text("example.pdf")?;
println!("Full text:\n{}", txt);

```

This entry point lives at lines 51-55 in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) and delegates to the core extraction logic in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs).

### In-Memory Extraction

When processing PDFs already loaded into memory (for example, from network requests), use `pdf_inspector::extractor::extract_text_mem(buffer)`. This accepts a `&[u8]` buffer instead of a file path, eliminating disk I/O overhead.

```rust
let pdf_bytes = std::fs::read("example.pdf")?;
let txt = pdf_inspector::extractor::extract_text_mem(&pdf_bytes)?;

```

The implementation resides in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) at lines 58-64.

## Position-Aware and Selective Extraction

### Text with Coordinates and Font Metadata

For downstream layout analysis, region cropping, or OCR-fallback decisions, extract `TextItem` structs containing **x/y coordinates**, **font information**, and **bounding boxes**. Call `pdf_inspector::extract_text_with_positions(path)` or its memory-based variant `extract_text_with_positions_mem(buffer)`.

```rust
let items = pdf_inspector::extract_text_with_positions("example.pdf")?;
for it in items {
    println!("Page {}: {} ({:.2},{:.2})", it.page, it.text, it.x, it.y);
}

```

This API is defined in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) at lines 74-84.

### Page-Specific Extraction

To save processing time when only specific pages are needed, use the page-filtered variants: `extract_text_with_positions_pages(path, Some(&page_set))` or `extract_text_with_positions_mem_pages(buffer, Some(&page_set))`. These accept a `HashSet<u32>` of 1-indexed page numbers.

```rust
use std::collections::HashSet;
let mut pages = HashSet::new();
pages.insert(1);
pages.insert(3);
let items = pdf_inspector::extract_text_with_positions_pages("example.pdf", Some(&pages))?;

```

### Password-Protected PDF Support

All position-aware extraction functions accept optional password parameters for encrypted documents. Use `extract_text_with_positions_pages_with_password(path, …, Some(password))` to decrypt before extraction. This functionality appears in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) at lines 95-105.

## Region-Based and Table Extraction

### Bounding Box Region Extraction

For extracting text only within specific coordinates—useful for form processing or invoice parsing—use `pdf_inspector::extract_text_in_regions_mem`. This returns a `Vec<PageRegionResult>` where each region reports whether `needs_ocr` is true, enabling hybrid OCR pipelines.

```rust
let regions = vec![
    (0, vec![[50.0, 50.0, 300.0, 200.0]]), // page 0, bounding box
];
let region_res = pdf_inspector::extract_text_in_regions_mem(&pdf_bytes, &regions)?;
for page_res in region_res {
    for (i, r) in page_res.regions.iter().enumerate() {
        println!("Region {}: {} (OCR needed: {})", i, r.text, r.needs_ocr);
    }
}

```

This API is exposed in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) at lines 66-71.

### Table-to-Markdown Conversion

When working with tabular data, `pdf_inspector::extract_tables_in_regions_mem` accepts the same bounding-box input but returns Markdown pipe-tables when the detector succeeds. The `RegionText` result contains the formatted table in the `text` field with `needs_ocr` set to false for successful detections. Find this in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) at lines 30-33.

## Full Processing Pipelines

### Complete PDF-to-Markdown Pipeline

The `pdf_inspector::process_pdf(path)` function runs the full detection → extraction → layout analysis → Markdown output pipeline. It handles multi-column layouts, tables, and headings automatically. For customization, use `process_pdf_with_options(path, opts)` with a `PdfOptions` builder.

```rust
let result = pdf_inspector::process_pdf("example.pdf")?;
println!("{}", result.markdown.unwrap());

```

### Detect-Only Mode for Fast Metadata

Use `pdf_inspector::detect_pdf(path)` or `PdfOptions::detect_only()` via the builder for rapid metadata extraction without text processing. This mode returns document type, page count, and OCR requirements—ideal for routing decisions before committing to heavy extraction. Located in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) at lines 73-78.

### Analyze-Only Mode for Layout Intelligence

The `ProcessMode::Analyze` option runs detection, extraction, and layout analysis while skipping the final Markdown rendering step. This yields a `PdfProcessResult` containing `pages_needing_ocr` and layout structures without the conversion overhead. Configure this via `PdfOptions::new().mode(ProcessMode::Analyze)` as shown in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) at lines 64-69.

## Configuration with PdfOptions Builder

The **`PdfOptions`** builder (defined in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) at lines 64-84) consolidates all extraction parameters into a single struct:

- **Detection configuration** (`DetectionConfig`)
- **Markdown formatting profiles** (`MarkdownOptions`)
- **Page filtering** (`HashSet<u32>`)
- **Password decryption**
- **Processing mode** (Full, Analyze, Detect)

```rust
let opts = pdf_inspector::PdfOptions::new()
    .mode(pdf_inspector::ProcessMode::Full)
    .markdown(pdf_inspector::MarkdownOptions::default()
        .profile(pdf_inspector::MarkdownProfile::Compact))
    .page_filter(Some([1, 2, 5].into_iter().collect()));
let result = pdf_inspector::process_pdf_with_options("example.pdf", opts)?;

```

## Command-Line Interface

The `pdf2md` binary in [`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) mirrors the library API and adds convenience flags: `--json`, `--items-json`, `--raw`, `--detect-only`, `--analyze`, `--pages`, `--select-pages`, and `--password`. This provides one-liner access to all extraction modes for shell scripting and debugging.

```bash
pdf2md example.pdf output.md --pages 1,3,5 --password secret

```

## Internal Extraction Architecture

All extraction functions share a common five-stage pipeline implemented across the codebase:

1. **Validation**: `validate_pdf_file` and `validate_pdf_bytes` check header integrity in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)
2. **Document Loading**: `load_document_from_path` or `load_document_from_mem` parses the PDF structure once
3. **CMap Handling**: `FontCMaps::from_doc` (full) or `from_doc_pages_fast` (fast) builds Unicode mapping caches in [`src/extractor/fonts.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/fonts.rs)
4. **Content-Stream Parsing**: `extract_page_text_items` in [`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs) walks PDF operators (`Tj`, `TJ`, `Td`, `Tm`) to emit `TextItem` structs
5. **Post-Processing**: `text_utils::fix_letterspaced_items` applies Canva-style join thresholds, while `region_overlaps_item` filters by bounding boxes for region-based calls

## Summary

- **Plain-text extraction** offers simple `String` output via `extract_text` and `extract_text_mem` in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) and [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs)
- **Position-aware extraction** provides coordinate and font metadata through `extract_text_with_positions` functions, with optional page filtering and password support
- **Region-based extraction** enables bounding-box cropping and OCR-needs detection via `extract_text_in_regions_mem`
- **Table extraction** converts detected tables to Markdown pipe-tables using `extract_tables_in_regions_mem`
- **Pipeline modes** include full Markdown conversion (`process_pdf`), fast metadata detection (`detect_pdf`), and layout analysis without rendering (`ProcessMode::Analyze`)
- **Configuration** is centralized through the `PdfOptions` builder supporting page filters, passwords, and Markdown profiles
- **CLI access** is provided by the `pdf2md` binary supporting all major extraction flags

## Frequently Asked Questions

### How do I extract text from specific pages only?

Use the page-filtered variants `extract_text_with_positions_pages` or `extract_text_with_positions_mem_pages`, passing a `HashSet<u32>` containing the 1-indexed page numbers you need. For the full pipeline, use `PdfOptions::new().page_filter(Some(page_set))` before calling `process_pdf_with_options`.

### Can pdf-inspector handle encrypted PDFs?

Yes. All position-aware extraction functions accept optional password parameters. Use `extract_text_with_positions_pages_with_password` or set the password via the `PdfOptions` builder when using `process_pdf_with_options`. The decryption happens before any text extraction begins.

### What is the difference between Detect and Analyze modes?

**Detect-only mode** (`detect_pdf` or `PdfOptions::detect_only()`) performs fast metadata extraction—returning document type, page count, and OCR requirements—without parsing content streams. **Analyze mode** (`ProcessMode::Analyze`) runs the full extraction and layout analysis but stops before Markdown rendering, returning structured data about text items, regions, and OCR needs without the final string conversion.

### How do I integrate region-based extraction with OCR pipelines?

Call `extract_text_in_regions_mem` with your desired bounding boxes. Each returned `PageRegionResult` contains a `needs_ocr` boolean flag. When this flag is true, pipe the region's image data to your OCR engine; when false, use the extracted text directly. This hybrid approach optimizes processing by only running OCR on regions where text extraction failed or returned low-confidence results.