PDF Text Extraction Options in pdf-inspector: A Complete Guide to Rust-Powered PDF Parsing

pdf-inspector provides nine distinct PDF text extraction modes ranging from simple string output to coordinate-aware item extraction, region-based cropping, table-to-Markdown conversion, and hybrid OCR-ready pipelines, all exposed through a Rust API with Python and WebAssembly bindings.

The firecrawl/pdf-inspector library ships a flexible extraction framework that handles everything from basic text retrieval to complex layout analysis. Whether you need raw strings for indexing, positional data for UI rendering, or structured Markdown for content management, the toolkit offers granular control over the extraction process through its Rust-based engine.

Basic Text Extraction Modes

Full Document Plain Text

For simple use cases requiring only the readable characters in reading order, use the high-level convenience functions in src/lib.rs. The pdf_inspector::extract_text(path) function returns a complete String containing all concatenated text from the document.

let txt = pdf_inspector::extract_text("example.pdf")?;
println!("Full text:\n{}", txt);

This entry point lives at lines 51-55 in src/lib.rs and delegates to the core extraction logic in src/extractor/mod.rs.

In-Memory Extraction

When processing PDFs already loaded into memory (for example, from network requests), use pdf_inspector::extractor::extract_text_mem(buffer). This accepts a &[u8] buffer instead of a file path, eliminating disk I/O overhead.

let pdf_bytes = std::fs::read("example.pdf")?;
let txt = pdf_inspector::extractor::extract_text_mem(&pdf_bytes)?;

The implementation resides in src/extractor/mod.rs at lines 58-64.

Position-Aware and Selective Extraction

Text with Coordinates and Font Metadata

For downstream layout analysis, region cropping, or OCR-fallback decisions, extract TextItem structs containing x/y coordinates, font information, and bounding boxes. Call pdf_inspector::extract_text_with_positions(path) or its memory-based variant extract_text_with_positions_mem(buffer).

let items = pdf_inspector::extract_text_with_positions("example.pdf")?;
for it in items {
    println!("Page {}: {} ({:.2},{:.2})", it.page, it.text, it.x, it.y);
}

This API is defined in src/extractor/mod.rs at lines 74-84.

Page-Specific Extraction

To save processing time when only specific pages are needed, use the page-filtered variants: extract_text_with_positions_pages(path, Some(&page_set)) or extract_text_with_positions_mem_pages(buffer, Some(&page_set)). These accept a HashSet<u32> of 1-indexed page numbers.

use std::collections::HashSet;
let mut pages = HashSet::new();
pages.insert(1);
pages.insert(3);
let items = pdf_inspector::extract_text_with_positions_pages("example.pdf", Some(&pages))?;

Password-Protected PDF Support

All position-aware extraction functions accept optional password parameters for encrypted documents. Use extract_text_with_positions_pages_with_password(path, …, Some(password)) to decrypt before extraction. This functionality appears in src/extractor/mod.rs at lines 95-105.

Region-Based and Table Extraction

Bounding Box Region Extraction

For extracting text only within specific coordinates—useful for form processing or invoice parsing—use pdf_inspector::extract_text_in_regions_mem. This returns a Vec<PageRegionResult> where each region reports whether needs_ocr is true, enabling hybrid OCR pipelines.

let regions = vec![
    (0, vec![[50.0, 50.0, 300.0, 200.0]]), // page 0, bounding box
];
let region_res = pdf_inspector::extract_text_in_regions_mem(&pdf_bytes, &regions)?;
for page_res in region_res {
    for (i, r) in page_res.regions.iter().enumerate() {
        println!("Region {}: {} (OCR needed: {})", i, r.text, r.needs_ocr);
    }
}

This API is exposed in src/lib.rs at lines 66-71.

Table-to-Markdown Conversion

When working with tabular data, pdf_inspector::extract_tables_in_regions_mem accepts the same bounding-box input but returns Markdown pipe-tables when the detector succeeds. The RegionText result contains the formatted table in the text field with needs_ocr set to false for successful detections. Find this in src/lib.rs at lines 30-33.

Full Processing Pipelines

Complete PDF-to-Markdown Pipeline

The pdf_inspector::process_pdf(path) function runs the full detection → extraction → layout analysis → Markdown output pipeline. It handles multi-column layouts, tables, and headings automatically. For customization, use process_pdf_with_options(path, opts) with a PdfOptions builder.

let result = pdf_inspector::process_pdf("example.pdf")?;
println!("{}", result.markdown.unwrap());

Detect-Only Mode for Fast Metadata

Use pdf_inspector::detect_pdf(path) or PdfOptions::detect_only() via the builder for rapid metadata extraction without text processing. This mode returns document type, page count, and OCR requirements—ideal for routing decisions before committing to heavy extraction. Located in src/lib.rs at lines 73-78.

Analyze-Only Mode for Layout Intelligence

The ProcessMode::Analyze option runs detection, extraction, and layout analysis while skipping the final Markdown rendering step. This yields a PdfProcessResult containing pages_needing_ocr and layout structures without the conversion overhead. Configure this via PdfOptions::new().mode(ProcessMode::Analyze) as shown in src/lib.rs at lines 64-69.

Configuration with PdfOptions Builder

The PdfOptions builder (defined in src/lib.rs at lines 64-84) consolidates all extraction parameters into a single struct:

  • Detection configuration (DetectionConfig)
  • Markdown formatting profiles (MarkdownOptions)
  • Page filtering (HashSet<u32>)
  • Password decryption
  • Processing mode (Full, Analyze, Detect)
let opts = pdf_inspector::PdfOptions::new()
    .mode(pdf_inspector::ProcessMode::Full)
    .markdown(pdf_inspector::MarkdownOptions::default()
        .profile(pdf_inspector::MarkdownProfile::Compact))
    .page_filter(Some([1, 2, 5].into_iter().collect()));
let result = pdf_inspector::process_pdf_with_options("example.pdf", opts)?;

Command-Line Interface

The pdf2md binary in src/bin/pdf2md.rs mirrors the library API and adds convenience flags: --json, --items-json, --raw, --detect-only, --analyze, --pages, --select-pages, and --password. This provides one-liner access to all extraction modes for shell scripting and debugging.

pdf2md example.pdf output.md --pages 1,3,5 --password secret

Internal Extraction Architecture

All extraction functions share a common five-stage pipeline implemented across the codebase:

  1. Validation: validate_pdf_file and validate_pdf_bytes check header integrity in src/lib.rs
  2. Document Loading: load_document_from_path or load_document_from_mem parses the PDF structure once
  3. CMap Handling: FontCMaps::from_doc (full) or from_doc_pages_fast (fast) builds Unicode mapping caches in src/extractor/fonts.rs
  4. Content-Stream Parsing: extract_page_text_items in src/extractor/content_stream.rs walks PDF operators (Tj, TJ, Td, Tm) to emit TextItem structs
  5. Post-Processing: text_utils::fix_letterspaced_items applies Canva-style join thresholds, while region_overlaps_item filters by bounding boxes for region-based calls

Summary

  • Plain-text extraction offers simple String output via extract_text and extract_text_mem in src/lib.rs and src/extractor/mod.rs
  • Position-aware extraction provides coordinate and font metadata through extract_text_with_positions functions, with optional page filtering and password support
  • Region-based extraction enables bounding-box cropping and OCR-needs detection via extract_text_in_regions_mem
  • Table extraction converts detected tables to Markdown pipe-tables using extract_tables_in_regions_mem
  • Pipeline modes include full Markdown conversion (process_pdf), fast metadata detection (detect_pdf), and layout analysis without rendering (ProcessMode::Analyze)
  • Configuration is centralized through the PdfOptions builder supporting page filters, passwords, and Markdown profiles
  • CLI access is provided by the pdf2md binary supporting all major extraction flags

Frequently Asked Questions

How do I extract text from specific pages only?

Use the page-filtered variants extract_text_with_positions_pages or extract_text_with_positions_mem_pages, passing a HashSet<u32> containing the 1-indexed page numbers you need. For the full pipeline, use PdfOptions::new().page_filter(Some(page_set)) before calling process_pdf_with_options.

Can pdf-inspector handle encrypted PDFs?

Yes. All position-aware extraction functions accept optional password parameters. Use extract_text_with_positions_pages_with_password or set the password via the PdfOptions builder when using process_pdf_with_options. The decryption happens before any text extraction begins.

What is the difference between Detect and Analyze modes?

Detect-only mode (detect_pdf or PdfOptions::detect_only()) performs fast metadata extraction—returning document type, page count, and OCR requirements—without parsing content streams. Analyze mode (ProcessMode::Analyze) runs the full extraction and layout analysis but stops before Markdown rendering, returning structured data about text items, regions, and OCR needs without the final string conversion.

How do I integrate region-based extraction with OCR pipelines?

Call extract_text_in_regions_mem with your desired bounding boxes. Each returned PageRegionResult contains a needs_ocr boolean flag. When this flag is true, pipe the region's image data to your OCR engine; when false, use the extracted text directly. This hybrid approach optimizes processing by only running OCR on regions where text extraction failed or returned low-confidence results.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →