PDF Text Extraction Options in pdf-inspector: A Complete Guide to Rust-Powered PDF Parsing
pdf-inspector provides nine distinct PDF text extraction modes ranging from simple string output to coordinate-aware item extraction, region-based cropping, table-to-Markdown conversion, and hybrid OCR-ready pipelines, all exposed through a Rust API with Python and WebAssembly bindings.
The firecrawl/pdf-inspector library ships a flexible extraction framework that handles everything from basic text retrieval to complex layout analysis. Whether you need raw strings for indexing, positional data for UI rendering, or structured Markdown for content management, the toolkit offers granular control over the extraction process through its Rust-based engine.
Basic Text Extraction Modes
Full Document Plain Text
For simple use cases requiring only the readable characters in reading order, use the high-level convenience functions in src/lib.rs. The pdf_inspector::extract_text(path) function returns a complete String containing all concatenated text from the document.
let txt = pdf_inspector::extract_text("example.pdf")?;
println!("Full text:\n{}", txt);
This entry point lives at lines 51-55 in src/lib.rs and delegates to the core extraction logic in src/extractor/mod.rs.
In-Memory Extraction
When processing PDFs already loaded into memory (for example, from network requests), use pdf_inspector::extractor::extract_text_mem(buffer). This accepts a &[u8] buffer instead of a file path, eliminating disk I/O overhead.
let pdf_bytes = std::fs::read("example.pdf")?;
let txt = pdf_inspector::extractor::extract_text_mem(&pdf_bytes)?;
The implementation resides in src/extractor/mod.rs at lines 58-64.
Position-Aware and Selective Extraction
Text with Coordinates and Font Metadata
For downstream layout analysis, region cropping, or OCR-fallback decisions, extract TextItem structs containing x/y coordinates, font information, and bounding boxes. Call pdf_inspector::extract_text_with_positions(path) or its memory-based variant extract_text_with_positions_mem(buffer).
let items = pdf_inspector::extract_text_with_positions("example.pdf")?;
for it in items {
println!("Page {}: {} ({:.2},{:.2})", it.page, it.text, it.x, it.y);
}
This API is defined in src/extractor/mod.rs at lines 74-84.
Page-Specific Extraction
To save processing time when only specific pages are needed, use the page-filtered variants: extract_text_with_positions_pages(path, Some(&page_set)) or extract_text_with_positions_mem_pages(buffer, Some(&page_set)). These accept a HashSet<u32> of 1-indexed page numbers.
use std::collections::HashSet;
let mut pages = HashSet::new();
pages.insert(1);
pages.insert(3);
let items = pdf_inspector::extract_text_with_positions_pages("example.pdf", Some(&pages))?;
Password-Protected PDF Support
All position-aware extraction functions accept optional password parameters for encrypted documents. Use extract_text_with_positions_pages_with_password(path, …, Some(password)) to decrypt before extraction. This functionality appears in src/extractor/mod.rs at lines 95-105.
Region-Based and Table Extraction
Bounding Box Region Extraction
For extracting text only within specific coordinates—useful for form processing or invoice parsing—use pdf_inspector::extract_text_in_regions_mem. This returns a Vec<PageRegionResult> where each region reports whether needs_ocr is true, enabling hybrid OCR pipelines.
let regions = vec![
(0, vec![[50.0, 50.0, 300.0, 200.0]]), // page 0, bounding box
];
let region_res = pdf_inspector::extract_text_in_regions_mem(&pdf_bytes, ®ions)?;
for page_res in region_res {
for (i, r) in page_res.regions.iter().enumerate() {
println!("Region {}: {} (OCR needed: {})", i, r.text, r.needs_ocr);
}
}
This API is exposed in src/lib.rs at lines 66-71.
Table-to-Markdown Conversion
When working with tabular data, pdf_inspector::extract_tables_in_regions_mem accepts the same bounding-box input but returns Markdown pipe-tables when the detector succeeds. The RegionText result contains the formatted table in the text field with needs_ocr set to false for successful detections. Find this in src/lib.rs at lines 30-33.
Full Processing Pipelines
Complete PDF-to-Markdown Pipeline
The pdf_inspector::process_pdf(path) function runs the full detection → extraction → layout analysis → Markdown output pipeline. It handles multi-column layouts, tables, and headings automatically. For customization, use process_pdf_with_options(path, opts) with a PdfOptions builder.
let result = pdf_inspector::process_pdf("example.pdf")?;
println!("{}", result.markdown.unwrap());
Detect-Only Mode for Fast Metadata
Use pdf_inspector::detect_pdf(path) or PdfOptions::detect_only() via the builder for rapid metadata extraction without text processing. This mode returns document type, page count, and OCR requirements—ideal for routing decisions before committing to heavy extraction. Located in src/lib.rs at lines 73-78.
Analyze-Only Mode for Layout Intelligence
The ProcessMode::Analyze option runs detection, extraction, and layout analysis while skipping the final Markdown rendering step. This yields a PdfProcessResult containing pages_needing_ocr and layout structures without the conversion overhead. Configure this via PdfOptions::new().mode(ProcessMode::Analyze) as shown in src/lib.rs at lines 64-69.
Configuration with PdfOptions Builder
The PdfOptions builder (defined in src/lib.rs at lines 64-84) consolidates all extraction parameters into a single struct:
- Detection configuration (
DetectionConfig) - Markdown formatting profiles (
MarkdownOptions) - Page filtering (
HashSet<u32>) - Password decryption
- Processing mode (Full, Analyze, Detect)
let opts = pdf_inspector::PdfOptions::new()
.mode(pdf_inspector::ProcessMode::Full)
.markdown(pdf_inspector::MarkdownOptions::default()
.profile(pdf_inspector::MarkdownProfile::Compact))
.page_filter(Some([1, 2, 5].into_iter().collect()));
let result = pdf_inspector::process_pdf_with_options("example.pdf", opts)?;
Command-Line Interface
The pdf2md binary in src/bin/pdf2md.rs mirrors the library API and adds convenience flags: --json, --items-json, --raw, --detect-only, --analyze, --pages, --select-pages, and --password. This provides one-liner access to all extraction modes for shell scripting and debugging.
pdf2md example.pdf output.md --pages 1,3,5 --password secret
Internal Extraction Architecture
All extraction functions share a common five-stage pipeline implemented across the codebase:
- Validation:
validate_pdf_fileandvalidate_pdf_bytescheck header integrity insrc/lib.rs - Document Loading:
load_document_from_pathorload_document_from_memparses the PDF structure once - CMap Handling:
FontCMaps::from_doc(full) orfrom_doc_pages_fast(fast) builds Unicode mapping caches insrc/extractor/fonts.rs - Content-Stream Parsing:
extract_page_text_itemsinsrc/extractor/content_stream.rswalks PDF operators (Tj,TJ,Td,Tm) to emitTextItemstructs - Post-Processing:
text_utils::fix_letterspaced_itemsapplies Canva-style join thresholds, whileregion_overlaps_itemfilters by bounding boxes for region-based calls
Summary
- Plain-text extraction offers simple
Stringoutput viaextract_textandextract_text_meminsrc/lib.rsandsrc/extractor/mod.rs - Position-aware extraction provides coordinate and font metadata through
extract_text_with_positionsfunctions, with optional page filtering and password support - Region-based extraction enables bounding-box cropping and OCR-needs detection via
extract_text_in_regions_mem - Table extraction converts detected tables to Markdown pipe-tables using
extract_tables_in_regions_mem - Pipeline modes include full Markdown conversion (
process_pdf), fast metadata detection (detect_pdf), and layout analysis without rendering (ProcessMode::Analyze) - Configuration is centralized through the
PdfOptionsbuilder supporting page filters, passwords, and Markdown profiles - CLI access is provided by the
pdf2mdbinary supporting all major extraction flags
Frequently Asked Questions
How do I extract text from specific pages only?
Use the page-filtered variants extract_text_with_positions_pages or extract_text_with_positions_mem_pages, passing a HashSet<u32> containing the 1-indexed page numbers you need. For the full pipeline, use PdfOptions::new().page_filter(Some(page_set)) before calling process_pdf_with_options.
Can pdf-inspector handle encrypted PDFs?
Yes. All position-aware extraction functions accept optional password parameters. Use extract_text_with_positions_pages_with_password or set the password via the PdfOptions builder when using process_pdf_with_options. The decryption happens before any text extraction begins.
What is the difference between Detect and Analyze modes?
Detect-only mode (detect_pdf or PdfOptions::detect_only()) performs fast metadata extraction—returning document type, page count, and OCR requirements—without parsing content streams. Analyze mode (ProcessMode::Analyze) runs the full extraction and layout analysis but stops before Markdown rendering, returning structured data about text items, regions, and OCR needs without the final string conversion.
How do I integrate region-based extraction with OCR pipelines?
Call extract_text_in_regions_mem with your desired bounding boxes. Each returned PageRegionResult contains a needs_ocr boolean flag. When this flag is true, pipe the region's image data to your OCR engine; when false, use the extracted text directly. This hybrid approach optimizes processing by only running OCR on regions where text extraction failed or returned low-confidence results.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →