How to Use `extract_pages_markdown_mem` for Per-Page Markdown Extraction
The extract_pages_markdown_mem function processes PDF bytes through a staged pipeline to return a PagesExtractionResult containing individual PageMarkdown structs with OCR routing metadata for each requested page.
The firecrawl/pdf-inspector repository provides a Rust-native solution for extracting structured Markdown from PDF documents. The extract_pages_markdown_mem function serves as the low-level entry point when you need granular control over which pages to process and want to identify which pages require OCR fallback.
Understanding the Extraction Pipeline
The implementation follows a seven-stage pipeline that separates document-wide analysis from per-page generation. This design ensures that font statistics and layout complexity calculations happen once, making single-page extraction cheap and deterministic.
Document Validation and Loading
The process begins by validating the input bytes and loading the PDF structure. The function calls validate_pdf_bytes to check the byte slice integrity, then load_document_from_mem creates a lopdf::Document instance. These utilities are defined in src/lib.rs at lines 431 and 462 respectively.
Global Font and Layout Analysis
Before processing individual pages, the entire document is scanned to gather statistics independent of specific page requests. The extract_positioned_text_from_doc function in src/extractor/mod.rs (line 146) computes:
- Font statistics including the most common font size for consistent heading detection
- Layout complexity indicators for tables, columns, and chart regions
This document-wide pass ensures that per-page Markdown generation uses consistent formatting thresholds regardless of which subset of pages you request.
OCR Signal Detection
Three independent checks determine whether a page must be routed to OCR. The analyze_text_quality function in src/text_quality.rs (line 12) identifies garbled or low-quality text, while detector::page_ocr_signals in src/detector.rs (line 1800) flags pages containing large background images or vector-drawn text. These same signals power the higher-level process_pdf_mem API, ensuring consistent routing across the codebase.
Page Selection Logic
The function accepts an optional pages parameter that controls which pages to process:
- If
pagesisNone, every page is processed using 0-based indexing - If a slice is provided (e.g.,
&[0u32, 2, 5]), only those specific indices are visited while preserving the caller-specified order - Out-of-range indices return empty Markdown with
needs_ocr = true
The selection logic is implemented in src/lib.rs between lines 1080 and 1114.
Per-Page Markdown Generation
For each selected page, the function extracts items belonging to that page and builds a MarkdownOptions object reusing the document-wide most common font size. It then calls markdown::to_markdown_from_items_with_rects_and_lines in src/markdown/convert.rs (line 210) to produce the raw Markdown text.
Finalizing OCR Decisions
A page is marked needs_ocr when any of the following conditions are met, as implemented in src/lib.rs (lines 6000-6030):
- OCR signals from the detection phase are true
- Empty Markdown generation indicates missing extractable text
- GID-encoded fonts (
has_gid) prevent proper text extraction - Garbage text detection (
is_garbage_text) indicates corrupted encoding
The reasons are collected in a BTreeMap<u32, Vec<OcrReason>> and exposed as ocr_reasons_by_page in the final result.
Code Implementation Examples
Rust Implementation
Call extract_pages_markdown_mem with a byte slice and optional page indices to receive a PagesExtractionResult:
use pdf_inspector::{extract_pages_markdown_mem, PagesExtractionResult};
fn main() -> Result<(), pdf_inspector::PdfError> {
// Load PDF bytes from any source (file, network, etc.)
let pdf_bytes = std::fs::read("example.pdf")?;
// Request specific pages using 0-based indexing, or None for all pages
let pages_to_extract = Some(&[0u32, 2, 5][..]);
// Execute extraction
let result: PagesExtractionResult = extract_pages_markdown_mem(&pdf_bytes, pages_to_extract)?;
// Process each page
for page_md in result.pages {
println!("--- Page {} ---", page_md.page + 1);
if page_md.needs_ocr {
println!("Requires OCR - defer to external pipeline");
} else {
println!("{}", page_md.markdown);
}
}
println!("Pages requiring OCR: {:?}", result.pages_needing_ocr);
Ok(())
}
Key implementation details:
- The
pagesslice uses 0-based indexing; the returnedPageMarkdown.pagepreserves this index - When
needs_ocristrue, themarkdownfield is intentionally empty - Document-wide analysis runs once regardless of how many pages you request
Python via PyO3 Bindings
The Python wrapper extract_pages_markdown_bytes in src/python.rs (lines 595-640) exposes the same functionality:
import pdf_inspector
# Load PDF as bytes
with open("example.pdf", "rb") as f:
data = f.read()
# Request pages 1 and 3 (0-based indices 0 and 2)
pages = [0, 2]
result = pdf_inspector.extract_pages_markdown_bytes(data, pages)
for page in result.pages:
print(f"--- Page {page.page + 1} ---")
if page.needs_ocr:
print("Needs OCR handling")
else:
print(page.markdown)
print("OCR flagged pages:", result.pages_needing_ocr)
The Python API maintains the same 0-based indexing as the Rust implementation and returns a dataclass matching the Rust PagesExtractionResult structure.
Return Structure and Metadata
The PagesExtractionResult struct defined in src/lib.rs (line 389) contains:
pages: Vec<PageMarkdown>— Each entry holds the original 0-indexed page number, the Markdown string (empty if OCR required), and theneeds_ocrbooleanpages_needing_ocr: Vec<u32>— 1-indexed list of pages requiring OCR processingocr_reasons_by_page: Vec<PageOcrReasons>— Human-readable explanations for OCR routing decisionsis_complex: bool— True when any page includes tables or columnar layouts indicating complex document structure
Summary
extract_pages_markdown_meminsrc/lib.rsprovides low-level access to per-page Markdown extraction with OCR routing- The function performs document-wide analysis once (fonts, layout) then processes only requested pages, making it efficient for partial extraction
- OCR routing depends on text quality analysis, background images, vector text, GID fonts, and garbage text detection
- The API uses 0-based indexing for input pages and returns both Markdown content and metadata about OCR requirements
- Python bindings via
extract_pages_markdown_bytesinsrc/python.rsoffer identical functionality for Python workflows
Frequently Asked Questions
What is the difference between extract_pages_markdown_mem and process_pdf_mem?
extract_pages_markdown_mem is the low-level primitive that returns raw Markdown and OCR metadata without performing OCR itself. process_pdf_mem is the higher-level orchestrator that calls extract_pages_markdown_mem and automatically routes flagged pages through the OCR pipeline. Use the former when you want to handle OCR yourself or mix extraction methods; use the latter for fully automated processing.
Why is the page index 0-based in Rust but 1-based in the OCR results?
The input pages slice and the PageMarkdown.page field use 0-based indexing to align with PDF internal page structures and Rust iterator conventions. However, pages_needing_ocr uses 1-based indexing for human readability when reporting which physical pages need attention. Always pass 0-based indices to the extraction function.
How does the function handle pages with tables or complex layouts?
The is_complex boolean in PagesExtractionResult indicates whether the document contains tables or columnar layouts detected during the global analysis phase in src/extractor/mod.rs. While the Markdown converter attempts to preserve structure, complex layouts often trigger OCR recommendations because native PDF text extraction may lose tabular relationships or reading order.
Can I extract a single page without loading the entire document's metadata?
While you can request a single page via Some(&[5]), the function always performs document-wide font and layout analysis first to ensure consistent heading detection and formatting. This guarantees that page 5 uses the same font-size thresholds as it would in a full-document extraction, but it does mean some overhead is unavoidable for single-page extraction.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →