# How to Use `extract_pages_markdown_mem` for Per-Page Markdown Extraction

> Learn how to use extract_pages_markdown_mem to get per-page Markdown extraction from PDF bytes with OCR routing metadata. Get detailed page results efficiently.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-08

---

**The `extract_pages_markdown_mem` function processes PDF bytes through a staged pipeline to return a `PagesExtractionResult` containing individual `PageMarkdown` structs with OCR routing metadata for each requested page.**

The `firecrawl/pdf-inspector` repository provides a Rust-native solution for extracting structured Markdown from PDF documents. The `extract_pages_markdown_mem` function serves as the low-level entry point when you need granular control over which pages to process and want to identify which pages require OCR fallback.

## Understanding the Extraction Pipeline

The implementation follows a seven-stage pipeline that separates document-wide analysis from per-page generation. This design ensures that font statistics and layout complexity calculations happen once, making single-page extraction cheap and deterministic.

### Document Validation and Loading

The process begins by validating the input bytes and loading the PDF structure. The function calls `validate_pdf_bytes` to check the byte slice integrity, then `load_document_from_mem` creates a `lopdf::Document` instance. These utilities are defined in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) at lines 431 and 462 respectively.

### Global Font and Layout Analysis

Before processing individual pages, the entire document is scanned to gather statistics independent of specific page requests. The `extract_positioned_text_from_doc` function in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) (line 146) computes:

- **Font statistics** including the most common font size for consistent heading detection
- **Layout complexity** indicators for tables, columns, and chart regions

This document-wide pass ensures that per-page Markdown generation uses consistent formatting thresholds regardless of which subset of pages you request.

### OCR Signal Detection

Three independent checks determine whether a page must be routed to OCR. The `analyze_text_quality` function in [`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs) (line 12) identifies garbled or low-quality text, while `detector::page_ocr_signals` in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) (line 1800) flags pages containing large background images or vector-drawn text. These same signals power the higher-level `process_pdf_mem` API, ensuring consistent routing across the codebase.

### Page Selection Logic

The function accepts an optional `pages` parameter that controls which pages to process:

- If `pages` is `None`, every page is processed using 0-based indexing
- If a slice is provided (e.g., `&[0u32, 2, 5]`), only those specific indices are visited while preserving the caller-specified order
- Out-of-range indices return empty Markdown with `needs_ocr = true`

The selection logic is implemented in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) between lines 1080 and 1114.

### Per-Page Markdown Generation

For each selected page, the function extracts items belonging to that page and builds a `MarkdownOptions` object reusing the document-wide most common font size. It then calls `markdown::to_markdown_from_items_with_rects_and_lines` in [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) (line 210) to produce the raw Markdown text.

### Finalizing OCR Decisions

A page is marked `needs_ocr` when any of the following conditions are met, as implemented in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) (lines 6000-6030):

- **OCR signals** from the detection phase are true
- **Empty Markdown** generation indicates missing extractable text
- **GID-encoded fonts** (`has_gid`) prevent proper text extraction
- **Garbage text detection** (`is_garbage_text`) indicates corrupted encoding

The reasons are collected in a `BTreeMap<u32, Vec<OcrReason>>` and exposed as `ocr_reasons_by_page` in the final result.

## Code Implementation Examples

### Rust Implementation

Call `extract_pages_markdown_mem` with a byte slice and optional page indices to receive a `PagesExtractionResult`:

```rust
use pdf_inspector::{extract_pages_markdown_mem, PagesExtractionResult};

fn main() -> Result<(), pdf_inspector::PdfError> {
    // Load PDF bytes from any source (file, network, etc.)
    let pdf_bytes = std::fs::read("example.pdf")?;
    
    // Request specific pages using 0-based indexing, or None for all pages
    let pages_to_extract = Some(&[0u32, 2, 5][..]);
    
    // Execute extraction
    let result: PagesExtractionResult = extract_pages_markdown_mem(&pdf_bytes, pages_to_extract)?;
    
    // Process each page
    for page_md in result.pages {
        println!("--- Page {} ---", page_md.page + 1);
        if page_md.needs_ocr {
            println!("Requires OCR - defer to external pipeline");
        } else {
            println!("{}", page_md.markdown);
        }
    }
    
    println!("Pages requiring OCR: {:?}", result.pages_needing_ocr);
    Ok(())
}

```

**Key implementation details:**
- The `pages` slice uses **0-based indexing**; the returned `PageMarkdown.page` preserves this index
- When `needs_ocr` is `true`, the `markdown` field is intentionally empty
- Document-wide analysis runs once regardless of how many pages you request

### Python via PyO3 Bindings

The Python wrapper `extract_pages_markdown_bytes` in [`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs) (lines 595-640) exposes the same functionality:

```python
import pdf_inspector

# Load PDF as bytes

with open("example.pdf", "rb") as f:
    data = f.read()

# Request pages 1 and 3 (0-based indices 0 and 2)

pages = [0, 2]

result = pdf_inspector.extract_pages_markdown_bytes(data, pages)

for page in result.pages:
    print(f"--- Page {page.page + 1} ---")
    if page.needs_ocr:
        print("Needs OCR handling")
    else:
        print(page.markdown)

print("OCR flagged pages:", result.pages_needing_ocr)

```

The Python API maintains the same 0-based indexing as the Rust implementation and returns a dataclass matching the Rust `PagesExtractionResult` structure.

## Return Structure and Metadata

The `PagesExtractionResult` struct defined in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) (line 389) contains:

- **`pages: Vec<PageMarkdown>`** — Each entry holds the original 0-indexed page number, the Markdown string (empty if OCR required), and the `needs_ocr` boolean
- **`pages_needing_ocr: Vec<u32>`** — 1-indexed list of pages requiring OCR processing
- **`ocr_reasons_by_page: Vec<PageOcrReasons>`** — Human-readable explanations for OCR routing decisions
- **`is_complex: bool`** — True when any page includes tables or columnar layouts indicating complex document structure

## Summary

- **`extract_pages_markdown_mem`** in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) provides low-level access to per-page Markdown extraction with OCR routing
- The function performs **document-wide analysis once** (fonts, layout) then processes only requested pages, making it efficient for partial extraction
- **OCR routing** depends on text quality analysis, background images, vector text, GID fonts, and garbage text detection
- The API uses **0-based indexing** for input pages and returns both Markdown content and metadata about OCR requirements
- Python bindings via `extract_pages_markdown_bytes` in [`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs) offer identical functionality for Python workflows

## Frequently Asked Questions

### What is the difference between `extract_pages_markdown_mem` and `process_pdf_mem`?

`extract_pages_markdown_mem` is the low-level primitive that returns raw Markdown and OCR metadata without performing OCR itself. `process_pdf_mem` is the higher-level orchestrator that calls `extract_pages_markdown_mem` and automatically routes flagged pages through the OCR pipeline. Use the former when you want to handle OCR yourself or mix extraction methods; use the latter for fully automated processing.

### Why is the page index 0-based in Rust but 1-based in the OCR results?

The input `pages` slice and the `PageMarkdown.page` field use 0-based indexing to align with PDF internal page structures and Rust iterator conventions. However, `pages_needing_ocr` uses 1-based indexing for human readability when reporting which physical pages need attention. Always pass 0-based indices to the extraction function.

### How does the function handle pages with tables or complex layouts?

The `is_complex` boolean in `PagesExtractionResult` indicates whether the document contains tables or columnar layouts detected during the global analysis phase in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs). While the Markdown converter attempts to preserve structure, complex layouts often trigger OCR recommendations because native PDF text extraction may lose tabular relationships or reading order.

### Can I extract a single page without loading the entire document's metadata?

While you can request a single page via `Some(&[5])`, the function always performs document-wide font and layout analysis first to ensure consistent heading detection and formatting. This guarantees that page 5 uses the same font-size thresholds as it would in a full-document extraction, but it does mean some overhead is unavoidable for single-page extraction.