# Extract Per-Page Markdown with Layout Metadata from PDFs in Memory Using `extract_pages_markdown_mem`

> Extract per page markdown from PDFs in memory using `extract_pages_markdown_mem`. Get layout metadata, table/column flags, and OCR needs without saving files.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-06

---

**Use `pdf_inspector::extract_pages_markdown_mem(buffer, pages)` to convert PDF bytes into a `PagesExtractionResult` containing markdown strings, table/column detection flags, and OCR requirements for each page without writing files to disk.**

The `pdf-inspector` crate provides a high-performance, memory-based pipeline for **extracting per-page markdown with layout metadata** from PDF documents. This guide covers the `extract_pages_markdown_mem` function in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), explaining its signature, return structure, internal mechanics, and practical usage patterns for Rust applications processing PDFs directly from byte buffers.

---

## Function Signature and Parameters

The public API entry point is defined in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) with this signature:

```rust
pub fn extract_pages_markdown_mem(
    buffer: &[u8],
    pages: Option<&[u32]>,
) -> Result<PagesExtractionResult, PdfError>

```

**Parameters:**

- **`buffer`** – A byte slice containing the raw PDF data. Use `std::fs::read` for files, or pass HTTP response bodies directly without temporary files.
- **`pages`** – An optional slice of **0-indexed** page numbers:
  - `None` extracts every page in document order
  - `Some(&[0, 2, 4])` extracts only pages 1, 3, and 5 (caller controls order)

---

## Understanding the `PagesExtractionResult` Structure

The function returns a `PagesExtractionResult` struct (also defined in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)) containing both per-page content and document-wide layout metadata:

| Field | Type | Description |
|-------|------|-------------|
| `pages` | `Vec<PageMarkdown>` | Per-page markdown, OCR flags, and extraction status |
| `pages_with_tables` | `Vec<u32>` | **1-indexed** pages where table detection succeeded |
| `pages_with_columns` | `Vec<u32>` | **1-indexed** pages with multi-column layout detected |
| `pages_needing_ocr` | `Vec<u32>` | **1-indexed** pages requiring OCR processing |
| `ocr_reasons_by_page` | `HashMap<u32, Vec<OcrSignal>>` | Machine-readable OCR classification per page |
| `is_complex` | `bool` | Global flag set when any page contains tables or columns |

Each `PageMarkdown` in the `pages` vector provides:

- `page: u32` – 0-indexed page number
- `markdown: String` – Generated markdown content
- `needs_ocr: bool` – Whether this page failed text extraction quality checks
- `ocr_reason: Option<OcrReason>` – Detailed classification when OCR is needed

---

## The Internal Extraction Pipeline

The `extract_pages_markdown_mem` function orchestrates multiple subsystems in `pdf-inspector`. Here's how the pipeline executes:

**1. Document validation and loading** ([`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs))
- `validate_pdf_bytes` confirms PDF header integrity
- `load_document_from_mem` parses the document structure once

**2. Font preparation** ([`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs))
- `FontCMaps::from_doc` builds ToUnicode mapping caches for accurate text decoding

**3. Global text extraction** ([`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs))
- `extract_positioned_text_for_document_analysis` or `extract_positioned_text_from_doc` runs **once** on the entire document, producing:
  - `all_items` – raw `TextItem`s with positioning data
  - `all_rects` and `all_lines` – geometric hints for layout analysis
  - `gid_pages` – pages containing GID-encoded fonts (common in scanned PDFs)

**4. Quality analysis** ([`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs))
- `analyze_text_quality` flags pages with garbled text or encoding anomalies

**5. Noise reduction** ([`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs))
- `filter_markdown_page_numbers_with_removed_pages` eliminates spurious page numbers that disrupt layout detection

**6. Layout complexity detection** ([`src/markdown/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/mod.rs))
- `markdown::chart_regions_by_page` identifies structural regions
- `compute_layout_complexity_with_chart_regions` flags tables and columns

**7. Font statistics computation** ([`src/markdown/analysis.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/analysis.rs))
- `calculate_font_stats_from_items` builds a document-wide font-size histogram
- Supplies `base_font_size` for consistent heading hierarchy in markdown output

**8. Per-page rendering** ([`src/markdown/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/mod.rs))
For each requested page, the function:
- Filters items, rects, and GID flags to the target page
- Collects OCR signals via `detector::page_ocr_signals` ([`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs))
- Generates markdown via `to_markdown_from_items_with_rects_and_lines`
- Applies decoding checks (`is_cid_garbage`, `detect_encoding_issues`)
- Determines final `needs_ocr` status and `ocr_reason`

**9. Result assembly**
All per-page structures aggregate into `PagesExtractionResult` with global metadata.

Because extraction runs **once** on the whole document, subsequent per-page operations are lightweight and benefit from shared font statistics and layout context.

---

## Complete Code Examples

### Extract All Pages from a File

```rust
use pdf_inspector::{extract_pages_markdown_mem, PdfError};

fn main() -> Result<(), PdfError> {
    // Load PDF bytes from filesystem or network
    let pdf_bytes = std::fs::read("example.pdf")?;

    // Extract markdown for every page
    let result = extract_pages_markdown_mem(&pdf_bytes, None)?;

    for page_md in result.pages {
        println!("--- Page {} ---", page_md.page + 1);
        
        if page_md.needs_ocr {
            println!("⚠️  OCR required: {:?}", page_md.ocr_reason);
        } else {
            println!("{}", page_md.markdown);
        }
    }

    // Inspect layout metadata
    println!("Tables detected on pages: {:?}", result.pages_with_tables);
    println!("Multi-column pages: {:?}", result.pages_with_columns);
    println!("Complex document: {}", result.is_complex);
    
    Ok(())
}

```

### Extract Select Pages with OCR Routing

```rust
use pdf_inspector::{extract_pages_markdown_mem, PdfError};

fn process_select_pages(pdf_bytes: &[u8]) -> Result<(), PdfError> {
    // Request specific pages (0-indexed: pages 1, 3, 5)
    let selection = [0u32, 2, 4];
    
    let result = extract_pages_markdown_mem(pdf_bytes, Some(&selection))?;

    for page_md in result.pages {
        match page_md.needs_ocr {
            true => {
                // Queue for external OCR service
                eprintln!("Page {} → OCR pipeline", page_md.page + 1);
            }
            false => {
                // Use extracted markdown directly
                println!("Page {} content:\n{}", page_md.page + 1, page_md.markdown);
            }
        }
    }
    
    Ok(())
}

```

### Hybrid OCR Pipeline with Result Filtering

```rust
use pdf_inspector::{extract_pages_markdown_mem, PdfError};

struct ProcessedDocument {
    clean_markdown: Vec<(u32, String)>,
    ocr_candidates: Vec<(u32, OcrReason)>,
    layout_flags: LayoutMetadata,
}

struct LayoutMetadata {
    has_tables: bool,
    has_columns: bool,
    complex_pages: Vec<u32>,
}

fn classify_document(pdf_bytes: &[u8]) -> Result<ProcessedDocument, PdfError> {
    let extraction = extract_pages_markdown_mem(pdf_bytes, None)?;

    let clean_markdown: Vec<_> = extraction
        .pages
        .iter()
        .filter(|p| !p.needs_ocr)
        .map(|p| (p.page, p.markdown.clone()))
        .collect();

    let ocr_candidates: Vec<_> = extraction
        .pages
        .iter()
        .filter(|p| p.needs_ocr)
        .filter_map(|p| p.ocr_reason.map(|r| (p.page, r)))
        .collect();

    Ok(ProcessedDocument {
        clean_markdown,
        ocr_candidates,
        layout_flags: LayoutMetadata {
            has_tables: !extraction.pages_with_tables.is_empty(),
            has_columns: !extraction.pages_with_columns.is_empty(),
            complex_pages: if extraction.is_complex {
                extraction.pages_with_tables.clone()
            } else {
                vec![]
            },
        },
    })
}

```

---

## Key Source Files and Responsibilities

| File | Purpose | Key Components |
|------|---------|----------------|
| [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) | Public API and orchestration | `extract_pages_markdown_mem`, `PagesExtractionResult`, `PageMarkdown` |
| [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) | Core text extraction | `extract_positioned_text_for_document_analysis`, `extract_positioned_text_from_doc` |
| [`src/markdown/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/mod.rs) | Markdown generation | `to_markdown_from_items_with_rects_and_lines`, layout complexity functions |
| [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) | PDF type and OCR detection | `page_ocr_signals`, template image detection |
| [`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs) | Encoding quality analysis | `analyze_text_quality`, CID garbage detection |
| [`src/markdown/analysis.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/analysis.rs) | Typography statistics | `calculate_font_stats_from_items` for heading detection |

These modules collectively implement the extraction pipeline, as implemented in `firecrawl/pdf-inspector`.

---

## Summary

- **`extract_pages_markdown_mem`** processes PDF bytes directly—no temporary files required
- The **0-indexed `pages` parameter** enables selective extraction; `None` processes all pages
- **Single-pass document analysis** provides global font statistics and layout context for consistent per-page output
- **`PagesExtractionResult`** combines markdown content with structural metadata: tables, columns, and OCR requirements
- **Layout detection** uses geometric analysis (`all_rects`, `all_lines`) rather than heuristic parsing
- **OCR classification** distinguishes between scanned images, GID-encoded fonts, and garbled text encoding

---

## Frequently Asked Questions

### What is the difference between 0-indexed and 1-indexed page numbers in the result?

The `pages` parameter and `PageMarkdown.page` field use **0-indexing** (consistent with Rust conventions). All layout metadata fields (`pages_with_tables`, `pages_with_columns`, `pages_needing_ocr`) use **1-indexing** for human-readable reporting. Convert between them by adding or subtracting 1.

### How does `pdf-inspector` determine if a page needs OCR?

The `needs_ocr` flag combines multiple signals: GID-encoded fonts (via `gid_pages` from [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs)), template image detection, garbled CID text (`is_cid_garbage`), and encoding anomalies (`detect_encoding_issues` in [`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs)). The `ocr_reason` field provides the specific classification.

### Can I use this with PDFs from HTTP responses without saving to disk?

Yes. Pass the response body bytes directly: `extract_pages_markdown_mem(&response_body, None)`. The function never writes to the filesystem—validation, parsing, and extraction all operate on the provided `&[u8]` buffer.

### What causes `is_complex` to return true?

The `is_complex` boolean in `PagesExtractionResult` activates when **any** page contains detected tables (`pages_with_tables` non-empty) or multi-column layouts (`pages_with_columns` non-empty). This flag helps downstream systems allocate appropriate processing resources for structured document layouts.