Extract Per-Page Markdown with Layout Metadata from PDFs in Memory Using `extract_pages_markdown_mem`

Use pdf_inspector::extract_pages_markdown_mem(buffer, pages) to convert PDF bytes into a PagesExtractionResult containing markdown strings, table/column detection flags, and OCR requirements for each page without writing files to disk.

The pdf-inspector crate provides a high-performance, memory-based pipeline for extracting per-page markdown with layout metadata from PDF documents. This guide covers the extract_pages_markdown_mem function in src/lib.rs, explaining its signature, return structure, internal mechanics, and practical usage patterns for Rust applications processing PDFs directly from byte buffers.


Function Signature and Parameters

The public API entry point is defined in src/lib.rs with this signature:

pub fn extract_pages_markdown_mem(
    buffer: &[u8],
    pages: Option<&[u32]>,
) -> Result<PagesExtractionResult, PdfError>

Parameters:

  • buffer – A byte slice containing the raw PDF data. Use std::fs::read for files, or pass HTTP response bodies directly without temporary files.
  • pages – An optional slice of 0-indexed page numbers:
    • None extracts every page in document order
    • Some(&[0, 2, 4]) extracts only pages 1, 3, and 5 (caller controls order)

Understanding the PagesExtractionResult Structure

The function returns a PagesExtractionResult struct (also defined in src/lib.rs) containing both per-page content and document-wide layout metadata:

Field Type Description
pages Vec<PageMarkdown> Per-page markdown, OCR flags, and extraction status
pages_with_tables Vec<u32> 1-indexed pages where table detection succeeded
pages_with_columns Vec<u32> 1-indexed pages with multi-column layout detected
pages_needing_ocr Vec<u32> 1-indexed pages requiring OCR processing
ocr_reasons_by_page HashMap<u32, Vec<OcrSignal>> Machine-readable OCR classification per page
is_complex bool Global flag set when any page contains tables or columns

Each PageMarkdown in the pages vector provides:

  • page: u32 – 0-indexed page number
  • markdown: String – Generated markdown content
  • needs_ocr: bool – Whether this page failed text extraction quality checks
  • ocr_reason: Option<OcrReason> – Detailed classification when OCR is needed

The Internal Extraction Pipeline

The extract_pages_markdown_mem function orchestrates multiple subsystems in pdf-inspector. Here's how the pipeline executes:

1. Document validation and loading (src/lib.rs)

  • validate_pdf_bytes confirms PDF header integrity
  • load_document_from_mem parses the document structure once

2. Font preparation (src/lib.rs)

  • FontCMaps::from_doc builds ToUnicode mapping caches for accurate text decoding

3. Global text extraction (src/extractor/mod.rs)

  • extract_positioned_text_for_document_analysis or extract_positioned_text_from_doc runs once on the entire document, producing:
    • all_items – raw TextItems with positioning data
    • all_rects and all_lines – geometric hints for layout analysis
    • gid_pages – pages containing GID-encoded fonts (common in scanned PDFs)

4. Quality analysis (src/text_quality.rs)

  • analyze_text_quality flags pages with garbled text or encoding anomalies

5. Noise reduction (src/lib.rs)

  • filter_markdown_page_numbers_with_removed_pages eliminates spurious page numbers that disrupt layout detection

6. Layout complexity detection (src/markdown/mod.rs)

  • markdown::chart_regions_by_page identifies structural regions
  • compute_layout_complexity_with_chart_regions flags tables and columns

7. Font statistics computation (src/markdown/analysis.rs)

  • calculate_font_stats_from_items builds a document-wide font-size histogram
  • Supplies base_font_size for consistent heading hierarchy in markdown output

8. Per-page rendering (src/markdown/mod.rs) For each requested page, the function:

  • Filters items, rects, and GID flags to the target page
  • Collects OCR signals via detector::page_ocr_signals (src/detector.rs)
  • Generates markdown via to_markdown_from_items_with_rects_and_lines
  • Applies decoding checks (is_cid_garbage, detect_encoding_issues)
  • Determines final needs_ocr status and ocr_reason

9. Result assembly All per-page structures aggregate into PagesExtractionResult with global metadata.

Because extraction runs once on the whole document, subsequent per-page operations are lightweight and benefit from shared font statistics and layout context.


Complete Code Examples

Extract All Pages from a File

use pdf_inspector::{extract_pages_markdown_mem, PdfError};

fn main() -> Result<(), PdfError> {
    // Load PDF bytes from filesystem or network
    let pdf_bytes = std::fs::read("example.pdf")?;

    // Extract markdown for every page
    let result = extract_pages_markdown_mem(&pdf_bytes, None)?;

    for page_md in result.pages {
        println!("--- Page {} ---", page_md.page + 1);
        
        if page_md.needs_ocr {
            println!("⚠️  OCR required: {:?}", page_md.ocr_reason);
        } else {
            println!("{}", page_md.markdown);
        }
    }

    // Inspect layout metadata
    println!("Tables detected on pages: {:?}", result.pages_with_tables);
    println!("Multi-column pages: {:?}", result.pages_with_columns);
    println!("Complex document: {}", result.is_complex);
    
    Ok(())
}

Extract Select Pages with OCR Routing

use pdf_inspector::{extract_pages_markdown_mem, PdfError};

fn process_select_pages(pdf_bytes: &[u8]) -> Result<(), PdfError> {
    // Request specific pages (0-indexed: pages 1, 3, 5)
    let selection = [0u32, 2, 4];
    
    let result = extract_pages_markdown_mem(pdf_bytes, Some(&selection))?;

    for page_md in result.pages {
        match page_md.needs_ocr {
            true => {
                // Queue for external OCR service
                eprintln!("Page {} → OCR pipeline", page_md.page + 1);
            }
            false => {
                // Use extracted markdown directly
                println!("Page {} content:\n{}", page_md.page + 1, page_md.markdown);
            }
        }
    }
    
    Ok(())
}

Hybrid OCR Pipeline with Result Filtering

use pdf_inspector::{extract_pages_markdown_mem, PdfError};

struct ProcessedDocument {
    clean_markdown: Vec<(u32, String)>,
    ocr_candidates: Vec<(u32, OcrReason)>,
    layout_flags: LayoutMetadata,
}

struct LayoutMetadata {
    has_tables: bool,
    has_columns: bool,
    complex_pages: Vec<u32>,
}

fn classify_document(pdf_bytes: &[u8]) -> Result<ProcessedDocument, PdfError> {
    let extraction = extract_pages_markdown_mem(pdf_bytes, None)?;

    let clean_markdown: Vec<_> = extraction
        .pages
        .iter()
        .filter(|p| !p.needs_ocr)
        .map(|p| (p.page, p.markdown.clone()))
        .collect();

    let ocr_candidates: Vec<_> = extraction
        .pages
        .iter()
        .filter(|p| p.needs_ocr)
        .filter_map(|p| p.ocr_reason.map(|r| (p.page, r)))
        .collect();

    Ok(ProcessedDocument {
        clean_markdown,
        ocr_candidates,
        layout_flags: LayoutMetadata {
            has_tables: !extraction.pages_with_tables.is_empty(),
            has_columns: !extraction.pages_with_columns.is_empty(),
            complex_pages: if extraction.is_complex {
                extraction.pages_with_tables.clone()
            } else {
                vec![]
            },
        },
    })
}

Key Source Files and Responsibilities

File Purpose Key Components
src/lib.rs Public API and orchestration extract_pages_markdown_mem, PagesExtractionResult, PageMarkdown
src/extractor/mod.rs Core text extraction extract_positioned_text_for_document_analysis, extract_positioned_text_from_doc
src/markdown/mod.rs Markdown generation to_markdown_from_items_with_rects_and_lines, layout complexity functions
src/detector.rs PDF type and OCR detection page_ocr_signals, template image detection
src/text_quality.rs Encoding quality analysis analyze_text_quality, CID garbage detection
src/markdown/analysis.rs Typography statistics calculate_font_stats_from_items for heading detection

These modules collectively implement the extraction pipeline, as implemented in firecrawl/pdf-inspector.


Summary

  • extract_pages_markdown_mem processes PDF bytes directly—no temporary files required
  • The 0-indexed pages parameter enables selective extraction; None processes all pages
  • Single-pass document analysis provides global font statistics and layout context for consistent per-page output
  • PagesExtractionResult combines markdown content with structural metadata: tables, columns, and OCR requirements
  • Layout detection uses geometric analysis (all_rects, all_lines) rather than heuristic parsing
  • OCR classification distinguishes between scanned images, GID-encoded fonts, and garbled text encoding

Frequently Asked Questions

What is the difference between 0-indexed and 1-indexed page numbers in the result?

The pages parameter and PageMarkdown.page field use 0-indexing (consistent with Rust conventions). All layout metadata fields (pages_with_tables, pages_with_columns, pages_needing_ocr) use 1-indexing for human-readable reporting. Convert between them by adding or subtracting 1.

How does pdf-inspector determine if a page needs OCR?

The needs_ocr flag combines multiple signals: GID-encoded fonts (via gid_pages from src/extractor/mod.rs), template image detection, garbled CID text (is_cid_garbage), and encoding anomalies (detect_encoding_issues in src/text_quality.rs). The ocr_reason field provides the specific classification.

Can I use this with PDFs from HTTP responses without saving to disk?

Yes. Pass the response body bytes directly: extract_pages_markdown_mem(&response_body, None). The function never writes to the filesystem—validation, parsing, and extraction all operate on the provided &[u8] buffer.

What causes is_complex to return true?

The is_complex boolean in PagesExtractionResult activates when any page contains detected tables (pages_with_tables non-empty) or multi-column layouts (pages_with_columns non-empty). This flag helps downstream systems allocate appropriate processing resources for structured document layouts.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →