Extract Per-Page Markdown with Layout Metadata from PDFs in Memory Using `extract_pages_markdown_mem`
Use pdf_inspector::extract_pages_markdown_mem(buffer, pages) to convert PDF bytes into a PagesExtractionResult containing markdown strings, table/column detection flags, and OCR requirements for each page without writing files to disk.
The pdf-inspector crate provides a high-performance, memory-based pipeline for extracting per-page markdown with layout metadata from PDF documents. This guide covers the extract_pages_markdown_mem function in src/lib.rs, explaining its signature, return structure, internal mechanics, and practical usage patterns for Rust applications processing PDFs directly from byte buffers.
Function Signature and Parameters
The public API entry point is defined in src/lib.rs with this signature:
pub fn extract_pages_markdown_mem(
buffer: &[u8],
pages: Option<&[u32]>,
) -> Result<PagesExtractionResult, PdfError>
Parameters:
buffer– A byte slice containing the raw PDF data. Usestd::fs::readfor files, or pass HTTP response bodies directly without temporary files.pages– An optional slice of 0-indexed page numbers:Noneextracts every page in document orderSome(&[0, 2, 4])extracts only pages 1, 3, and 5 (caller controls order)
Understanding the PagesExtractionResult Structure
The function returns a PagesExtractionResult struct (also defined in src/lib.rs) containing both per-page content and document-wide layout metadata:
| Field | Type | Description |
|---|---|---|
pages |
Vec<PageMarkdown> |
Per-page markdown, OCR flags, and extraction status |
pages_with_tables |
Vec<u32> |
1-indexed pages where table detection succeeded |
pages_with_columns |
Vec<u32> |
1-indexed pages with multi-column layout detected |
pages_needing_ocr |
Vec<u32> |
1-indexed pages requiring OCR processing |
ocr_reasons_by_page |
HashMap<u32, Vec<OcrSignal>> |
Machine-readable OCR classification per page |
is_complex |
bool |
Global flag set when any page contains tables or columns |
Each PageMarkdown in the pages vector provides:
page: u32– 0-indexed page numbermarkdown: String– Generated markdown contentneeds_ocr: bool– Whether this page failed text extraction quality checksocr_reason: Option<OcrReason>– Detailed classification when OCR is needed
The Internal Extraction Pipeline
The extract_pages_markdown_mem function orchestrates multiple subsystems in pdf-inspector. Here's how the pipeline executes:
1. Document validation and loading (src/lib.rs)
validate_pdf_bytesconfirms PDF header integrityload_document_from_memparses the document structure once
2. Font preparation (src/lib.rs)
FontCMaps::from_docbuilds ToUnicode mapping caches for accurate text decoding
3. Global text extraction (src/extractor/mod.rs)
extract_positioned_text_for_document_analysisorextract_positioned_text_from_docruns once on the entire document, producing:all_items– rawTextItems with positioning dataall_rectsandall_lines– geometric hints for layout analysisgid_pages– pages containing GID-encoded fonts (common in scanned PDFs)
4. Quality analysis (src/text_quality.rs)
analyze_text_qualityflags pages with garbled text or encoding anomalies
5. Noise reduction (src/lib.rs)
filter_markdown_page_numbers_with_removed_pageseliminates spurious page numbers that disrupt layout detection
6. Layout complexity detection (src/markdown/mod.rs)
markdown::chart_regions_by_pageidentifies structural regionscompute_layout_complexity_with_chart_regionsflags tables and columns
7. Font statistics computation (src/markdown/analysis.rs)
calculate_font_stats_from_itemsbuilds a document-wide font-size histogram- Supplies
base_font_sizefor consistent heading hierarchy in markdown output
8. Per-page rendering (src/markdown/mod.rs)
For each requested page, the function:
- Filters items, rects, and GID flags to the target page
- Collects OCR signals via
detector::page_ocr_signals(src/detector.rs) - Generates markdown via
to_markdown_from_items_with_rects_and_lines - Applies decoding checks (
is_cid_garbage,detect_encoding_issues) - Determines final
needs_ocrstatus andocr_reason
9. Result assembly
All per-page structures aggregate into PagesExtractionResult with global metadata.
Because extraction runs once on the whole document, subsequent per-page operations are lightweight and benefit from shared font statistics and layout context.
Complete Code Examples
Extract All Pages from a File
use pdf_inspector::{extract_pages_markdown_mem, PdfError};
fn main() -> Result<(), PdfError> {
// Load PDF bytes from filesystem or network
let pdf_bytes = std::fs::read("example.pdf")?;
// Extract markdown for every page
let result = extract_pages_markdown_mem(&pdf_bytes, None)?;
for page_md in result.pages {
println!("--- Page {} ---", page_md.page + 1);
if page_md.needs_ocr {
println!("⚠️ OCR required: {:?}", page_md.ocr_reason);
} else {
println!("{}", page_md.markdown);
}
}
// Inspect layout metadata
println!("Tables detected on pages: {:?}", result.pages_with_tables);
println!("Multi-column pages: {:?}", result.pages_with_columns);
println!("Complex document: {}", result.is_complex);
Ok(())
}
Extract Select Pages with OCR Routing
use pdf_inspector::{extract_pages_markdown_mem, PdfError};
fn process_select_pages(pdf_bytes: &[u8]) -> Result<(), PdfError> {
// Request specific pages (0-indexed: pages 1, 3, 5)
let selection = [0u32, 2, 4];
let result = extract_pages_markdown_mem(pdf_bytes, Some(&selection))?;
for page_md in result.pages {
match page_md.needs_ocr {
true => {
// Queue for external OCR service
eprintln!("Page {} → OCR pipeline", page_md.page + 1);
}
false => {
// Use extracted markdown directly
println!("Page {} content:\n{}", page_md.page + 1, page_md.markdown);
}
}
}
Ok(())
}
Hybrid OCR Pipeline with Result Filtering
use pdf_inspector::{extract_pages_markdown_mem, PdfError};
struct ProcessedDocument {
clean_markdown: Vec<(u32, String)>,
ocr_candidates: Vec<(u32, OcrReason)>,
layout_flags: LayoutMetadata,
}
struct LayoutMetadata {
has_tables: bool,
has_columns: bool,
complex_pages: Vec<u32>,
}
fn classify_document(pdf_bytes: &[u8]) -> Result<ProcessedDocument, PdfError> {
let extraction = extract_pages_markdown_mem(pdf_bytes, None)?;
let clean_markdown: Vec<_> = extraction
.pages
.iter()
.filter(|p| !p.needs_ocr)
.map(|p| (p.page, p.markdown.clone()))
.collect();
let ocr_candidates: Vec<_> = extraction
.pages
.iter()
.filter(|p| p.needs_ocr)
.filter_map(|p| p.ocr_reason.map(|r| (p.page, r)))
.collect();
Ok(ProcessedDocument {
clean_markdown,
ocr_candidates,
layout_flags: LayoutMetadata {
has_tables: !extraction.pages_with_tables.is_empty(),
has_columns: !extraction.pages_with_columns.is_empty(),
complex_pages: if extraction.is_complex {
extraction.pages_with_tables.clone()
} else {
vec![]
},
},
})
}
Key Source Files and Responsibilities
| File | Purpose | Key Components |
|---|---|---|
src/lib.rs |
Public API and orchestration | extract_pages_markdown_mem, PagesExtractionResult, PageMarkdown |
src/extractor/mod.rs |
Core text extraction | extract_positioned_text_for_document_analysis, extract_positioned_text_from_doc |
src/markdown/mod.rs |
Markdown generation | to_markdown_from_items_with_rects_and_lines, layout complexity functions |
src/detector.rs |
PDF type and OCR detection | page_ocr_signals, template image detection |
src/text_quality.rs |
Encoding quality analysis | analyze_text_quality, CID garbage detection |
src/markdown/analysis.rs |
Typography statistics | calculate_font_stats_from_items for heading detection |
These modules collectively implement the extraction pipeline, as implemented in firecrawl/pdf-inspector.
Summary
extract_pages_markdown_memprocesses PDF bytes directly—no temporary files required- The 0-indexed
pagesparameter enables selective extraction;Noneprocesses all pages - Single-pass document analysis provides global font statistics and layout context for consistent per-page output
PagesExtractionResultcombines markdown content with structural metadata: tables, columns, and OCR requirements- Layout detection uses geometric analysis (
all_rects,all_lines) rather than heuristic parsing - OCR classification distinguishes between scanned images, GID-encoded fonts, and garbled text encoding
Frequently Asked Questions
What is the difference between 0-indexed and 1-indexed page numbers in the result?
The pages parameter and PageMarkdown.page field use 0-indexing (consistent with Rust conventions). All layout metadata fields (pages_with_tables, pages_with_columns, pages_needing_ocr) use 1-indexing for human-readable reporting. Convert between them by adding or subtracting 1.
How does pdf-inspector determine if a page needs OCR?
The needs_ocr flag combines multiple signals: GID-encoded fonts (via gid_pages from src/extractor/mod.rs), template image detection, garbled CID text (is_cid_garbage), and encoding anomalies (detect_encoding_issues in src/text_quality.rs). The ocr_reason field provides the specific classification.
Can I use this with PDFs from HTTP responses without saving to disk?
Yes. Pass the response body bytes directly: extract_pages_markdown_mem(&response_body, None). The function never writes to the filesystem—validation, parsing, and extraction all operate on the provided &[u8] buffer.
What causes is_complex to return true?
The is_complex boolean in PagesExtractionResult activates when any page contains detected tables (pages_with_tables non-empty) or multi-column layouts (pages_with_columns non-empty). This flag helps downstream systems allocate appropriate processing resources for structured document layouts.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →