How to Extract Images from PDFs Using Firecrawl pdf‑inspector
Enable the include_images option in pdf‑inspector to detect Image XObjects and emit structured image metadata with bounding boxes.
Firecrawl pdf‑inspector treats every embedded image as an Image XObject. When you extract images from PDFs using Firecrawl pdf‑inspector, the engine scans page resources, creates positional placeholders, and optionally converts them into Markdown image tags. This article explains the complete extraction pipeline with code examples for Rust, Python, and CLI usage.
How Image Extraction Works in pdf‑inspector
The extraction pipeline follows a four-stage flow through specific source files:
src/extractor/xobjects.rs– Detects Image XObjects by scanning the page-level Resources dictionary for entries of subtype/Imagesrc/extractor/mod.rs(lines 511–531) – Emits aTextItemwithItemType::Imagecontaining the calculated page-space bounding boxsrc/markdown/mod.rs– Gates output via theMarkdownOptions.include_imagesflag (defaults tofalse)src/markdown/convert.rs– Converts placeholders to Markdown image syntax viato_markdown_from_lines_with_tables_and_images
The extracted metadata includes:
- Name – Generated placeholder (e.g.,
Image: Im0) - Page – PDF page number where the image appears
- BBox – Axis-aligned bounding box in PDF points as
[x0, y0, x1, y1]
CLI Method: Extract Images with pdf2md
The pdf2md binary supports image extraction through environment variables or flags.
# Using environment variable
PDF_INSPECTOR_INCLUDE_IMAGES=1 pdf2md report.pdf --json > out.json
# Query extracted images
jq '.images[] | "\(.name) on page \(.page) at \(.bbox)"' out.json
The --json flag outputs structured data including an images array when image extraction is enabled.
Rust Library Method: Programmatic Image Extraction
Use process_pdf_with_options from src/lib.rs with an Options struct:
use pdf_inspector::{process_pdf_with_options, Options};
let pdf_bytes = std::fs::read("report.pdf")?;
let opts = Options {
include_images: true,
..Default::default()
};
let result = process_pdf_with_options(&pdf_bytes, opts)?;
println!("Extracted {} images", result.markdown.images.len());
for img in result.markdown.images {
println!("Image {} on page {} at {:?}", img.name, img.page, img.bbox);
}
The PdfResult returned contains markdown.images with complete image metadata.
Python Bindings Method: Extract Images from PDFs
The Python interface in src/python.rs exposes the same functionality:
import pdf_inspector
result = pdf_inspector.process_pdf("report.pdf", include_images=True)
print(f"Found {len(result['images'])} images")
for img in result["images"]:
print(f"{img['name']} – page {img['page']} – bbox {img['bbox']}")
The returned dictionary includes an images key with a list of {name, bbox, page} objects.
Accessing Raw Image Raster Data
The placeholders provide location metadata. To retrieve actual pixel data, access the Image XObject stream directly via the lopdf crate:
- Locate the XObject by its dictionary name (
/Im0,/Im1, etc.) - Fetch the stream with
lopdf::Object::Stream - Decode using the appropriate filter (typically
/FlateDecodeor/DCTDecode)
This separation of concerns lets pdf‑inspector handle layout and positioning while you manage image decoding as needed.
Summary
- Enable extraction – Set
include_images: trueinOptionsorPDF_INSPECTOR_INCLUDE_IMAGES=1in CLI - Detection happens in –
src/extractor/xobjects.rsfor XObject scanning,src/extractor/mod.rsfor placeholder emission - Output formats – Rust structs, Python dictionaries, or JSON via CLI
- Metadata provided – Name, page number, and bounding box for each image
- Raw data access – Use
lopdfto fetch streams by XObject name when needed
Frequently Asked Questions
What image formats can pdf‑inspector detect?
pdf‑inspector detects any Image XObject defined in the PDF Resources dictionary, regardless of encoding. The detection logic in src/extractor/xobjects.rs identifies the /Image subtype; actual format decoding (JPEG, PNG, etc.) depends on the XObject's /Filter entry and requires separate handling via lopdf.
Does image extraction work with nested Form XObjects?
Yes. The extractor in src/extractor/xobjects.rs recursively handles nested Form XObjects, ensuring images embedded within complex page structures are discovered and reported with accurate bounding boxes.
Can I get image dimensions in pixels instead of PDF points?
The bbox values from pdf‑inspector are in PDF points (1/72 inch). To convert to pixels, multiply by the image's effective resolution. The raw XObject stream contains width/height in samples via /Width and /Height entries; access these through lopdf after locating the XObject by name.
Is there a performance penalty for enabling include_images?
Yes. According to the source in src/markdown/mod.rs, enabling include_images triggers additional processing in to_markdown_from_lines_with_tables_and_images to preserve reading order and insert image tags. The scan in src/extractor/xobjects.rs also adds overhead proportional to page resource complexity.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →