How to Extract Images from PDFs Using Firecrawl pdf‑inspector

Enable the include_images option in pdf‑inspector to detect Image XObjects and emit structured image metadata with bounding boxes.

Firecrawl pdf‑inspector treats every embedded image as an Image XObject. When you extract images from PDFs using Firecrawl pdf‑inspector, the engine scans page resources, creates positional placeholders, and optionally converts them into Markdown image tags. This article explains the complete extraction pipeline with code examples for Rust, Python, and CLI usage.

How Image Extraction Works in pdf‑inspector

The extraction pipeline follows a four-stage flow through specific source files:

  • src/extractor/xobjects.rs – Detects Image XObjects by scanning the page-level Resources dictionary for entries of subtype /Image
  • src/extractor/mod.rs (lines 511–531) – Emits a TextItem with ItemType::Image containing the calculated page-space bounding box
  • src/markdown/mod.rs – Gates output via the MarkdownOptions.include_images flag (defaults to false)
  • src/markdown/convert.rs – Converts placeholders to Markdown image syntax via to_markdown_from_lines_with_tables_and_images

The extracted metadata includes:

  • Name – Generated placeholder (e.g., Image: Im0)
  • Page – PDF page number where the image appears
  • BBox – Axis-aligned bounding box in PDF points as [x0, y0, x1, y1]

CLI Method: Extract Images with pdf2md

The pdf2md binary supports image extraction through environment variables or flags.


# Using environment variable

PDF_INSPECTOR_INCLUDE_IMAGES=1 pdf2md report.pdf --json > out.json

# Query extracted images

jq '.images[] | "\(.name) on page \(.page) at \(.bbox)"' out.json

The --json flag outputs structured data including an images array when image extraction is enabled.

Rust Library Method: Programmatic Image Extraction

Use process_pdf_with_options from src/lib.rs with an Options struct:

use pdf_inspector::{process_pdf_with_options, Options};

let pdf_bytes = std::fs::read("report.pdf")?;
let opts = Options {
    include_images: true,
    ..Default::default()
};

let result = process_pdf_with_options(&pdf_bytes, opts)?;
println!("Extracted {} images", result.markdown.images.len());

for img in result.markdown.images {
    println!("Image {} on page {} at {:?}", img.name, img.page, img.bbox);
}

The PdfResult returned contains markdown.images with complete image metadata.

Python Bindings Method: Extract Images from PDFs

The Python interface in src/python.rs exposes the same functionality:

import pdf_inspector

result = pdf_inspector.process_pdf("report.pdf", include_images=True)

print(f"Found {len(result['images'])} images")
for img in result["images"]:
    print(f"{img['name']} – page {img['page']} – bbox {img['bbox']}")

The returned dictionary includes an images key with a list of {name, bbox, page} objects.

Accessing Raw Image Raster Data

The placeholders provide location metadata. To retrieve actual pixel data, access the Image XObject stream directly via the lopdf crate:

  1. Locate the XObject by its dictionary name (/Im0, /Im1, etc.)
  2. Fetch the stream with lopdf::Object::Stream
  3. Decode using the appropriate filter (typically /FlateDecode or /DCTDecode)

This separation of concerns lets pdf‑inspector handle layout and positioning while you manage image decoding as needed.

Summary

  • Enable extraction – Set include_images: true in Options or PDF_INSPECTOR_INCLUDE_IMAGES=1 in CLI
  • Detection happens in – src/extractor/xobjects.rs for XObject scanning, src/extractor/mod.rs for placeholder emission
  • Output formats – Rust structs, Python dictionaries, or JSON via CLI
  • Metadata provided – Name, page number, and bounding box for each image
  • Raw data access – Use lopdf to fetch streams by XObject name when needed

Frequently Asked Questions

What image formats can pdf‑inspector detect?

pdf‑inspector detects any Image XObject defined in the PDF Resources dictionary, regardless of encoding. The detection logic in src/extractor/xobjects.rs identifies the /Image subtype; actual format decoding (JPEG, PNG, etc.) depends on the XObject's /Filter entry and requires separate handling via lopdf.

Does image extraction work with nested Form XObjects?

Yes. The extractor in src/extractor/xobjects.rs recursively handles nested Form XObjects, ensuring images embedded within complex page structures are discovered and reported with accurate bounding boxes.

Can I get image dimensions in pixels instead of PDF points?

The bbox values from pdf‑inspector are in PDF points (1/72 inch). To convert to pixels, multiply by the image's effective resolution. The raw XObject stream contains width/height in samples via /Width and /Height entries; access these through lopdf after locating the XObject by name.

Is there a performance penalty for enabling include_images?

Yes. According to the source in src/markdown/mod.rs, enabling include_images triggers additional processing in to_markdown_from_lines_with_tables_and_images to preserve reading order and insert image tags. The scan in src/extractor/xobjects.rs also adds overhead proportional to page resource complexity.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →