How to Parse PDF Metadata Using firecrawl pdf-inspector: A Complete Guide

Use process_pdf() to read a PDF once and access the metadata field on the returned PdfResult struct, which exposes standard PDF Info dictionary entries like title, author, subject, and creation_date.

The firecrawl pdf-inspector crate provides fast, unified PDF introspection across Rust, Python, Node.js, and CLI environments. When you need to extract document properties without performing full content extraction, the library's single-load pipeline reads the PDF once and surfaces all metadata fields in a consistent structure. This approach avoids the overhead of redundant I/O operations while giving you immediate access to the Info dictionary that every compliant PDF contains.

Core API: The process_pdf() Entry Point

In src/lib.rs, the process_pdf() function serves as the primary interface for all bindings. It accepts a file path and PdfOptions, then returns a PdfResult containing both detection metadata and optional full-text content.

// src/lib.rs, lines 64-73
pub fn process_pdf<P: AsRef<Path>>(
    path: P,
    options: PdfOptions,
) -> Result<PdfResult, PdfError> {
    let doc = load_document_from_path(path.as_ref())?;
    // Detection and metadata extraction happen here
    let detector_result = detector::detect(&doc)?;
    // ...
}

The PdfResult struct includes a metadata: Option<PdfMetadata> field. This PdfMetadata type mirrors the PDF Info dictionary with strongly-typed fields for common entries.

Accessing Metadata Fields in Rust

When working directly with the Rust crate, pattern match on the metadata option to access individual fields:

use pdf_inspector::{process_pdf, PdfOptions};

fn main() -> Result<(), pdf_inspector::PdfError> {
    let opts = PdfOptions::default();              // Fast detection, no OCR
    let result = process_pdf("contract.pdf", opts)?;

    if let Some(meta) = result.metadata {
        println!("Title:      {}", meta.title.unwrap_or("Untitled"));
        println!("Author:     {}", meta.author.unwrap_or("Unknown"));
        println!("Subject:    {}", meta.subject.unwrap_or("N/A"));
        println!("Created:    {:?}", meta.creation_date);
        println!("Modified:   {:?}", meta.modification_date);
        println!("Producer:   {}", meta.producer.unwrap_or("N/A"));
        println!("PDF Version: {}", meta.pdf_version);
    } else {
        println!("No metadata dictionary found in PDF");
    }

    Ok(())
}

The PdfMetadata struct implements Default, so missing fields return None rather than causing errors. This design ensures robust handling of malformed or minimal PDFs.

Python Binding: Direct Metadata Access

The PyO3 wrapper in src/python.rs exposes the same metadata attribute as a Python object with named properties. The docstring explicitly notes "Title from PDF metadata" as the primary use case for document classification.

import pdf_inspector

# Single call loads PDF and extracts metadata

result = pdf_inspector.process_pdf("whitepaper.pdf")

# Access metadata fields directly

meta = result.metadata
print(f"Document: {meta.title or 'No title'}")
print(f"By:       {meta.author or 'Anonymous'}")
print(f"About:    {meta.subject or 'No subject'}")
print(f"Created:  {meta.creation_date}")

# Check if PDF is text-based or scanned before OCR

print(f"PDF type: {result.pdf_type}")  # "text", "scanned", "hybrid", etc.

The Python binding preserves the null-safety of the Rust API: absent fields appear as None rather than raising AttributeError.

Node.js (N-API) Integration

For JavaScript and TypeScript environments, the N-API layer in napi/src/lib.rs maps the Rust PdfResult to a plain JavaScript object. The metadata property passes through without transformation, maintaining identical field names.

import { readFileSync } from 'fs';
import { processPdf } from '@firecrawl/pdf-inspector';

// Buffer input allows processing without file system round-trips
const pdfBuffer = readFileSync('report.pdf');
const result = processPdf(pdfBuffer);

// Destructure metadata with fallback defaults
const { title = 'Untitled', author = 'Unknown', creationDate } = result.metadata;

console.log(`Document: ${title}`);
console.log(`Author:   ${author}`);
console.log(`Created:  ${creationDate?.toISOString?.() || 'N/A'}`);

// Use detection metadata to route processing
if (result.pdfType === 'scanned') {
  console.warn('Scanned PDF detected — consider OCR for text extraction');
}

The Node.js binding accepts both Buffer and string path inputs, matching the flexibility of the underlying Rust API.

CLI: Metadata-Only Detection with JSON Output

The detect-pdf binary provides a command-line interface optimized for scripting and automation. Pass --json to receive structured output including the full metadata object.


# Basic metadata extraction

detect-pdf invoice.pdf --json

# Filter with jq for specific fields

detect-pdf proposal.pdf --json | jq '.metadata | {title, author, created}'

# Combine with other tools in pipelines

for pdf in *.pdf; do
  detect-pdf "$pdf" --json | jq -r '[.metadata.title, .metadata.author] | @tsv'
done > documents.tsv

The CLI implementation in src/bin/detect_pdf.rs uses the same process_pdf() path as all other interfaces, ensuring consistent behavior across environments.

Metadata Extraction Pipeline Internals

Understanding the internal flow helps optimize usage:

  1. load_document_from_path() — Opens the PDF file and parses the cross-reference table to locate the Info dictionary without rendering page content.

  2. detector::detect() — Reads the Info dictionary entries (/Title, /Author, /Subject, /CreationDate, /ModDate, /Producer, /Creator) and populates the PdfMetadata struct. This detector also classifies the PDF as text-based, scanned, or hybrid using heuristics applied to the first few pages.

  3. Conditional full-text extraction — Only if PdfOptions requests content (e.g., markdown output) does the pipeline proceed to text extraction. Pure metadata detection skips this expensive phase.

This architecture means metadata-only queries complete in milliseconds even for multi-hundred-page documents, as the parser never touches page content streams.

Common Metadata Fields Reference

Field PDF Key Rust Type Description
title /Title Option<String> Document title as set by the authoring application
author /Author Option<String> Creator or author name
subject /Subject Option<String> Subject or abstract summary
keywords /Keywords Option<String> Comma-separated keywords
creator /Creator Option<String> Application that created the original document
producer /Producer Option<String> PDF conversion tool or library
creation_date /CreationDate Option<DateTime<Utc>> When the PDF was first created (PDF date format)
modification_date /ModDate Option<DateTime<Utc>> Last modification timestamp
pdf_version header String PDF version from file header (e.g., "1.4")

Dates are parsed from PDF's proprietary date string format (D:YYYYMMDDHHmmSSOHH'mm') and normalized to UTC DateTime objects in Rust, then converted to native date types in Python and JavaScript bindings.

Handling Missing or Malformed Metadata

Not all PDFs contain complete Info dictionaries. The pdf-inspector handles these cases gracefully:

  • Absent dictionary: metadata field is None (Rust) / None (Python) / null (JS)
  • Missing individual fields: Specific properties are None/null
  • Invalid date strings: Parsed as None rather than causing parse errors
  • Binary or corrupted fields: Sanitized to valid UTF-8, with replacement characters for unrecoverable sequences

This permissive parsing ensures your application processes real-world PDFs without crashing on edge cases.

Performance Characteristics

  • Metadata-only detection: ~1-5ms for typical documents (< 10MB)
  • Memory footprint: O(1) relative to document size; only the Info dictionary and trailer are loaded
  • Parallel processing: The PdfOptions type is Send + Sync, enabling concurrent metadata extraction across file collections

For bulk processing pipelines, disable full-text extraction in PdfOptions to maximize throughput:

let opts = PdfOptions {
    extract_text: false,
    extract_images: false,
    ocr_enabled: false,
    ..PdfOptions::default()
};

Summary

  • process_pdf() in src/lib.rs provides the single entry point for metadata extraction across all language bindings
  • The PdfMetadata struct exposes standard PDF Info dictionary fields with null-safe Option types
  • Python, Node.js, and CLI interfaces mirror the Rust API with idiomatic naming conventions
  • Metadata extraction uses a fast path that avoids page content parsing, completing in milliseconds
  • All bindings handle absent or malformed metadata gracefully without raising errors

Frequently Asked Questions

How do I extract PDF metadata without installing the full Rust toolchain?

Use the Python package (pip install pdf-inspector) or Node.js package (npm install @firecrawl/pdf-inspector). Both provide pre-compiled binaries for major platforms and expose the same metadata interface as the Rust crate. The CLI binary can also be downloaded as a standalone release from the GitHub repository.

Why is the creation_date field returning null for some PDFs?

PDF creators are not required to populate the /CreationDate entry in the Info dictionary. Some scanning software and legacy tools omit this field entirely. Additionally, malformed date strings that violate PDF date format specifications are parsed as None rather than causing errors. Always handle creation_date and modification_date as optional in your application logic.

Can I modify PDF metadata using pdf-inspector?

No. The current pdf-inspector release is read-only for metadata operations. The PdfMetadata struct and all binding interfaces expose only getter methods. For metadata editing, you would need to use a PDF manipulation library such as pikepdf (Python) or lopdf (Rust) in conjunction with pdf-inspector for initial reading.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →