How to Parse PDF Metadata Using firecrawl pdf-inspector: A Complete Guide
Use process_pdf() to read a PDF once and access the metadata field on the returned PdfResult struct, which exposes standard PDF Info dictionary entries like title, author, subject, and creation_date.
The firecrawl pdf-inspector crate provides fast, unified PDF introspection across Rust, Python, Node.js, and CLI environments. When you need to extract document properties without performing full content extraction, the library's single-load pipeline reads the PDF once and surfaces all metadata fields in a consistent structure. This approach avoids the overhead of redundant I/O operations while giving you immediate access to the Info dictionary that every compliant PDF contains.
Core API: The process_pdf() Entry Point
In src/lib.rs, the process_pdf() function serves as the primary interface for all bindings. It accepts a file path and PdfOptions, then returns a PdfResult containing both detection metadata and optional full-text content.
// src/lib.rs, lines 64-73
pub fn process_pdf<P: AsRef<Path>>(
path: P,
options: PdfOptions,
) -> Result<PdfResult, PdfError> {
let doc = load_document_from_path(path.as_ref())?;
// Detection and metadata extraction happen here
let detector_result = detector::detect(&doc)?;
// ...
}
The PdfResult struct includes a metadata: Option<PdfMetadata> field. This PdfMetadata type mirrors the PDF Info dictionary with strongly-typed fields for common entries.
Accessing Metadata Fields in Rust
When working directly with the Rust crate, pattern match on the metadata option to access individual fields:
use pdf_inspector::{process_pdf, PdfOptions};
fn main() -> Result<(), pdf_inspector::PdfError> {
let opts = PdfOptions::default(); // Fast detection, no OCR
let result = process_pdf("contract.pdf", opts)?;
if let Some(meta) = result.metadata {
println!("Title: {}", meta.title.unwrap_or("Untitled"));
println!("Author: {}", meta.author.unwrap_or("Unknown"));
println!("Subject: {}", meta.subject.unwrap_or("N/A"));
println!("Created: {:?}", meta.creation_date);
println!("Modified: {:?}", meta.modification_date);
println!("Producer: {}", meta.producer.unwrap_or("N/A"));
println!("PDF Version: {}", meta.pdf_version);
} else {
println!("No metadata dictionary found in PDF");
}
Ok(())
}
The PdfMetadata struct implements Default, so missing fields return None rather than causing errors. This design ensures robust handling of malformed or minimal PDFs.
Python Binding: Direct Metadata Access
The PyO3 wrapper in src/python.rs exposes the same metadata attribute as a Python object with named properties. The docstring explicitly notes "Title from PDF metadata" as the primary use case for document classification.
import pdf_inspector
# Single call loads PDF and extracts metadata
result = pdf_inspector.process_pdf("whitepaper.pdf")
# Access metadata fields directly
meta = result.metadata
print(f"Document: {meta.title or 'No title'}")
print(f"By: {meta.author or 'Anonymous'}")
print(f"About: {meta.subject or 'No subject'}")
print(f"Created: {meta.creation_date}")
# Check if PDF is text-based or scanned before OCR
print(f"PDF type: {result.pdf_type}") # "text", "scanned", "hybrid", etc.
The Python binding preserves the null-safety of the Rust API: absent fields appear as None rather than raising AttributeError.
Node.js (N-API) Integration
For JavaScript and TypeScript environments, the N-API layer in napi/src/lib.rs maps the Rust PdfResult to a plain JavaScript object. The metadata property passes through without transformation, maintaining identical field names.
import { readFileSync } from 'fs';
import { processPdf } from '@firecrawl/pdf-inspector';
// Buffer input allows processing without file system round-trips
const pdfBuffer = readFileSync('report.pdf');
const result = processPdf(pdfBuffer);
// Destructure metadata with fallback defaults
const { title = 'Untitled', author = 'Unknown', creationDate } = result.metadata;
console.log(`Document: ${title}`);
console.log(`Author: ${author}`);
console.log(`Created: ${creationDate?.toISOString?.() || 'N/A'}`);
// Use detection metadata to route processing
if (result.pdfType === 'scanned') {
console.warn('Scanned PDF detected — consider OCR for text extraction');
}
The Node.js binding accepts both Buffer and string path inputs, matching the flexibility of the underlying Rust API.
CLI: Metadata-Only Detection with JSON Output
The detect-pdf binary provides a command-line interface optimized for scripting and automation. Pass --json to receive structured output including the full metadata object.
# Basic metadata extraction
detect-pdf invoice.pdf --json
# Filter with jq for specific fields
detect-pdf proposal.pdf --json | jq '.metadata | {title, author, created}'
# Combine with other tools in pipelines
for pdf in *.pdf; do
detect-pdf "$pdf" --json | jq -r '[.metadata.title, .metadata.author] | @tsv'
done > documents.tsv
The CLI implementation in src/bin/detect_pdf.rs uses the same process_pdf() path as all other interfaces, ensuring consistent behavior across environments.
Metadata Extraction Pipeline Internals
Understanding the internal flow helps optimize usage:
-
load_document_from_path()— Opens the PDF file and parses the cross-reference table to locate the Info dictionary without rendering page content. -
detector::detect()— Reads the Info dictionary entries (/Title,/Author,/Subject,/CreationDate,/ModDate,/Producer,/Creator) and populates thePdfMetadatastruct. This detector also classifies the PDF as text-based, scanned, or hybrid using heuristics applied to the first few pages. -
Conditional full-text extraction — Only if
PdfOptionsrequests content (e.g., markdown output) does the pipeline proceed to text extraction. Pure metadata detection skips this expensive phase.
This architecture means metadata-only queries complete in milliseconds even for multi-hundred-page documents, as the parser never touches page content streams.
Common Metadata Fields Reference
| Field | PDF Key | Rust Type | Description |
|---|---|---|---|
title |
/Title |
Option<String> |
Document title as set by the authoring application |
author |
/Author |
Option<String> |
Creator or author name |
subject |
/Subject |
Option<String> |
Subject or abstract summary |
keywords |
/Keywords |
Option<String> |
Comma-separated keywords |
creator |
/Creator |
Option<String> |
Application that created the original document |
producer |
/Producer |
Option<String> |
PDF conversion tool or library |
creation_date |
/CreationDate |
Option<DateTime<Utc>> |
When the PDF was first created (PDF date format) |
modification_date |
/ModDate |
Option<DateTime<Utc>> |
Last modification timestamp |
pdf_version |
header | String |
PDF version from file header (e.g., "1.4") |
Dates are parsed from PDF's proprietary date string format (D:YYYYMMDDHHmmSSOHH'mm') and normalized to UTC DateTime objects in Rust, then converted to native date types in Python and JavaScript bindings.
Handling Missing or Malformed Metadata
Not all PDFs contain complete Info dictionaries. The pdf-inspector handles these cases gracefully:
- Absent dictionary:
metadatafield isNone(Rust) /None(Python) /null(JS) - Missing individual fields: Specific properties are
None/null - Invalid date strings: Parsed as
Nonerather than causing parse errors - Binary or corrupted fields: Sanitized to valid UTF-8, with replacement characters for unrecoverable sequences
This permissive parsing ensures your application processes real-world PDFs without crashing on edge cases.
Performance Characteristics
- Metadata-only detection: ~1-5ms for typical documents (< 10MB)
- Memory footprint: O(1) relative to document size; only the Info dictionary and trailer are loaded
- Parallel processing: The
PdfOptionstype isSend + Sync, enabling concurrent metadata extraction across file collections
For bulk processing pipelines, disable full-text extraction in PdfOptions to maximize throughput:
let opts = PdfOptions {
extract_text: false,
extract_images: false,
ocr_enabled: false,
..PdfOptions::default()
};
Summary
process_pdf()insrc/lib.rsprovides the single entry point for metadata extraction across all language bindings- The
PdfMetadatastruct exposes standard PDF Info dictionary fields with null-safeOptiontypes - Python, Node.js, and CLI interfaces mirror the Rust API with idiomatic naming conventions
- Metadata extraction uses a fast path that avoids page content parsing, completing in milliseconds
- All bindings handle absent or malformed metadata gracefully without raising errors
Frequently Asked Questions
How do I extract PDF metadata without installing the full Rust toolchain?
Use the Python package (pip install pdf-inspector) or Node.js package (npm install @firecrawl/pdf-inspector). Both provide pre-compiled binaries for major platforms and expose the same metadata interface as the Rust crate. The CLI binary can also be downloaded as a standalone release from the GitHub repository.
Why is the creation_date field returning null for some PDFs?
PDF creators are not required to populate the /CreationDate entry in the Info dictionary. Some scanning software and legacy tools omit this field entirely. Additionally, malformed date strings that violate PDF date format specifications are parsed as None rather than causing errors. Always handle creation_date and modification_date as optional in your application logic.
Can I modify PDF metadata using pdf-inspector?
No. The current pdf-inspector release is read-only for metadata operations. The PdfMetadata struct and all binding interfaces expose only getter methods. For metadata editing, you would need to use a PDF manipulation library such as pikepdf (Python) or lopdf (Rust) in conjunction with pdf-inspector for initial reading.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →