# How to Parse PDF Metadata Using firecrawl pdf-inspector: A Complete Guide

> Learn to parse PDF metadata with firecrawl pdf-inspector. Easily extract title author subject and creation date using the process_pdf() function for your technical projects.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-07

---

**Use `process_pdf()` to read a PDF once and access the `metadata` field on the returned `PdfResult` struct, which exposes standard PDF Info dictionary entries like `title`, `author`, `subject`, and `creation_date`.**

The **firecrawl pdf-inspector** crate provides fast, unified PDF introspection across Rust, Python, Node.js, and CLI environments. When you need to extract document properties without performing full content extraction, the library's single-load pipeline reads the PDF once and surfaces all metadata fields in a consistent structure. This approach avoids the overhead of redundant I/O operations while giving you immediate access to the Info dictionary that every compliant PDF contains.

## Core API: The `process_pdf()` Entry Point

In [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), the `process_pdf()` function serves as the primary interface for all bindings. It accepts a file path and `PdfOptions`, then returns a `PdfResult` containing both detection metadata and optional full-text content.

```rust
// src/lib.rs, lines 64-73
pub fn process_pdf<P: AsRef<Path>>(
    path: P,
    options: PdfOptions,
) -> Result<PdfResult, PdfError> {
    let doc = load_document_from_path(path.as_ref())?;
    // Detection and metadata extraction happen here
    let detector_result = detector::detect(&doc)?;
    // ...
}

```

The `PdfResult` struct includes a `metadata: Option<PdfMetadata>` field. This `PdfMetadata` type mirrors the PDF Info dictionary with strongly-typed fields for common entries.

## Accessing Metadata Fields in Rust

When working directly with the Rust crate, pattern match on the `metadata` option to access individual fields:

```rust
use pdf_inspector::{process_pdf, PdfOptions};

fn main() -> Result<(), pdf_inspector::PdfError> {
    let opts = PdfOptions::default();              // Fast detection, no OCR
    let result = process_pdf("contract.pdf", opts)?;

    if let Some(meta) = result.metadata {
        println!("Title:      {}", meta.title.unwrap_or("Untitled"));
        println!("Author:     {}", meta.author.unwrap_or("Unknown"));
        println!("Subject:    {}", meta.subject.unwrap_or("N/A"));
        println!("Created:    {:?}", meta.creation_date);
        println!("Modified:   {:?}", meta.modification_date);
        println!("Producer:   {}", meta.producer.unwrap_or("N/A"));
        println!("PDF Version: {}", meta.pdf_version);
    } else {
        println!("No metadata dictionary found in PDF");
    }

    Ok(())
}

```

The `PdfMetadata` struct implements `Default`, so missing fields return `None` rather than causing errors. This design ensures robust handling of malformed or minimal PDFs.

## Python Binding: Direct Metadata Access

The PyO3 wrapper in [`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs) exposes the same `metadata` attribute as a Python object with named properties. The docstring explicitly notes "Title from PDF metadata" as the primary use case for document classification.

```python
import pdf_inspector

# Single call loads PDF and extracts metadata

result = pdf_inspector.process_pdf("whitepaper.pdf")

# Access metadata fields directly

meta = result.metadata
print(f"Document: {meta.title or 'No title'}")
print(f"By:       {meta.author or 'Anonymous'}")
print(f"About:    {meta.subject or 'No subject'}")
print(f"Created:  {meta.creation_date}")

# Check if PDF is text-based or scanned before OCR

print(f"PDF type: {result.pdf_type}")  # "text", "scanned", "hybrid", etc.

```

The Python binding preserves the null-safety of the Rust API: absent fields appear as `None` rather than raising `AttributeError`.

## Node.js (N-API) Integration

For JavaScript and TypeScript environments, the N-API layer in [`napi/src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/napi/src/lib.rs) maps the Rust `PdfResult` to a plain JavaScript object. The `metadata` property passes through without transformation, maintaining identical field names.

```javascript
import { readFileSync } from 'fs';
import { processPdf } from '@firecrawl/pdf-inspector';

// Buffer input allows processing without file system round-trips
const pdfBuffer = readFileSync('report.pdf');
const result = processPdf(pdfBuffer);

// Destructure metadata with fallback defaults
const { title = 'Untitled', author = 'Unknown', creationDate } = result.metadata;

console.log(`Document: ${title}`);
console.log(`Author:   ${author}`);
console.log(`Created:  ${creationDate?.toISOString?.() || 'N/A'}`);

// Use detection metadata to route processing
if (result.pdfType === 'scanned') {
  console.warn('Scanned PDF detected — consider OCR for text extraction');
}

```

The Node.js binding accepts both `Buffer` and `string` path inputs, matching the flexibility of the underlying Rust API.

## CLI: Metadata-Only Detection with JSON Output

The `detect-pdf` binary provides a command-line interface optimized for scripting and automation. Pass `--json` to receive structured output including the full metadata object.

```bash

# Basic metadata extraction

detect-pdf invoice.pdf --json

# Filter with jq for specific fields

detect-pdf proposal.pdf --json | jq '.metadata | {title, author, created}'

# Combine with other tools in pipelines

for pdf in *.pdf; do
  detect-pdf "$pdf" --json | jq -r '[.metadata.title, .metadata.author] | @tsv'
done > documents.tsv

```

The CLI implementation in [`src/bin/detect_pdf.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs) uses the same `process_pdf()` path as all other interfaces, ensuring consistent behavior across environments.

## Metadata Extraction Pipeline Internals

Understanding the internal flow helps optimize usage:

1. **`load_document_from_path()`** — Opens the PDF file and parses the cross-reference table to locate the **Info dictionary** without rendering page content.

2. **`detector::detect()`** — Reads the Info dictionary entries (`/Title`, `/Author`, `/Subject`, `/CreationDate`, `/ModDate`, `/Producer`, `/Creator`) and populates the `PdfMetadata` struct. This detector also classifies the PDF as text-based, scanned, or hybrid using heuristics applied to the first few pages.

3. **Conditional full-text extraction** — Only if `PdfOptions` requests content (e.g., markdown output) does the pipeline proceed to text extraction. Pure metadata detection skips this expensive phase.

This architecture means metadata-only queries complete in milliseconds even for multi-hundred-page documents, as the parser never touches page content streams.

## Common Metadata Fields Reference

| Field | PDF Key | Rust Type | Description |
|-------|---------|-----------|-------------|
| `title` | `/Title` | `Option<String>` | Document title as set by the authoring application |
| `author` | `/Author` | `Option<String>` | Creator or author name |
| `subject` | `/Subject` | `Option<String>` | Subject or abstract summary |
| `keywords` | `/Keywords` | `Option<String>` | Comma-separated keywords |
| `creator` | `/Creator` | `Option<String>` | Application that created the original document |
| `producer` | `/Producer` | `Option<String>` | PDF conversion tool or library |
| `creation_date` | `/CreationDate` | `Option<DateTime<Utc>>` | When the PDF was first created (PDF date format) |
| `modification_date` | `/ModDate` | `Option<DateTime<Utc>>` | Last modification timestamp |
| `pdf_version` | header | `String` | PDF version from file header (e.g., "1.4") |

Dates are parsed from PDF's proprietary date string format (`D:YYYYMMDDHHmmSSOHH'mm'`) and normalized to UTC `DateTime` objects in Rust, then converted to native date types in Python and JavaScript bindings.

## Handling Missing or Malformed Metadata

Not all PDFs contain complete Info dictionaries. The pdf-inspector handles these cases gracefully:

- **Absent dictionary**: `metadata` field is `None` (Rust) / `None` (Python) / `null` (JS)
- **Missing individual fields**: Specific properties are `None`/`null`
- **Invalid date strings**: Parsed as `None` rather than causing parse errors
- **Binary or corrupted fields**: Sanitized to valid UTF-8, with replacement characters for unrecoverable sequences

This permissive parsing ensures your application processes real-world PDFs without crashing on edge cases.

## Performance Characteristics

- **Metadata-only detection**: ~1-5ms for typical documents (< 10MB)
- **Memory footprint**: O(1) relative to document size; only the Info dictionary and trailer are loaded
- **Parallel processing**: The `PdfOptions` type is `Send + Sync`, enabling concurrent metadata extraction across file collections

For bulk processing pipelines, disable full-text extraction in `PdfOptions` to maximize throughput:

```rust
let opts = PdfOptions {
    extract_text: false,
    extract_images: false,
    ocr_enabled: false,
    ..PdfOptions::default()
};

```

## Summary

- **`process_pdf()`** in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) provides the single entry point for metadata extraction across all language bindings
- The **`PdfMetadata`** struct exposes standard PDF Info dictionary fields with null-safe `Option` types
- **Python**, **Node.js**, and **CLI** interfaces mirror the Rust API with idiomatic naming conventions
- Metadata extraction uses a **fast path** that avoids page content parsing, completing in milliseconds
- All bindings handle **absent or malformed metadata** gracefully without raising errors

## Frequently Asked Questions

### How do I extract PDF metadata without installing the full Rust toolchain?

Use the **Python package** (`pip install pdf-inspector`) or **Node.js package** (`npm install @firecrawl/pdf-inspector`). Both provide pre-compiled binaries for major platforms and expose the same `metadata` interface as the Rust crate. The CLI binary can also be downloaded as a standalone release from the GitHub repository.

### Why is the `creation_date` field returning null for some PDFs?

PDF creators are not required to populate the `/CreationDate` entry in the Info dictionary. Some scanning software and legacy tools omit this field entirely. Additionally, malformed date strings that violate PDF date format specifications are parsed as `None` rather than causing errors. Always handle `creation_date` and `modification_date` as optional in your application logic.

### Can I modify PDF metadata using pdf-inspector?

No. The current `pdf-inspector` release is **read-only** for metadata operations. The `PdfMetadata` struct and all binding interfaces expose only getter methods. For metadata editing, you would need to use a PDF manipulation library such as `pikepdf` (Python) or `lopdf` (Rust) in conjunction with pdf-inspector for initial reading.