How to Extract Metadata from PDFs Using pdf-inspector: A Complete Guide

To extract metadata from PDFs using pdf-inspector, call the detect_pdf_type function (or language-specific equivalents) which parses the document trailer to retrieve the title, page count, and classification without performing full text extraction.

The firecrawl/pdf-inspector library provides a high-performance Rust-based toolkit for analyzing PDF documents. When you need to extract metadata from PDFs—such as page counts, titles, or document classifications—the library offers a lightweight detection pipeline that operates without invoking costly OCR or text extraction processes.

How pdf-inspector Extracts PDF Metadata

The extraction process relies on parsing the PDF's internal document structure rather than rendering pages or analyzing content streams.

The Detection Pipeline Entry Points

According to the firecrawl/pdf-inspector source code, the public API exposes two primary functions in src/lib.rs: detect_pdf_type for file paths and detect_pdf_type_mem for in-memory buffers. These functions load the document using load_document_from_path or load_document_from_mem respectively, then delegate processing to the internal detect_from_document function defined in src/detector.rs.

Reading the Document Trailer and Info Dictionary

Inside src/detector.rs, the detect_from_document function queries the PDF's trailer dictionary via doc.trailer.get(b"Info") to extract the optional title string. The total page count is retrieved through Document::load_metadata(). This approach accesses only the document's cross-reference table and trailer, avoiding the need to parse page content streams.

Metadata Fields in the PdfTypeResult Structure

The detection pipeline returns a PdfTypeResult structure defined in src/detector.rs that bundles the following metadata fields:

  • pdf_type – Classification of the document (e.g., TextBased, Scanned, or Mixed)
  • page_count – Total number of pages as reported by the PDF metadata
  • title – Optional document title extracted from the Info dictionary
  • confidence – Classifier confidence score ranging from 0.0 to 1.0
  • ocr_recommended – Boolean indicating whether OCR processing is advised
  • pages_needing_ocr – List of specific page numbers requiring OCR
  • ocr_reasons_by_page – Machine-readable explanations for OCR recommendations

Implementation Examples by Language

The library provides consistent metadata extraction interfaces across multiple programming environments.

Rust Implementation

Use the detect_pdf_type function from the crate root:

use pdf_inspector::{detect_pdf_type, PdfTypeResult};

fn main() -> Result<(), pdf_inspector::PdfError> {
    let info: PdfTypeResult = detect_pdf_type("example.pdf")?;
    println!("Title:   {:?}", info.title.unwrap_or("<none>".into()));
    println!("Pages:   {}", info.page_count);
    println!("Type:    {:?}", info.pdf_type);
    println!("Confidence: {:.2}", info.confidence);
    Ok(())
}

Python Implementation

The PyO3 bindings in src/python.rs expose the detect_pdf function:

import pdf_inspector

info = pdf_inspector.detect_pdf("example.pdf")
print("Title:", info.get("title", "<none>"))
print("Pages:", info["page_count"])
print("Type:", info["pdf_type"])
print("Confidence:", info["confidence"])

Node.js Implementation

For JavaScript environments, use the N-API wrapper:

import { readFileSync } from "fs";
import { detectPdf } from "@firecrawl/pdf-inspector";

const buffer = readFileSync("example.pdf");
const info = detectPdf(buffer);
console.log("Title:", info.title ?? "<none>");
console.log("Pages:", info.pageCount);
console.log("Type:", info.pdfType);
console.log("Confidence:", info.confidence);

Command-Line Interface

The CLI binary implemented in src/bin/detect_pdf.rs provides JSON output:

detect-pdf example.pdf --json

# {"pdf_type":"TextBased","page_count":12,"title":"Annual Report 2025",...}

Performance Benefits of Metadata-Only Extraction

Because the detection path never walks content streams for text extraction, it executes in milliseconds even for large PDFs. The document loads once through Document::load_metadata(), allowing the resulting PdfTypeResult to be reused by subsequent extraction stages without re-parsing the file. This architecture makes pdf-inspector ideal for high-throughput workflows that require document classification or cataloging without full content analysis.

Summary

  • Primary API: Use detect_pdf_type (file) or detect_pdf_type_mem (buffer) from src/lib.rs to extract metadata from PDFs.
  • Core Logic: The detect_from_document function in src/detector.rs queries doc.trailer.get(b"Info") for titles and Document::load_metadata() for page counts.
  • Result Structure: PdfTypeResult provides pdf_type, page_count, title, confidence, and OCR recommendations.
  • Multi-Language Support: Interfaces available for Rust, Python, Node.js, and CLI with consistent field naming.
  • Performance: Metadata extraction operates without OCR or content stream parsing, enabling sub-second analysis of large documents.

Frequently Asked Questions

What metadata fields does pdf-inspector extract from PDFs?

pdf-inspector returns the pdf_type classification, page_count, optional title from the Info dictionary, a confidence score, and OCR-related flags including ocr_recommended and pages_needing_ocr. These fields are packaged in the PdfTypeResult structure defined in src/detector.rs.

Does pdf-inspector require OCR to extract metadata?

No. The metadata extraction pipeline implemented in detect_from_document reads only the PDF trailer and metadata dictionaries. It does not perform text extraction or OCR, making it significantly faster than full-content analysis.

How do I extract metadata from a PDF in memory using pdf-inspector?

Use the detect_pdf_type_mem function in Rust, which accepts a byte buffer instead of a file path. In Python, the detect_pdf function can accept file-like objects or paths depending on the binding implementation. The Node.js API accepts Buffer objects directly via the detectPdf function.

Where is the PDF metadata extraction logic implemented in the source code?

The extraction logic resides primarily in src/detector.rs within the detect_from_document function, which is called by the public API functions in src/lib.rs. The type definitions for the returned metadata structure are located in src/detector.rs (for PdfTypeResult) and src/types.rs (for PdfType enumerations).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →