# How to Extract Metadata from PDFs Using pdf-inspector: A Complete Guide

> Easily extract PDF metadata like title and page count using pdf-inspector. Learn how to call the detect_pdf_type function for quick, efficient metadata retrieval without full text parsing.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-04

---

**To extract metadata from PDFs using pdf-inspector, call the `detect_pdf_type` function (or language-specific equivalents) which parses the document trailer to retrieve the title, page count, and classification without performing full text extraction.**

The `firecrawl/pdf-inspector` library provides a high-performance Rust-based toolkit for analyzing PDF documents. When you need to extract metadata from PDFs—such as page counts, titles, or document classifications—the library offers a lightweight detection pipeline that operates without invoking costly OCR or text extraction processes.

## How pdf-inspector Extracts PDF Metadata

The extraction process relies on parsing the PDF's internal document structure rather than rendering pages or analyzing content streams.

### The Detection Pipeline Entry Points

According to the firecrawl/pdf-inspector source code, the public API exposes two primary functions in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs): `detect_pdf_type` for file paths and `detect_pdf_type_mem` for in-memory buffers. These functions load the document using `load_document_from_path` or `load_document_from_mem` respectively, then delegate processing to the internal `detect_from_document` function defined in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs).

### Reading the Document Trailer and Info Dictionary

Inside [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs), the `detect_from_document` function queries the PDF's trailer dictionary via `doc.trailer.get(b"Info")` to extract the optional title string. The total page count is retrieved through `Document::load_metadata()`. This approach accesses only the document's cross-reference table and trailer, avoiding the need to parse page content streams.

## Metadata Fields in the PdfTypeResult Structure

The detection pipeline returns a **`PdfTypeResult`** structure defined in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) that bundles the following metadata fields:

- **`pdf_type`** – Classification of the document (e.g., `TextBased`, `Scanned`, or `Mixed`)
- **`page_count`** – Total number of pages as reported by the PDF metadata
- **`title`** – Optional document title extracted from the Info dictionary
- **`confidence`** – Classifier confidence score ranging from 0.0 to 1.0
- **`ocr_recommended`** – Boolean indicating whether OCR processing is advised
- **`pages_needing_ocr`** – List of specific page numbers requiring OCR
- **`ocr_reasons_by_page`** – Machine-readable explanations for OCR recommendations

## Implementation Examples by Language

The library provides consistent metadata extraction interfaces across multiple programming environments.

### Rust Implementation

Use the `detect_pdf_type` function from the crate root:

```rust
use pdf_inspector::{detect_pdf_type, PdfTypeResult};

fn main() -> Result<(), pdf_inspector::PdfError> {
    let info: PdfTypeResult = detect_pdf_type("example.pdf")?;
    println!("Title:   {:?}", info.title.unwrap_or("<none>".into()));
    println!("Pages:   {}", info.page_count);
    println!("Type:    {:?}", info.pdf_type);
    println!("Confidence: {:.2}", info.confidence);
    Ok(())
}

```

### Python Implementation

The PyO3 bindings in [`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs) expose the `detect_pdf` function:

```python
import pdf_inspector

info = pdf_inspector.detect_pdf("example.pdf")
print("Title:", info.get("title", "<none>"))
print("Pages:", info["page_count"])
print("Type:", info["pdf_type"])
print("Confidence:", info["confidence"])

```

### Node.js Implementation

For JavaScript environments, use the N-API wrapper:

```javascript
import { readFileSync } from "fs";
import { detectPdf } from "@firecrawl/pdf-inspector";

const buffer = readFileSync("example.pdf");
const info = detectPdf(buffer);
console.log("Title:", info.title ?? "<none>");
console.log("Pages:", info.pageCount);
console.log("Type:", info.pdfType);
console.log("Confidence:", info.confidence);

```

### Command-Line Interface

The CLI binary implemented in [`src/bin/detect_pdf.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs) provides JSON output:

```bash
detect-pdf example.pdf --json

# {"pdf_type":"TextBased","page_count":12,"title":"Annual Report 2025",...}

```

## Performance Benefits of Metadata-Only Extraction

Because the detection path never walks content streams for text extraction, it executes in milliseconds even for large PDFs. The document loads once through `Document::load_metadata()`, allowing the resulting `PdfTypeResult` to be reused by subsequent extraction stages without re-parsing the file. This architecture makes pdf-inspector ideal for high-throughput workflows that require document classification or cataloging without full content analysis.

## Summary

- **Primary API**: Use `detect_pdf_type` (file) or `detect_pdf_type_mem` (buffer) from [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) to extract metadata from PDFs.
- **Core Logic**: The `detect_from_document` function in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) queries `doc.trailer.get(b"Info")` for titles and `Document::load_metadata()` for page counts.
- **Result Structure**: `PdfTypeResult` provides `pdf_type`, `page_count`, `title`, `confidence`, and OCR recommendations.
- **Multi-Language Support**: Interfaces available for Rust, Python, Node.js, and CLI with consistent field naming.
- **Performance**: Metadata extraction operates without OCR or content stream parsing, enabling sub-second analysis of large documents.

## Frequently Asked Questions

### What metadata fields does pdf-inspector extract from PDFs?

pdf-inspector returns the `pdf_type` classification, `page_count`, optional `title` from the Info dictionary, a `confidence` score, and OCR-related flags including `ocr_recommended` and `pages_needing_ocr`. These fields are packaged in the `PdfTypeResult` structure defined in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs).

### Does pdf-inspector require OCR to extract metadata?

No. The metadata extraction pipeline implemented in `detect_from_document` reads only the PDF trailer and metadata dictionaries. It does not perform text extraction or OCR, making it significantly faster than full-content analysis.

### How do I extract metadata from a PDF in memory using pdf-inspector?

Use the `detect_pdf_type_mem` function in Rust, which accepts a byte buffer instead of a file path. In Python, the `detect_pdf` function can accept file-like objects or paths depending on the binding implementation. The Node.js API accepts Buffer objects directly via the `detectPdf` function.

### Where is the PDF metadata extraction logic implemented in the source code?

The extraction logic resides primarily in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) within the `detect_from_document` function, which is called by the public API functions in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs). The type definitions for the returned metadata structure are located in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) (for `PdfTypeResult`) and [`src/types.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs) (for `PdfType` enumerations).