How to Use the pdf-inspector API for Custom Integrations

The pdf-inspector API provides a layered Rust interface—available with optional Python bindings—that exposes high-level functions like process_pdf and detect_pdf alongside a configurable PdfOptions builder for fine-grained control over PDF detection, text extraction, and Markdown conversion.

The firecrawl/pdf-inspector repository delivers a robust toolkit for programmatic PDF analysis. Whether you are constructing document processing pipelines or embedding PDF capabilities into existing applications, the pdf-inspector API for custom integrations offers clean abstractions that handle complex parsing operations through a consistent, single-load architecture.

Architecture and Pipeline Overview

The library organizes functionality into distinct layers that share a single Document instance, guaranteeing O(1) loading cost and consistent page counts across operations.

  • Public façade (src/lib.rs): Provides simple "one-call" helpers including process_pdf, detect_pdf, and process_pdf_with_options for common use cases.
  • Options builder (src/lib.rs): The PdfOptions struct enables fine-grained control via a fluent API for setting processing modes, page selection, and Markdown formatting.
  • Detection layer (src/detector.rs): Determines PDF type (text-based, scanned, or mixed) and analyzes layout complexity through detect_pdf_type and DetectionConfig.
  • Extraction engine (src/extractor/mod.rs): Parses text operators, fonts, and positions using functions like extract_text_with_positions, emitting structured TextItem objects with coordinates.
  • Markdown conversion (src/markdown/convert.rs): Transforms extracted items into clean Markdown or JSON output via to_markdown with configurable MarkdownProfile settings.
  • Python bindings (src/python.rs): Optional feature-gated module exposing the same API to Python applications.

The pipeline executes sequentially: load the PDF from a file path or buffer, run detection to classify the document, extract text with positional data if required, and convert the results to Markdown according to specified options.

Core API Entry Points in src/lib.rs

The primary integration surface resides in src/lib.rs, which exports three essential functions for Rust consumers.

process_pdf executes the full pipeline—detection, extraction, and Markdown conversion—in a single call. It accepts a file path and returns a result struct containing the detected PDF type, page count, and optional Markdown output.

detect_pdf performs metadata-only analysis without text extraction, ideal for quickly classifying documents or determining page counts before committing to full processing.

process_pdf_with_options accepts a PdfOptions builder instance, allowing customization of every pipeline stage while maintaining the same simple interface.

Configuring Extraction with PdfOptions

The PdfOptions builder pattern, defined in src/lib.rs (lines 164–190), provides granular control over processing behavior.

Key configuration methods include:

  • .mode(ProcessMode::Analyze) – Stops execution after the detection phase.
  • .pages([2, 4, 6]) – Restricts processing to specific page indices, reducing computation for large documents.
  • .markdown(MarkdownProfile::default().with_hard_breaks(true)) – Customizes Markdown output formatting.
  • .password("secret") – Supplies decryption credentials for protected PDFs.
  • .detection(DetectionConfig { ... }) – Configures scan strategies and layout analysis parameters.

Implementation Examples

Full Processing Pipeline

The simplest integration calls process_pdf to handle detection, extraction, and conversion automatically:

use pdf_inspector::process_pdf;

fn main() -> Result<(), pdf_inspector::PdfError> {
    let result = process_pdf("sample.pdf")?;
    println!("PDF type: {:?}", result.pdf_type);
    if let Some(md) = result.markdown {
        println!("Markdown output:\n{md}");
    }
    Ok(())
}

This invokes the public façade function in src/lib.rs (lines 65–71), which orchestrates the complete pipeline.

Fast Metadata Detection

For applications requiring only document classification without text extraction, use detect_pdf:

use pdf_inspector::detect_pdf;

fn main() -> Result<(), pdf_inspector::PdfError> {
    let info = detect_pdf("sample.pdf")?;
    println!("Detected type: {:?}, pages: {}", info.pdf_type, info.page_count);
    Ok(())
}

Implemented in src/lib.rs (lines 73–78), this function runs only the detection layer.

Custom Page Selection and Markdown Profiles

Target specific pages and customize output formatting using the PdfOptions builder:

use pdf_inspector::{PdfOptions, ProcessMode, MarkdownProfile};

fn main() -> Result<(), pdf_inspector::PdfError> {
    let opts = PdfOptions::new()
        .mode(ProcessMode::Analyze)
        .pages([2, 4, 6])
        .markdown(MarkdownProfile::default().with_hard_breaks(true));

    let result = pdf_inspector::process_pdf_with_options("sample.pdf", opts)?;
    println!("Pages needing OCR: {:?}", result.pages_needing_ocr);
    Ok(())
}

The option builder resides in src/lib.rs (lines 164–190), enabling precise control over the extraction and conversion stages.

Raw Text Extraction with Positions

Access low-level text items with coordinates for custom layout analysis by calling the extractor directly:

use pdf_inspector::extractor::extract_text_with_positions;

fn main() -> Result<(), pdf_inspector::PdfError> {
    let items = extract_text_with_positions("sample.pdf")?;
    for item in items {
        println!("{} @ ({}, {})", item.text, item.x, item.y);
    }
    Ok(())
}

This function is defined in src/extractor/mod.rs (lines 51–59), returning TextItem structs containing text content and positional rectangles.

Python Integration

Enable the python feature flag during compilation to access the API from Python:

import pdf_inspector

# Full pipeline, returns a dict with markdown and metadata

result = pdf_inspector.process_pdf("sample.pdf")
print(result["markdown"])

The Python module entry point is src/python.rs, exposing the same Rust functionality through PyO3 bindings.

Key Source Files and Their Responsibilities

Understanding the codebase structure aids in debugging and extending functionality:

  • src/lib.rs: Public API façade, PdfOptions builder, and top-level result structs.
  • src/extractor/mod.rs: Core text extraction logic, position handling, and page filtering.
  • src/detector.rs: PDF-type detection algorithms and layout-complexity analysis.
  • src/markdown/convert.rs: Markdown generation, JSON output serialization, and profile handling.
  • src/types.rs: Core data structures including TextItem, PdfLine, and PdfRect used across all layers.
  • src/python.rs: Optional Python bindings compiled when the python feature is enabled.

Summary

  • pdf-inspector exposes a layered architecture with a simple public façade in src/lib.rs and specialized modules for detection, extraction, and conversion.
  • Three main entry points—process_pdf, detect_pdf, and process_pdf_with_options—cover most integration scenarios.
  • PdfOptions builder provides fine-grained control over page selection, processing modes, Markdown formatting, and password decryption.
  • Python bindings are available via the python feature flag, offering identical functionality to the Rust API.
  • Shared document loading ensures O(1) initialization costs and consistent page counts across all operations.

Frequently Asked Questions

What is the difference between process_pdf and detect_pdf?

process_pdf executes the complete pipeline including text extraction and Markdown conversion, while detect_pdf stops after the classification phase, returning only metadata such as PDF type and page count without extracting text content. Use detect_pdf for lightweight document screening and process_pdf when you need the actual content.

How do I process only specific pages using the pdf-inspector API?

Pass a slice of page numbers to the .pages() method on the PdfOptions builder before calling process_pdf_with_options. For example, PdfOptions::new().pages([0, 2, 5]) processes only the first, third, and sixth pages, significantly reducing computation time for large documents.

Can I use pdf-inspector in Python applications?

Yes, compile the library with the python feature flag enabled to generate Python bindings via src/python.rs. The resulting module exposes functions like pdf_inspector.process_pdf() that accept file paths and return dictionaries containing Markdown output and metadata, identical to the Rust API behavior.

Where does the heavy text extraction logic live in the codebase?

The core extraction engine resides in src/extractor/mod.rs, which parses PDF text operators, fonts, and positions to generate structured TextItem objects. This module handles the low-level details of content streams and coordinate transformations, feeding data to the Markdown converter in src/markdown/convert.rs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →