# How to Use the pdf-inspector API for Custom Integrations

> Integrate pdf-inspector API into your projects. Customize PDF processing, text extraction, and Markdown conversion with this powerful Rust and Python interface.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-04

---

**The pdf-inspector API provides a layered Rust interface—available with optional Python bindings—that exposes high-level functions like `process_pdf` and `detect_pdf` alongside a configurable `PdfOptions` builder for fine-grained control over PDF detection, text extraction, and Markdown conversion.**

The `firecrawl/pdf-inspector` repository delivers a robust toolkit for programmatic PDF analysis. Whether you are constructing document processing pipelines or embedding PDF capabilities into existing applications, the **pdf-inspector API for custom integrations** offers clean abstractions that handle complex parsing operations through a consistent, single-load architecture.

## Architecture and Pipeline Overview

The library organizes functionality into distinct layers that share a single `Document` instance, guaranteeing **O(1)** loading cost and consistent page counts across operations.

- **Public façade** ([`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)): Provides simple "one-call" helpers including `process_pdf`, `detect_pdf`, and `process_pdf_with_options` for common use cases.
- **Options builder** ([`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)): The `PdfOptions` struct enables fine-grained control via a fluent API for setting processing modes, page selection, and Markdown formatting.
- **Detection layer** ([`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)): Determines PDF type (text-based, scanned, or mixed) and analyzes layout complexity through `detect_pdf_type` and `DetectionConfig`.
- **Extraction engine** ([`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs)): Parses text operators, fonts, and positions using functions like `extract_text_with_positions`, emitting structured `TextItem` objects with coordinates.
- **Markdown conversion** ([`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs)): Transforms extracted items into clean Markdown or JSON output via `to_markdown` with configurable `MarkdownProfile` settings.
- **Python bindings** ([`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs)): Optional feature-gated module exposing the same API to Python applications.

The pipeline executes sequentially: load the PDF from a file path or buffer, run detection to classify the document, extract text with positional data if required, and convert the results to Markdown according to specified options.

## Core API Entry Points in src/lib.rs

The primary integration surface resides in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), which exports three essential functions for Rust consumers.

**`process_pdf`** executes the full pipeline—detection, extraction, and Markdown conversion—in a single call. It accepts a file path and returns a result struct containing the detected PDF type, page count, and optional Markdown output.

**`detect_pdf`** performs metadata-only analysis without text extraction, ideal for quickly classifying documents or determining page counts before committing to full processing.

**`process_pdf_with_options`** accepts a `PdfOptions` builder instance, allowing customization of every pipeline stage while maintaining the same simple interface.

## Configuring Extraction with PdfOptions

The `PdfOptions` builder pattern, defined in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) (lines 164–190), provides granular control over processing behavior.

Key configuration methods include:

- `.mode(ProcessMode::Analyze)` – Stops execution after the detection phase.
- `.pages([2, 4, 6])` – Restricts processing to specific page indices, reducing computation for large documents.
- `.markdown(MarkdownProfile::default().with_hard_breaks(true))` – Customizes Markdown output formatting.
- `.password("secret")` – Supplies decryption credentials for protected PDFs.
- `.detection(DetectionConfig { ... })` – Configures scan strategies and layout analysis parameters.

## Implementation Examples

### Full Processing Pipeline

The simplest integration calls `process_pdf` to handle detection, extraction, and conversion automatically:

```rust
use pdf_inspector::process_pdf;

fn main() -> Result<(), pdf_inspector::PdfError> {
    let result = process_pdf("sample.pdf")?;
    println!("PDF type: {:?}", result.pdf_type);
    if let Some(md) = result.markdown {
        println!("Markdown output:\n{md}");
    }
    Ok(())
}

```

*This invokes the public façade function in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) (lines 65–71), which orchestrates the complete pipeline.*

### Fast Metadata Detection

For applications requiring only document classification without text extraction, use `detect_pdf`:

```rust
use pdf_inspector::detect_pdf;

fn main() -> Result<(), pdf_inspector::PdfError> {
    let info = detect_pdf("sample.pdf")?;
    println!("Detected type: {:?}, pages: {}", info.pdf_type, info.page_count);
    Ok(())
}

```

*Implemented in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) (lines 73–78), this function runs only the detection layer.*

### Custom Page Selection and Markdown Profiles

Target specific pages and customize output formatting using the `PdfOptions` builder:

```rust
use pdf_inspector::{PdfOptions, ProcessMode, MarkdownProfile};

fn main() -> Result<(), pdf_inspector::PdfError> {
    let opts = PdfOptions::new()
        .mode(ProcessMode::Analyze)
        .pages([2, 4, 6])
        .markdown(MarkdownProfile::default().with_hard_breaks(true));

    let result = pdf_inspector::process_pdf_with_options("sample.pdf", opts)?;
    println!("Pages needing OCR: {:?}", result.pages_needing_ocr);
    Ok(())
}

```

*The option builder resides in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) (lines 164–190), enabling precise control over the extraction and conversion stages.*

### Raw Text Extraction with Positions

Access low-level text items with coordinates for custom layout analysis by calling the extractor directly:

```rust
use pdf_inspector::extractor::extract_text_with_positions;

fn main() -> Result<(), pdf_inspector::PdfError> {
    let items = extract_text_with_positions("sample.pdf")?;
    for item in items {
        println!("{} @ ({}, {})", item.text, item.x, item.y);
    }
    Ok(())
}

```

*This function is defined in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) (lines 51–59), returning `TextItem` structs containing text content and positional rectangles.*

### Python Integration

Enable the `python` feature flag during compilation to access the API from Python:

```python
import pdf_inspector

# Full pipeline, returns a dict with markdown and metadata

result = pdf_inspector.process_pdf("sample.pdf")
print(result["markdown"])

```

*The Python module entry point is [`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs), exposing the same Rust functionality through PyO3 bindings.*

## Key Source Files and Their Responsibilities

Understanding the codebase structure aids in debugging and extending functionality:

- **[`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)**: Public API façade, `PdfOptions` builder, and top-level result structs.
- **[`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs)**: Core text extraction logic, position handling, and page filtering.
- **[`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)**: PDF-type detection algorithms and layout-complexity analysis.
- **[`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs)**: Markdown generation, JSON output serialization, and profile handling.
- **[`src/types.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs)**: Core data structures including `TextItem`, `PdfLine`, and `PdfRect` used across all layers.
- **[`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs)**: Optional Python bindings compiled when the `python` feature is enabled.

## Summary

- **pdf-inspector** exposes a layered architecture with a simple public façade in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) and specialized modules for detection, extraction, and conversion.
- **Three main entry points**—`process_pdf`, `detect_pdf`, and `process_pdf_with_options`—cover most integration scenarios.
- **PdfOptions builder** provides fine-grained control over page selection, processing modes, Markdown formatting, and password decryption.
- **Python bindings** are available via the `python` feature flag, offering identical functionality to the Rust API.
- **Shared document loading** ensures O(1) initialization costs and consistent page counts across all operations.

## Frequently Asked Questions

### What is the difference between `process_pdf` and `detect_pdf`?

`process_pdf` executes the complete pipeline including text extraction and Markdown conversion, while `detect_pdf` stops after the classification phase, returning only metadata such as PDF type and page count without extracting text content. Use `detect_pdf` for lightweight document screening and `process_pdf` when you need the actual content.

### How do I process only specific pages using the pdf-inspector API?

Pass a slice of page numbers to the `.pages()` method on the `PdfOptions` builder before calling `process_pdf_with_options`. For example, `PdfOptions::new().pages([0, 2, 5])` processes only the first, third, and sixth pages, significantly reducing computation time for large documents.

### Can I use pdf-inspector in Python applications?

Yes, compile the library with the `python` feature flag enabled to generate Python bindings via [`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs). The resulting module exposes functions like `pdf_inspector.process_pdf()` that accept file paths and return dictionaries containing Markdown output and metadata, identical to the Rust API behavior.

### Where does the heavy text extraction logic live in the codebase?

The core extraction engine resides in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs), which parses PDF text operators, fonts, and positions to generate structured `TextItem` objects. This module handles the low-level details of content streams and coordinate transformations, feeding data to the Markdown converter in [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs).