# How to Extract Text from PDFs Using pdf-inspector: A Complete Technical Guide

> Learn how to extract text from PDFs using pdf-inspector. This Rust library offers a complete technical guide to raw text and structured Markdown extraction.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-04

---

**pdf-inspector is a Rust library (with Python bindings) that extracts raw text and structured Markdown from PDFs through a multi-stage pipeline involving type detection, positioned text extraction, and quality analysis.**

The `firecrawl/pdf-inspector` repository provides a robust, programmatic solution to extract text from PDFs with high fidelity. Built in Rust with optional Python bindings via N-API, this library parses PDF content streams, resolves font encodings and ToUnicode CMaps, and converts extracted text into plain text or Markdown format while automatically flagging pages that require OCR.

## The Text Extraction Pipeline

pdf-inspector implements a five-stage extraction flow that handles everything from document loading to Markdown conversion. Each stage is designed to maximize text accuracy while identifying documents or pages that need alternative processing.

### 1. Document Loading and Reuse

The pipeline begins by loading the PDF into memory once and reusing that handle for both detection and extraction. In [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), the `load_document_from_path_with_password` function (lines 82–95) creates a `lopdf::Document` instance that subsequent stages consume. This approach avoids redundant I/O when running both type detection and text extraction on the same file.

### 2. PDF Type Detection

Before extracting content, the library classifies the document to determine extraction strategy. The `detect_pdf_type` function, exported from [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) (lines 47–50) and implemented in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs), categorizes files as *TextBased*, *Scanned*, or *Mixed*. This classification determines which pages may require OCR and informs downstream formatting decisions.

### 3. Positioned Text Extraction

The core extraction logic resides in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs). The `extract_positioned_text_from_doc` function (lines 166–187) parses PDF content streams, resolves fonts and ToUnicode CMaps, and builds a vector of `TextItem` structs containing the text content, page coordinates, font metadata, and size information. The low-level operator handling—processing PDF commands like `Tj`, `TJ`, `Td`, and `Tm`—is implemented as a state machine in [`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs) (lines 30–60).

### 4. Text Quality Analysis

After extraction, pdf-inspector runs heuristics to detect garbled text, CID-encoding issues, and other corruption signals. The `analyze_text_quality` function in [`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs) (lines 5–12) evaluates the extracted content and sets `needs_ocr = true` for pages that fail quality checks, allowing downstream pipelines to route those specific pages to OCR engines.

### 5. Markdown Conversion (Optional)

When the caller requests full processing via `ProcessMode::Full`, the extracted `TextItem` structs feed into the markdown formatter. The `to_markdown_from_items_with_rects_and_page_count` function in [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) (lines 55–66) converts positioned text into Markdown, preserving structural elements like tables and headings based on layout analysis.

## Public API Methods for Text Extraction

pdf-inspector exposes several entry points in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) tailored to different use cases:

- **`extract_text(path)`** (lines 52–53): Extracts plain text from a PDF file without positional or formatting metadata. Returns a single `String` containing the document's textual content.

- **`extract_text_with_positions(path)`** (lines 52–53): Returns text together with page-wise coordinates. This is useful for region-based processing pipelines that need to know exactly where each text element appears on the page.

- **`extract_text_mem(buffer)`** (lines 98–100): Performs the same extraction as `extract_text` but operates on an in-memory PDF buffer (`&[u8]`) rather than a file path, ideal for web services processing uploaded files.

- **`extract_pages_markdown(path, pages)`** (lines 13–16): Extracts per-page Markdown with layout metadata. This method returns OCR flags and structural information alongside the formatted text.

- **`pdf2md` CLI**: A ready-to-run binary implemented in [`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) that orchestrates the full pipeline—detection, extraction, quality checks, and Markdown conversion—from the command line.

## Code Examples

### Rust: Basic Text Extraction

Extract plain text from a file using the high-level API:

```rust
use pdf_inspector::{extract_text, PdfError};

fn main() -> Result<(), PdfError> {
    let txt = extract_text("sample.pdf")?;
    println!("Extracted {} bytes of text:\n{}", txt.len(), txt);
    Ok(())
}

```

### Rust: Extract Text with Positions

For downstream layout processing, extract coordinates alongside text content:

```rust
use pdf_inspector::{extract_text_with_positions, PdfError};

fn main() -> Result<(), PdfError> {
    let (items, _rects, _lines) = extract_text_with_positions("sample.pdf")?;
    for item in items {
        println!("Page {} – ({:.1},{:.1}) \"{}\"",
                 item.page, item.x, item.y, item.text);
    }
    Ok(())
}

```

### Rust: Per-Page Markdown Extraction

Detect tables, columns, and OCR requirements while converting to Markdown:

```rust
use pdf_inspector::{extract_pages_markdown, PdfError};

fn main() -> Result<(), PdfError> {
    let result = extract_pages_markdown("sample.pdf", None)?;
    for page in result.pages {
        println!("--- Page {} ---\n{}", page.page + 1, page.markdown);
        if page.needs_ocr {
            eprintln!("⚠️  Page {} needs OCR!", page.page + 1);
        }
    }
    Ok(())
}

```

### Python: Quick Text Extraction

Using the N-API bindings:

```python
import pdf_inspector

txt = pdf_inspector.extract_text("sample.pdf")
print(txt)

```

### Command Line: Full Pipeline

Convert a PDF to Markdown via the included binary:

```bash

# Full detection + extraction, output markdown to stdout

pdf2md sample.pdf

# Only detect PDF type without text extraction

detect-pdf --json sample.pdf

```

## Key Implementation Files

Understanding the codebase structure helps when customizing or debugging text extraction:

- **[`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)**: Contains public API entry points including `process_pdf`, `extract_text`, and memory-based variants.
- **[`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)**: Implements PDF type classification and OCR-need detection heuristics.
- **[`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs)**: Orchestrates text extraction and manages font CMap resolution.
- **[`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs)**: Low-level PDF content-stream parser handling text-positioning operators.
- **[`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs)**: Heuristics for detecting garbled text, encoding issues, and OCR flags.
- **[`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs)**: Converts extracted `TextItem` structs into Markdown with layout preservation.
- **[`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs)**: CLI binary implementing the complete detection-to-markdown pipeline.

## Summary

- pdf-inspector extracts text from PDFs through a reusable `lopdf::Document` pipeline that minimizes I/O overhead.
- The library classifies PDFs by type (*TextBased*, *Scanned*, *Mixed*) before extraction to optimize processing strategy.
- Text extraction occurs at the content-stream level in [`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs), resolving fonts and CMaps to produce accurate `TextItem` structs.
- Built-in quality analysis in [`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs) automatically flags pages requiring OCR, enabling hybrid extraction workflows.
- Multiple API tiers support plain text extraction, positioned text retrieval, and full Markdown conversion with layout metadata.
- Python bindings and CLI binaries (`pdf2md`, `detect-pdf`) make the library accessible beyond the Rust ecosystem.

## Frequently Asked Questions

### Can pdf-inspector extract text from scanned PDFs?

pdf-inspector detects scanned content during the type-detection phase but does not perform OCR itself. When the `analyze_text_quality` function identifies pages with garbled text or CID-encoding issues, it marks them with `needs_ocr = true`. You can then route those specific pages to an OCR engine like Tesseract while processing the text-based pages natively.

### How does pdf-inspector handle font encoding issues?

The library resolves font encodings and ToUnicode CMaps during the extraction phase in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs). The low-level state machine in [`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs) processes PDF text-showing operators (`Tj`, `TJ`) while mapping character IDs to Unicode values, significantly reducing garbled output compared to naive text extraction.

### Is pdf-inspector available for Python projects?

Yes. pdf-inspector provides Python bindings via N-API, exposing functions like `extract_text()` directly to Python code without requiring a separate Rust toolchain. Install the package and call `pdf_inspector.extract_text("file.pdf")` to extract text from PDFs within Python applications.

### What is the difference between `extract_text` and `extract_pages_markdown`?

`extract_text` returns a single plain-text string without formatting or positional data, suitable for simple full-text indexing. `extract_pages_markdown` returns per-page Markdown strings with layout metadata, including tables, headings, and OCR flags, making it ideal for document conversion pipelines that need to preserve structure.