How to Extract Text from PDFs Using pdf-inspector: A Complete Technical Guide

pdf-inspector is a Rust library (with Python bindings) that extracts raw text and structured Markdown from PDFs through a multi-stage pipeline involving type detection, positioned text extraction, and quality analysis.

The firecrawl/pdf-inspector repository provides a robust, programmatic solution to extract text from PDFs with high fidelity. Built in Rust with optional Python bindings via N-API, this library parses PDF content streams, resolves font encodings and ToUnicode CMaps, and converts extracted text into plain text or Markdown format while automatically flagging pages that require OCR.

The Text Extraction Pipeline

pdf-inspector implements a five-stage extraction flow that handles everything from document loading to Markdown conversion. Each stage is designed to maximize text accuracy while identifying documents or pages that need alternative processing.

1. Document Loading and Reuse

The pipeline begins by loading the PDF into memory once and reusing that handle for both detection and extraction. In src/lib.rs, the load_document_from_path_with_password function (lines 82–95) creates a lopdf::Document instance that subsequent stages consume. This approach avoids redundant I/O when running both type detection and text extraction on the same file.

2. PDF Type Detection

Before extracting content, the library classifies the document to determine extraction strategy. The detect_pdf_type function, exported from src/lib.rs (lines 47–50) and implemented in src/detector.rs, categorizes files as TextBased, Scanned, or Mixed. This classification determines which pages may require OCR and informs downstream formatting decisions.

3. Positioned Text Extraction

The core extraction logic resides in src/extractor/mod.rs. The extract_positioned_text_from_doc function (lines 166–187) parses PDF content streams, resolves fonts and ToUnicode CMaps, and builds a vector of TextItem structs containing the text content, page coordinates, font metadata, and size information. The low-level operator handling—processing PDF commands like Tj, TJ, Td, and Tm—is implemented as a state machine in src/extractor/content_stream.rs (lines 30–60).

4. Text Quality Analysis

After extraction, pdf-inspector runs heuristics to detect garbled text, CID-encoding issues, and other corruption signals. The analyze_text_quality function in src/text_quality.rs (lines 5–12) evaluates the extracted content and sets needs_ocr = true for pages that fail quality checks, allowing downstream pipelines to route those specific pages to OCR engines.

5. Markdown Conversion (Optional)

When the caller requests full processing via ProcessMode::Full, the extracted TextItem structs feed into the markdown formatter. The to_markdown_from_items_with_rects_and_page_count function in src/markdown/convert.rs (lines 55–66) converts positioned text into Markdown, preserving structural elements like tables and headings based on layout analysis.

Public API Methods for Text Extraction

pdf-inspector exposes several entry points in src/lib.rs tailored to different use cases:

  • extract_text(path) (lines 52–53): Extracts plain text from a PDF file without positional or formatting metadata. Returns a single String containing the document's textual content.

  • extract_text_with_positions(path) (lines 52–53): Returns text together with page-wise coordinates. This is useful for region-based processing pipelines that need to know exactly where each text element appears on the page.

  • extract_text_mem(buffer) (lines 98–100): Performs the same extraction as extract_text but operates on an in-memory PDF buffer (&[u8]) rather than a file path, ideal for web services processing uploaded files.

  • extract_pages_markdown(path, pages) (lines 13–16): Extracts per-page Markdown with layout metadata. This method returns OCR flags and structural information alongside the formatted text.

  • pdf2md CLI: A ready-to-run binary implemented in src/bin/pdf2md.rs that orchestrates the full pipeline—detection, extraction, quality checks, and Markdown conversion—from the command line.

Code Examples

Rust: Basic Text Extraction

Extract plain text from a file using the high-level API:

use pdf_inspector::{extract_text, PdfError};

fn main() -> Result<(), PdfError> {
    let txt = extract_text("sample.pdf")?;
    println!("Extracted {} bytes of text:\n{}", txt.len(), txt);
    Ok(())
}

Rust: Extract Text with Positions

For downstream layout processing, extract coordinates alongside text content:

use pdf_inspector::{extract_text_with_positions, PdfError};

fn main() -> Result<(), PdfError> {
    let (items, _rects, _lines) = extract_text_with_positions("sample.pdf")?;
    for item in items {
        println!("Page {} – ({:.1},{:.1}) \"{}\"",
                 item.page, item.x, item.y, item.text);
    }
    Ok(())
}

Rust: Per-Page Markdown Extraction

Detect tables, columns, and OCR requirements while converting to Markdown:

use pdf_inspector::{extract_pages_markdown, PdfError};

fn main() -> Result<(), PdfError> {
    let result = extract_pages_markdown("sample.pdf", None)?;
    for page in result.pages {
        println!("--- Page {} ---\n{}", page.page + 1, page.markdown);
        if page.needs_ocr {
            eprintln!("⚠️  Page {} needs OCR!", page.page + 1);
        }
    }
    Ok(())
}

Python: Quick Text Extraction

Using the N-API bindings:

import pdf_inspector

txt = pdf_inspector.extract_text("sample.pdf")
print(txt)

Command Line: Full Pipeline

Convert a PDF to Markdown via the included binary:


# Full detection + extraction, output markdown to stdout

pdf2md sample.pdf

# Only detect PDF type without text extraction

detect-pdf --json sample.pdf

Key Implementation Files

Understanding the codebase structure helps when customizing or debugging text extraction:

Summary

  • pdf-inspector extracts text from PDFs through a reusable lopdf::Document pipeline that minimizes I/O overhead.
  • The library classifies PDFs by type (TextBased, Scanned, Mixed) before extraction to optimize processing strategy.
  • Text extraction occurs at the content-stream level in src/extractor/content_stream.rs, resolving fonts and CMaps to produce accurate TextItem structs.
  • Built-in quality analysis in src/text_quality.rs automatically flags pages requiring OCR, enabling hybrid extraction workflows.
  • Multiple API tiers support plain text extraction, positioned text retrieval, and full Markdown conversion with layout metadata.
  • Python bindings and CLI binaries (pdf2md, detect-pdf) make the library accessible beyond the Rust ecosystem.

Frequently Asked Questions

Can pdf-inspector extract text from scanned PDFs?

pdf-inspector detects scanned content during the type-detection phase but does not perform OCR itself. When the analyze_text_quality function identifies pages with garbled text or CID-encoding issues, it marks them with needs_ocr = true. You can then route those specific pages to an OCR engine like Tesseract while processing the text-based pages natively.

How does pdf-inspector handle font encoding issues?

The library resolves font encodings and ToUnicode CMaps during the extraction phase in src/extractor/mod.rs. The low-level state machine in src/extractor/content_stream.rs processes PDF text-showing operators (Tj, TJ) while mapping character IDs to Unicode values, significantly reducing garbled output compared to naive text extraction.

Is pdf-inspector available for Python projects?

Yes. pdf-inspector provides Python bindings via N-API, exposing functions like extract_text() directly to Python code without requiring a separate Rust toolchain. Install the package and call pdf_inspector.extract_text("file.pdf") to extract text from PDFs within Python applications.

What is the difference between extract_text and extract_pages_markdown?

extract_text returns a single plain-text string without formatting or positional data, suitable for simple full-text indexing. extract_pages_markdown returns per-page Markdown strings with layout metadata, including tables, headings, and OCR flags, making it ideal for document conversion pipelines that need to preserve structure.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →