How firecrawl pdf-inspector Extracts Text from PDFs: A Deep Dive into the Rust Pipeline

Firecrawl pdf-inspector extracts text from PDFs by parsing content streams with a custom state‑machine, mapping raw glyph codes to Unicode via cached CMaps, and applying graphics‑state transformations to produce positioned TextItem objects.

The firecrawl/pdf-inspector repository implements a Rust‑based PDF text extraction engine that goes beyond simple string dumping. It reconstructs layout by tracking coordinate transforms, handles complex font encodings, and post‑processes the output for clean downstream consumption. This article walks through the exact pipeline used by the library to turn binary PDF content into structured text.


Step‑by‑Step PDF Text Extraction Pipeline

The extraction process follows eight distinct stages, from file ingestion through final output normalization.

1. Load and Validate the PDF Document

Extraction begins in src/lib.rs with load_document_from_path (or load_document_from_mem for byte buffers). This function validates the file header and returns a lopdf::Document — the foundational data structure representing the PDF's object graph.

use pdf_inspector::extract_text;

let text = extract_text("report.pdf")?;

The library relies on the lopdf crate for low‑level PDF parsing, but wraps it to handle edge cases encountered in real‑world documents.

2. Build Font Encoding Lookup Tables

Before processing any page content, build_font_encodings, build_font_widths, and build_type3_scales in src/extractor/fonts.rs construct caches for:

  • ToUnicode CMaps — maps glyph IDs to Unicode strings
  • Font width dictionaries — for accurate character spacing
  • Type 3 font scaling factors — custom vector fonts that require coordinate multiplication

This preprocessing step ensures that when show‑text operators appear later, raw byte sequences can be instantly decoded to readable strings without repeated dictionary lookups.

3. Strip PDF Comments from Content Streams

Some PDF generators inject comments (lines starting with %) that confuse lopdf's parser. The strip_pdf_comments function in src/extractor/content_stream.rs (lines 25‑78) sanitizes the raw content stream bytes before decoding.

4. Decode Content Stream Operations

The cleaned bytes pass through lopdf::content::Content::decode, which produces a sequence of PDF operators: BT (begin text), Tj (show text), TJ (show text with positioning), Td (move text position), and dozens more.

5. State‑Machine Traversal of Graphics State

The core extraction logic lives in extract_page_text_items (src/extractor/content_stream.rs, lines 40‑106). This function maintains a complete graphics state including:

State Component Purpose
CTM (Current Transformation Matrix) Page‑to‑device coordinate mapping
text_matrix / line_matrix Text positioning within the content stream
font / font_size Active font resource
character_spacing (Tc) / word_spacing (Tw) / text_rise (Ts) Fine‑grained spacing control

For every operator in the stream, the state updates accordingly. When show‑text operators (Tj, TJ) appear, the machine:

  1. Calls extract_text_from_operand to decode bytes using the CMap cache
  2. Applies matrix multiplication (multiply_matrices) to obtain final page coordinates
  3. Adjusts for text rise with rise_adjusted
  4. Emits a TextItem containing the text, position, font info, and markup flags

6. Handle Special PDF Constructs

The extractor recognizes several non‑text elements that affect output:

Images (Do operator with Image XObject) : Converted to placeholder TextItem objects with content like [Image: name], preserving document structure.

Marked Content with ActualText (BDC/EMC operators) : Captures accessibility text without emitting intermediate glyphs, then re‑emits as a single TextItem with width derived from surrounding matrices. Critical for screen‑reader‑friendly PDFs where displayed and actual text differ.

Underline/Strikeout Detection (re and line operators) : Geometry recorded on path construction, confirmed when paint operators execute; deferred processing enables accurate bounding‑box calculation.

7. Post‑Process and Normalize Output

After all pages complete, two merging passes clean the results:

  • merge_text_items — joins adjacent items on the same line, respecting spacing thresholds, tracking runs, and detecting RTL (right‑to‑left) text segments
  • merge_subscript_items — collapses numeric subscripts/superscripts into preceding tokens (e.g., "H₂O" stays unified rather than fragmented)

These functions reside in src/extractor/mod.rs (lines 52‑89 and 87‑110). The final output is either a flat Vec<TextItem> with positional metadata, or a PageExtraction tuple including detected rectangles and line segments for table detection pipelines.

8. Public API Entry Points

The high‑level interface in src/extractor/mod.rs provides thin wrappers around the full pipeline:

Function Return Type Use Case
extract_text(path) String Simple plain‑text extraction
extract_text_with_positions(path) Vec<TextItem> Layout‑aware processing, OCR hybrid pipelines
process_pdf(path) ProcessResult Full detection → extraction → markdown conversion

Code Examples: Using pdf-inspector in Practice

Plain Text Extraction

use pdf_inspector::extract_text;

fn main() -> Result<(), pdf_inspector::PdfError> {
    let txt = extract_text("example.pdf")?;
    println!("Full text:\n{txt}");
    Ok(())
}

Positioned Extraction for Layout Analysis

use pdf_inspector::extract_text_with_positions;

fn main() -> Result<(), pdf_inspector::PdfError> {
    let items = extract_text_with_positions("example.pdf")?;
    for it in items {
        println!(
            "Page {} – ({:.1},{:.1}) – \"{}\"  (font: {}, size: {:.1})",
            it.page, it.x, it.y, it.text, it.font, it.font_size,
        );
    }
    Ok(())
}

Full Processing Pipeline

use pdf_inspector::process_pdf;

fn main() -> Result<(), pdf_inspector::PdfError> {
    let result = process_pdf("example.pdf")?;
    println!("Detected type: {:?}", result.pdf_type);
    if let Some(md) = result.markdown {
        println!("Markdown output:\n{md}");
    }
    Ok(())
}

Key Source Files and Responsibilities

File Role
src/lib.rs Public API surface (process_pdf, detect_pdf, option builders)
src/extractor/mod.rs High‑level extraction orchestration and post‑processing
src/extractor/content_stream.rs State‑machine parser for PDF content streams
src/extractor/fonts.rs Font encoding, width tables, and Type 3 scaling
src/tounicode.rs ToUnicode CMap decoding with stream and fallback handling
src/text_utils.rs Ligature expansion, RTL detection, merging thresholds
src/types.rs Core structures: TextItem, PdfRect, PdfLine

Summary

  • pdf-inspector extracts text by parsing PDF content streams operator‑by‑operator, not by scraping rendered output.
  • The graphics state machine in src/extractor/content_stream.rs tracks coordinate transforms to produce page‑accurate positions.
  • Font encoding caching in src/extractor/fonts.rs enables fast Unicode mapping without repeated dictionary traversal.
  • Post‑processing merges adjacent tokens and handles subscripts, producing clean output for markdown conversion or downstream OCR pipelines.
  • The public API offers three extraction modes: plain text, positioned items, or full document processing with type detection.

Frequently Asked Questions

How does pdf-inspector handle PDFs with custom fonts or missing Unicode mappings?

The library builds CMap caches during initialization using build_font_encodings. When a font lacks a ToUnicode entry, it falls back to encoding dictionaries and Adobe Glyph List heuristics defined in src/tounicode.rs. Type 3 fonts receive special handling via build_type3_scales to account for their custom coordinate systems. This multi‑layer approach successfully extracts text from documents where simpler tools fail.

What is the difference between extract_text and extract_text_with_positions?

Both functions execute the same core pipeline, but return different outputs. extract_text runs merge_text_items and concatenates results into a single String, discarding positional metadata. extract_text_with_positions returns Vec<TextItem> where each item contains page, x, y, font, font_size, and markup flags. Use the latter when building layout‑aware applications like table extractors or PDF‑to‑HTML converters.

Can pdf-inspector detect tables, images, or other non‑text elements?

Yes. Images are emitted as placeholder TextItem objects with content like [Image: name]. The PageExtraction type from process_pdf includes rectangles and line_segments vectors populated by detecting path‑painting operators. While the library does not automatically reconstruct table structures, the positional text and geometric primitives provide sufficient data for downstream table detection algorithms.

Is pdf-inspector suitable for scanned PDFs or OCR workflows?

As implemented in firecrawl/pdf-inspector, the tool extracts embedded text content only — it does not perform raster image analysis. However, the positioned output format (extract_text_with_positions) is designed to hybridize with OCR engines: you can identify text‑free regions from the [Image] placeholders and coordinate data, run external OCR on those bounding boxes, and merge results back into the TextItem stream.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →