# How firecrawl pdf-inspector Extracts Text from PDFs: A Deep Dive into the Rust Pipeline

> Discover how firecrawl pdf-inspector extracts text from PDFs using its Rust pipeline. Learn about state machines CMaps and graphics transformations for accurate text extraction.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: deep-dive
- Published: 2026-08-07

---

**Firecrawl pdf-inspector extracts text from PDFs by parsing content streams with a custom state‑machine, mapping raw glyph codes to Unicode via cached CMaps, and applying graphics‑state transformations to produce positioned `TextItem` objects.**

The `firecrawl/pdf-inspector` repository implements a Rust‑based PDF text extraction engine that goes beyond simple string dumping. It reconstructs layout by tracking coordinate transforms, handles complex font encodings, and post‑processes the output for clean downstream consumption. This article walks through the **exact pipeline** used by the library to turn binary PDF content into structured text.

---

## Step‑by‑Step PDF Text Extraction Pipeline

The extraction process follows **eight distinct stages**, from file ingestion through final output normalization.

### 1. Load and Validate the PDF Document

Extraction begins in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) with `load_document_from_path` (or `load_document_from_mem` for byte buffers). This function validates the file header and returns a `lopdf::Document` — the foundational data structure representing the PDF's object graph.

```rust
use pdf_inspector::extract_text;

let text = extract_text("report.pdf")?;

```

The library relies on the **`lopdf`** crate for low‑level PDF parsing, but wraps it to handle edge cases encountered in real‑world documents.

### 2. Build Font Encoding Lookup Tables

Before processing any page content, `build_font_encodings`, `build_font_widths`, and `build_type3_scales` in [`src/extractor/fonts.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/fonts.rs) construct caches for:

- **ToUnicode CMaps** — maps glyph IDs to Unicode strings
- **Font width dictionaries** — for accurate character spacing
- **Type 3 font scaling factors** — custom vector fonts that require coordinate multiplication

This preprocessing step ensures that when show‑text operators appear later, raw byte sequences can be **instantly decoded to readable strings** without repeated dictionary lookups.

### 3. Strip PDF Comments from Content Streams

Some PDF generators inject comments (lines starting with `%`) that confuse lopdf's parser. The `strip_pdf_comments` function in [`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs) (lines 25‑78) sanitizes the raw content stream bytes before decoding.

### 4. Decode Content Stream Operations

The cleaned bytes pass through `lopdf::content::Content::decode`, which produces a sequence of **PDF operators**: `BT` (begin text), `Tj` (show text), `TJ` (show text with positioning), `Td` (move text position), and dozens more.

### 5. State‑Machine Traversal of Graphics State

The core extraction logic lives in `extract_page_text_items` ([`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs), lines 40‑106). This function maintains a **complete graphics state** including:

| State Component | Purpose |
|-----------------|---------|
| `CTM` (Current Transformation Matrix) | Page‑to‑device coordinate mapping |
| `text_matrix` / `line_matrix` | Text positioning within the content stream |
| `font` / `font_size` | Active font resource |
| `character_spacing` (`Tc`) / `word_spacing` (`Tw`) / `text_rise` (`Ts`) | Fine‑grained spacing control |

For every operator in the stream, the state updates accordingly. When **show‑text operators** (`Tj`, `TJ`) appear, the machine:

1. Calls `extract_text_from_operand` to decode bytes using the CMap cache
2. Applies matrix multiplication (`multiply_matrices`) to obtain final page coordinates
3. Adjusts for text rise with `rise_adjusted`
4. Emits a `TextItem` containing the text, position, font info, and markup flags

### 6. Handle Special PDF Constructs

The extractor recognizes several **non‑text elements** that affect output:

**Images** (`Do` operator with Image XObject)
: Converted to placeholder `TextItem` objects with content like `[Image: name]`, preserving document structure.

**Marked Content with ActualText** (`BDC`/`EMC` operators)
: Captures accessibility text without emitting intermediate glyphs, then re‑emits as a single `TextItem` with width derived from surrounding matrices. Critical for screen‑reader‑friendly PDFs where displayed and actual text differ.

**Underline/Strikeout Detection** (`re` and line operators)
: Geometry recorded on path construction, confirmed when paint operators execute; deferred processing enables accurate bounding‑box calculation.

### 7. Post‑Process and Normalize Output

After all pages complete, two merging passes clean the results:

- **`merge_text_items`** — joins adjacent items on the same line, respecting spacing thresholds, tracking runs, and detecting **RTL** (right‑to‑left) text segments
- **`merge_subscript_items`** — collapses numeric subscripts/superscripts into preceding tokens (e.g., "H₂O" stays unified rather than fragmented)

These functions reside in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) (lines 52‑89 and 87‑110). The final output is either a flat `Vec<TextItem>` with positional metadata, or a `PageExtraction` tuple including detected rectangles and line segments for **table detection** pipelines.

### 8. Public API Entry Points

The high‑level interface in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) provides thin wrappers around the full pipeline:

| Function | Return Type | Use Case |
|----------|-------------|----------|
| `extract_text(path)` | `String` | Simple plain‑text extraction |
| `extract_text_with_positions(path)` | `Vec<TextItem>` | Layout‑aware processing, OCR hybrid pipelines |
| `process_pdf(path)` | `ProcessResult` | Full detection → extraction → markdown conversion |

---

## Code Examples: Using pdf-inspector in Practice

### Plain Text Extraction

```rust
use pdf_inspector::extract_text;

fn main() -> Result<(), pdf_inspector::PdfError> {
    let txt = extract_text("example.pdf")?;
    println!("Full text:\n{txt}");
    Ok(())
}

```

### Positioned Extraction for Layout Analysis

```rust
use pdf_inspector::extract_text_with_positions;

fn main() -> Result<(), pdf_inspector::PdfError> {
    let items = extract_text_with_positions("example.pdf")?;
    for it in items {
        println!(
            "Page {} – ({:.1},{:.1}) – \"{}\"  (font: {}, size: {:.1})",
            it.page, it.x, it.y, it.text, it.font, it.font_size,
        );
    }
    Ok(())
}

```

### Full Processing Pipeline

```rust
use pdf_inspector::process_pdf;

fn main() -> Result<(), pdf_inspector::PdfError> {
    let result = process_pdf("example.pdf")?;
    println!("Detected type: {:?}", result.pdf_type);
    if let Some(md) = result.markdown {
        println!("Markdown output:\n{md}");
    }
    Ok(())
}

```

---

## Key Source Files and Responsibilities

| File | Role |
|------|------|
| [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) | Public API surface (`process_pdf`, `detect_pdf`, option builders) |
| [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) | High‑level extraction orchestration and post‑processing |
| [`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs) | State‑machine parser for PDF content streams |
| [`src/extractor/fonts.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/fonts.rs) | Font encoding, width tables, and Type 3 scaling |
| [`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs) | ToUnicode CMap decoding with stream and fallback handling |
| [`src/text_utils.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_utils.rs) | Ligature expansion, RTL detection, merging thresholds |
| [`src/types.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs) | Core structures: `TextItem`, `PdfRect`, `PdfLine` |

---

## Summary

- **pdf-inspector** extracts text by parsing PDF **content streams** operator‑by‑operator, not by scraping rendered output.
- The **graphics state machine** in [`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs) tracks coordinate transforms to produce **page‑accurate positions**.
- **Font encoding caching** in [`src/extractor/fonts.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/fonts.rs) enables fast Unicode mapping without repeated dictionary traversal.
- **Post‑processing merges** adjacent tokens and handles subscripts, producing clean output for **markdown conversion** or **downstream OCR** pipelines.
- The public API offers **three extraction modes**: plain text, positioned items, or full document processing with type detection.

---

## Frequently Asked Questions

### How does pdf-inspector handle PDFs with custom fonts or missing Unicode mappings?

The library builds **CMap caches** during initialization using `build_font_encodings`. When a font lacks a ToUnicode entry, it falls back to **encoding dictionaries** and **Adobe Glyph List** heuristics defined in [`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs). Type 3 fonts receive special handling via `build_type3_scales` to account for their custom coordinate systems. This multi‑layer approach successfully extracts text from documents where simpler tools fail.

### What is the difference between `extract_text` and `extract_text_with_positions`?

Both functions execute the same core pipeline, but **return different outputs**. `extract_text` runs `merge_text_items` and concatenates results into a single `String`, discarding positional metadata. `extract_text_with_positions` returns `Vec<TextItem>` where each item contains `page`, `x`, `y`, `font`, `font_size`, and markup flags. Use the latter when building **layout‑aware applications** like table extractors or PDF‑to‑HTML converters.

### Can pdf-inspector detect tables, images, or other non‑text elements?

Yes. **Images** are emitted as placeholder `TextItem` objects with content like `[Image: name]`. The `PageExtraction` type from `process_pdf` includes `rectangles` and `line_segments` vectors populated by detecting path‑painting operators. While the library does not automatically reconstruct table structures, the **positional text and geometric primitives** provide sufficient data for downstream table detection algorithms.

### Is pdf-inspector suitable for scanned PDFs or OCR workflows?

As implemented in `firecrawl/pdf-inspector`, the tool extracts **embedded text content only** — it does not perform raster image analysis. However, the positioned output format (`extract_text_with_positions`) is designed to **hybridize with OCR engines**: you can identify text‑free regions from the `[Image]` placeholders and coordinate data, run external OCR on those bounding boxes, and merge results back into the `TextItem` stream.