# How pdf-inspector Extracts Text from PDF Content Streams: A Deep Dive into Its State-Machine Parser

> Learn how pdf-inspector extracts text from PDF content streams using a state-machine parser. Discover its glyph ID mapping and positional metadata for accurate text extraction.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: deep-dive
- Published: 2026-08-31

---

**pdf-inspector extracts text by parsing PDF content streams with a state-machine interpreter that tracks graphics and text states, maps glyph IDs to Unicode via ToUnicode CMaps, and emits structured `TextItem`s with precise positional metadata.**

The `pdf-inspector` library from Firecrawl provides a Rust-based pipeline for converting PDF documents to Markdown or structured data. Understanding how it extracts text from PDF content streams reveals the complexity hidden behind seemingly simple documents—coordinate transformations, font encoding tables, bidirectional text, and invisible OCR layers all require careful handling.

## The Content Stream Processing Pipeline

PDF documents store page content in *content streams*—sequences of operators and operands that describe what to draw and where. The core interpreter lives in **[`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs)**, where a state-machine walks through each operator and maintains the rendering context needed to extract meaningful text.

### Stream Preparation and Decoding

Before parsing operators, pdf-inspector prepares the raw stream bytes through two steps:

1. **Strip PDF comments** — the `strip_pdf_comments` function removes top-level `%` comments that could interfere with parsing
2. **Bounded decoding** — `decode_content_bounded` decompresses and tokenizes the stream while limiting the total operator count to prevent denial-of-service from pathological PDFs

This defensive approach ensures the extractor remains robust against malformed or malicious inputs.

## Graphics State Tracking

PDF rendering depends on a stack-based **graphics state** that transforms coordinates and controls appearance. Operators in the content stream modify this state:

| Operator | Action |
|----------|--------|
| `q` / `Q` | Push / pop graphics state stack |
| `cm` | Concatenate matrix to current transformation matrix (CTM) |
| `w` | Set line width |
| `Do` | Invoke named XObject (image or form) |

The CTM—accessed via `multiply_matrices` in the source—converts local text coordinates to final page coordinates, accounting for rotation, scaling, and translation.

## Text State Management Inside BT/ET Blocks

Text extraction only occurs within **BT** (begin text) and **ET** (end text) operators. Inside these blocks, pdf-inspector tracks:

- **Text matrix** (`Tm`) and **line matrix** — current cursor position
- **Font** (`Tf`) and **font size**
- **Character spacing** (`Tc`) and **word spacing** (`Tw`)
- **Text rise** (`Ts`) — vertical offset for superscripts/subscripts
- **Rendering mode** (`Tr`) — determines visibility (mode 3 = invisible)

Positioning operators adjust these values:

- `Td`, `TD`, `Tm` — move to absolute/relative position
- `T*` — move to next line
- `TL` — set leading (line height)

## Text Operator Handling and Unicode Extraction

When the interpreter encounters text-showing operators, it converts raw glyph IDs to Unicode through a multi-stage process.

### The Tj Operator (Single String)

The **`Tj`** operator shows a single text string. The extraction path involves:

1. Raw bytes passed to `extract_text_from_operand` in [`content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/content_stream.rs)
2. Font's **ToUnicode CMap** consulted via `tounicode::FontCMaps` for glyph-to-Unicode mapping
3. Encoding differences and width tables applied from cached font data
4. **Ligature expansion** via `expand_ligatures` in [`text_utils.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/text_utils.rs)
5. Final `TextItem` created with transformed position: `multiply_matrices(&rise_adjusted(&text_matrix, text_rise), &ctm)`

### The TJ Operator (Mixed Array)

The **`TJ`** operator accepts an array containing both strings and spacing adjustments. The interpreter:

- Walks the array sequentially
- Accumulates advance widths for precise positioning
- Detects column gaps and back-tracks for RTL runs
- Emits separate `TextItem`s for each text sub-run

This operator commonly appears in justified text or documents with kerning adjustments.

### The ' Operator (Line Feed + Text)

The **`'`** operator combines `T*` (line feed) followed by `Tj`, providing a compact way to move to the next line and show text.

## Font Encoding and Caching Architecture

Before extracting text, pdf-inspector builds per-font helper structures in [`src/extractor/fonts.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/fonts.rs):

- **`build_font_encodings`** — parses Differences arrays to map character names to glyph IDs
- **`build_font_widths`** — constructs tables for exact glyph-advance calculation
- **`get_font_file2_obj_num`** — TrueType CMap fallback for fonts without explicit encoding

These structures are cached to avoid repeated parsing of ToUnicode streams, which can be large embedded CMap resources.

## Handling Invisible Text and OCR Layers

The text-rendering mode (`Tr`) value **3** designates invisible text—typically used for OCR text layers underneath scanned images. By default, pdf-inspector skips invisible text but **still advances the text matrix** to maintain correct positioning for subsequent visible text.

Enable the `include_invisible` option to recover OCR-layer text:

```rust
use pdf_inspector::lib::process_pdf_with_options;
use pdf_inspector::process_mode::ProcessMode;
use pdf_inspector::options::ExtractionOptions;

fn main() -> Result<(), pdf_inspector::PdfError> {
    let opts = ExtractionOptions {
        include_invisible: true,  // Extract OCR layer text
        ..Default::default()
    };
    let (extraction, _warnings) = process_pdf_with_options(
        "scanned_document.pdf",
        ProcessMode::Extract,
        opts,
    )?;
    Ok(())
}

```

## Bidirectional Text and Rotation Detection

During extraction, pdf-inspector records metadata for post-processing:

- **`rtl_visual_candidates`** — positions suggesting right-to-left visual order
- **`rtl_logical_ops`** — back-track operations indicating logical-order encoding

This data feeds into [`text_utils.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/text_utils.rs) algorithms that reconstruct proper reading order for Arabic, Hebrew, and mixed-direction documents.

## Recursive XObject Processing

The **`Do`** operator invokes XObjects. The interpreter handles two types via **`extract_form_xobject_text`**:

- **Image XObjects** — become placeholder `PdfRect` items marking image boundaries
- **Form XObjects** — recursively processed as nested content streams with inherited graphics state

This recursion enables extraction from complex documents with template-based page elements.

## Command-Line and Library Usage

Extract text to Markdown using the CLI:

```bash

# Default Markdown output

pdf2md mydocument.pdf > mydocument.md

# Structured JSON with full TextItem metadata

pdf2md --json mydocument.pdf > mydocument.json

```

Programmatic access through the Rust API provides granular control:

```rust
use pdf_inspector::lib::process_pdf_with_options;
use pdf_inspector::process_mode::ProcessMode;

fn main() -> Result<(), pdf_inspector::PdfError> {
    let (extraction, _warnings) = process_pdf_with_options(
        "samples/example.pdf",
        ProcessMode::Extract,
        Default::default(),
    )?;
    
    for item in extraction.0 {
        println!(
            "{} @ ({:.2}, {:.2})",
            item.text, item.x, item.y
        );
    }
    Ok(())
}

```

## Key Source Files Reference

| File | Purpose |
|------|---------|
| [`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs) | Core state-machine interpreter walking PDF operators |
| [`src/extractor/fonts.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/fonts.rs) | Font encoding, width tables, TrueType fallback |
| [`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs) | ToUnicode CMap parsing for glyph-to-Unicode mapping |
| [`src/text_utils.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_utils.rs) | Ligature expansion, RTL detection, style heuristics |
| [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) | Public API entry (`process_pdf_with_options`) |
| [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) | `TextItem` to Markdown conversion |

## Summary

- **State-machine architecture** — pdf-inspector uses a state-machine in [`content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/content_stream.rs) to track graphics and text states while walking PDF operators sequentially
- **Unicode via ToUnicode CMaps** — glyph IDs map to Unicode through font-specific CMaps with encoding differences and width tables for precise positioning
- **BT/ET block handling** — text extraction occurs only within these blocks, tracking matrices, spacing, and rendering modes
- **Invisible text support** — rendering mode 3 text is skipped by default but can be enabled via `include_invisible` for OCR layer recovery
- **RTL and bidirectional support** — runtime detection of visual vs. logical order with post-processing correction in [`text_utils.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/text_utils.rs)

## Frequently Asked Questions

### What is a PDF content stream and why does text extraction require parsing it?

A PDF content stream is a sequence of operators and operands that describes what to render on a page—text, images, vector graphics, and transformations. Text extraction requires parsing it because PDF stores text as positioned glyph IDs rather than Unicode characters; the extractor must interpret the same state changes a PDF viewer would use to reconstruct reading order and Unicode content.

### How does pdf-inspector handle fonts that lack ToUnicode CMaps?

When ToUnicode CMaps are absent, pdf-inspector falls back to encoding information parsed by `build_font_encodings` in [`src/extractor/fonts.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/fonts.rs). For TrueType fonts, it uses `get_font_file2_obj_num` to extract embedded font programs and derive character mappings. This fallback chain handles legacy PDFs or poorly generated documents.

### Can pdf-inspector preserve the exact position of every extracted character?

Yes. Each `TextItem` includes `x` and `y` coordinates computed through matrix multiplication: `multiply_matrices(&rise_adjusted(&text_matrix, text_rise), &ctm)`. This transforms local text coordinates through the CTM into final page coordinates, preserving sub-pixel precision including rotation and scaling effects.

### What causes missing or garbled text in extraction output?

Common causes include: missing or malformed ToUnicode CMaps (fallback to encoding tables may fail), subset fonts with incomplete character mappings, CID fonts with custom identity-H encodings, and encrypted PDFs where the crypt filter hasn't been properly removed. The `--json` output mode helps diagnose these by showing raw glyph IDs when Unicode mapping fails.