How firecrawl pdf-inspector Extracts Text from PDFs: A Deep Dive into the Rust Pipeline
Firecrawl pdf-inspector extracts text from PDFs by parsing content streams with a custom state‑machine, mapping raw glyph codes to Unicode via cached CMaps, and applying graphics‑state transformations to produce positioned TextItem objects.
The firecrawl/pdf-inspector repository implements a Rust‑based PDF text extraction engine that goes beyond simple string dumping. It reconstructs layout by tracking coordinate transforms, handles complex font encodings, and post‑processes the output for clean downstream consumption. This article walks through the exact pipeline used by the library to turn binary PDF content into structured text.
Step‑by‑Step PDF Text Extraction Pipeline
The extraction process follows eight distinct stages, from file ingestion through final output normalization.
1. Load and Validate the PDF Document
Extraction begins in src/lib.rs with load_document_from_path (or load_document_from_mem for byte buffers). This function validates the file header and returns a lopdf::Document — the foundational data structure representing the PDF's object graph.
use pdf_inspector::extract_text;
let text = extract_text("report.pdf")?;
The library relies on the lopdf crate for low‑level PDF parsing, but wraps it to handle edge cases encountered in real‑world documents.
2. Build Font Encoding Lookup Tables
Before processing any page content, build_font_encodings, build_font_widths, and build_type3_scales in src/extractor/fonts.rs construct caches for:
- ToUnicode CMaps — maps glyph IDs to Unicode strings
- Font width dictionaries — for accurate character spacing
- Type 3 font scaling factors — custom vector fonts that require coordinate multiplication
This preprocessing step ensures that when show‑text operators appear later, raw byte sequences can be instantly decoded to readable strings without repeated dictionary lookups.
3. Strip PDF Comments from Content Streams
Some PDF generators inject comments (lines starting with %) that confuse lopdf's parser. The strip_pdf_comments function in src/extractor/content_stream.rs (lines 25‑78) sanitizes the raw content stream bytes before decoding.
4. Decode Content Stream Operations
The cleaned bytes pass through lopdf::content::Content::decode, which produces a sequence of PDF operators: BT (begin text), Tj (show text), TJ (show text with positioning), Td (move text position), and dozens more.
5. State‑Machine Traversal of Graphics State
The core extraction logic lives in extract_page_text_items (src/extractor/content_stream.rs, lines 40‑106). This function maintains a complete graphics state including:
| State Component | Purpose |
|---|---|
CTM (Current Transformation Matrix) |
Page‑to‑device coordinate mapping |
text_matrix / line_matrix |
Text positioning within the content stream |
font / font_size |
Active font resource |
character_spacing (Tc) / word_spacing (Tw) / text_rise (Ts) |
Fine‑grained spacing control |
For every operator in the stream, the state updates accordingly. When show‑text operators (Tj, TJ) appear, the machine:
- Calls
extract_text_from_operandto decode bytes using the CMap cache - Applies matrix multiplication (
multiply_matrices) to obtain final page coordinates - Adjusts for text rise with
rise_adjusted - Emits a
TextItemcontaining the text, position, font info, and markup flags
6. Handle Special PDF Constructs
The extractor recognizes several non‑text elements that affect output:
Images (Do operator with Image XObject)
: Converted to placeholder TextItem objects with content like [Image: name], preserving document structure.
Marked Content with ActualText (BDC/EMC operators)
: Captures accessibility text without emitting intermediate glyphs, then re‑emits as a single TextItem with width derived from surrounding matrices. Critical for screen‑reader‑friendly PDFs where displayed and actual text differ.
Underline/Strikeout Detection (re and line operators)
: Geometry recorded on path construction, confirmed when paint operators execute; deferred processing enables accurate bounding‑box calculation.
7. Post‑Process and Normalize Output
After all pages complete, two merging passes clean the results:
merge_text_items— joins adjacent items on the same line, respecting spacing thresholds, tracking runs, and detecting RTL (right‑to‑left) text segmentsmerge_subscript_items— collapses numeric subscripts/superscripts into preceding tokens (e.g., "H₂O" stays unified rather than fragmented)
These functions reside in src/extractor/mod.rs (lines 52‑89 and 87‑110). The final output is either a flat Vec<TextItem> with positional metadata, or a PageExtraction tuple including detected rectangles and line segments for table detection pipelines.
8. Public API Entry Points
The high‑level interface in src/extractor/mod.rs provides thin wrappers around the full pipeline:
| Function | Return Type | Use Case |
|---|---|---|
extract_text(path) |
String |
Simple plain‑text extraction |
extract_text_with_positions(path) |
Vec<TextItem> |
Layout‑aware processing, OCR hybrid pipelines |
process_pdf(path) |
ProcessResult |
Full detection → extraction → markdown conversion |
Code Examples: Using pdf-inspector in Practice
Plain Text Extraction
use pdf_inspector::extract_text;
fn main() -> Result<(), pdf_inspector::PdfError> {
let txt = extract_text("example.pdf")?;
println!("Full text:\n{txt}");
Ok(())
}
Positioned Extraction for Layout Analysis
use pdf_inspector::extract_text_with_positions;
fn main() -> Result<(), pdf_inspector::PdfError> {
let items = extract_text_with_positions("example.pdf")?;
for it in items {
println!(
"Page {} – ({:.1},{:.1}) – \"{}\" (font: {}, size: {:.1})",
it.page, it.x, it.y, it.text, it.font, it.font_size,
);
}
Ok(())
}
Full Processing Pipeline
use pdf_inspector::process_pdf;
fn main() -> Result<(), pdf_inspector::PdfError> {
let result = process_pdf("example.pdf")?;
println!("Detected type: {:?}", result.pdf_type);
if let Some(md) = result.markdown {
println!("Markdown output:\n{md}");
}
Ok(())
}
Key Source Files and Responsibilities
| File | Role |
|---|---|
src/lib.rs |
Public API surface (process_pdf, detect_pdf, option builders) |
src/extractor/mod.rs |
High‑level extraction orchestration and post‑processing |
src/extractor/content_stream.rs |
State‑machine parser for PDF content streams |
src/extractor/fonts.rs |
Font encoding, width tables, and Type 3 scaling |
src/tounicode.rs |
ToUnicode CMap decoding with stream and fallback handling |
src/text_utils.rs |
Ligature expansion, RTL detection, merging thresholds |
src/types.rs |
Core structures: TextItem, PdfRect, PdfLine |
Summary
- pdf-inspector extracts text by parsing PDF content streams operator‑by‑operator, not by scraping rendered output.
- The graphics state machine in
src/extractor/content_stream.rstracks coordinate transforms to produce page‑accurate positions. - Font encoding caching in
src/extractor/fonts.rsenables fast Unicode mapping without repeated dictionary traversal. - Post‑processing merges adjacent tokens and handles subscripts, producing clean output for markdown conversion or downstream OCR pipelines.
- The public API offers three extraction modes: plain text, positioned items, or full document processing with type detection.
Frequently Asked Questions
How does pdf-inspector handle PDFs with custom fonts or missing Unicode mappings?
The library builds CMap caches during initialization using build_font_encodings. When a font lacks a ToUnicode entry, it falls back to encoding dictionaries and Adobe Glyph List heuristics defined in src/tounicode.rs. Type 3 fonts receive special handling via build_type3_scales to account for their custom coordinate systems. This multi‑layer approach successfully extracts text from documents where simpler tools fail.
What is the difference between extract_text and extract_text_with_positions?
Both functions execute the same core pipeline, but return different outputs. extract_text runs merge_text_items and concatenates results into a single String, discarding positional metadata. extract_text_with_positions returns Vec<TextItem> where each item contains page, x, y, font, font_size, and markup flags. Use the latter when building layout‑aware applications like table extractors or PDF‑to‑HTML converters.
Can pdf-inspector detect tables, images, or other non‑text elements?
Yes. Images are emitted as placeholder TextItem objects with content like [Image: name]. The PageExtraction type from process_pdf includes rectangles and line_segments vectors populated by detecting path‑painting operators. While the library does not automatically reconstruct table structures, the positional text and geometric primitives provide sufficient data for downstream table detection algorithms.
Is pdf-inspector suitable for scanned PDFs or OCR workflows?
As implemented in firecrawl/pdf-inspector, the tool extracts embedded text content only — it does not perform raster image analysis. However, the positioned output format (extract_text_with_positions) is designed to hybridize with OCR engines: you can identify text‑free regions from the [Image] placeholders and coordinate data, run external OCR on those bounding boxes, and merge results back into the TextItem stream.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →