How pdf-inspector Handles Vector Text (Path-Operator Text) Versus Actual Text in PDFs

pdf-inspector detects vector-drawn text by counting path operators versus text operators in the PDF content stream, then automatically falls back to OCR when glyphs are rendered as geometric paths rather than actual text objects.

PDF documents store text in two fundamentally different ways: as actual text objects using standard PDF text operators (Tj, TJ, Td, Tm), or as vector-drawn glyphs rendered through geometric path operators (m, l, c, re). The firecrawl/pdf-inspector library distinguishes between these approaches using a purpose-built detection heuristic, ensuring reliable text extraction regardless of how the content was originally encoded.

How Vector-Drawn Text Differs from Actual Text

Actual text in PDFs uses the Show operator family to display character codes mapped to fonts. This allows direct extraction by reading the content stream and decoding the glyphs.

Vector-drawn text converts characters into pure geometric outlines—lines, curves, and rectangles—without any text operators. This technique appears frequently in:

  • Scanned documents "cleaned" by converting text to outlines
  • Design-focused PDFs from graphic applications
  • Protected documents attempting to prevent simple copy-paste

Since no text operators exist, standard extraction returns empty results.

The Two-Stage Detection and Recovery Strategy

pdf-inspector implements a deliberate pipeline to identify and handle vector-outlined content without unnecessary performance overhead.

Stage 1: Vector-Text Detection

In src/detector.rs (approximately lines 530–560), the library scans each page's content stream and maintains counters for operator types. The heuristic, described in the comment "Whether the page has vector-outlined text (massive path ops, minimal text ops)", evaluates the ratio of path operators to text operators.

When path operations dominate and text operations are scarce, the detector flags the page via the has_vector_text return value. This check is computationally cheap—requiring only operator counting during the initial parse—avoiding expensive OCR on properly encoded text pages.

Stage 2: OCR Fallback Activation

The orchestration logic in src/lib.rs (around line 560, per the comment "Also covers vector-outlined text (glyphs drawn as paths, not…)") checks the detector's output. When needs_ocr evaluates true due to vector-text detection, the standard extraction pipeline is bypassed.

The library then invokes Tesseract OCR through the pdf-ocr crate, rasterizes the page, and performs optical character recognition. The extracted text merges into the PdfProcessResult, with fallback_reason recording that the "vector-outlined text" path was used.

Stage 3: Unified Post-Processing

OCR-derived text flows through identical downstream processing as native text extraction: layout analysis, table detection, and markdown conversion in src/extractor/* and src/markdown/* modules. This ensures output consistency—the final Markdown or JSON structure remains uniform regardless of extraction source.

Key Implementation Files

File Purpose
src/detector.rs Heuristic detection of vector-outlined text via operator counting
src/lib.rs Pipeline orchestration and OCR fallback decision logic
src/extractor/content_stream.rs Native text operator parsing (Tj, TJ, etc.)
src/extractor/fonts.rs Font decoding for genuine text objects
src/markdown/convert.rs Final text-to-Markdown conversion, source-agnostic

Practical Usage Examples

Command-Line Extraction with Auto-Detection


# Standard extraction—vector text detection runs automatically

pdf2md my_report.pdf > output.md

# Force OCR bypassing detection (useful for debugging)

pdf2md --ocr-only my_report.pdf > ocr_output.md

# Inspect detection results including vector-text flags

detect-pdf --json my_report.pdf

Rust API Integration

use pdf_inspector::process_pdf;

fn main() -> Result<(), pdf_inspector::PdfError> {
    // process_pdf handles vector-text detection and OCR fallback transparently
    let result = process_pdf("my_report.pdf")?;
    println!("{}", result.markdown);
    Ok(())
}

Why This Architecture Matters

  • Accuracy: Prevents silent data loss on vector-outlined pages that would otherwise return no extractable text
  • Efficiency: Operator-counting detection avoids OCR overhead on standard text-encoded PDFs
  • Consistency: Single downstream pipeline guarantees uniform Markdown output structure

Summary

  • pdf-inspector detects vector-drawn text in src/detector.rs by counting path operators versus text operators in the PDF content stream
  • Flagged pages trigger automatic OCR fallback in src/lib.rs, using Tesseract via the pdf-ocr crate
  • OCR results merge into the standard PdfProcessResult and flow through identical post-processing as native extraction
  • The approach balances accuracy, performance, and output consistency for documents with mixed or problematic encoding

Frequently Asked Questions

How can I tell if my PDF contains vector-outlined text?

Run detect-pdf --json your_file.pdf to see the detection flags. A has_vector_text: true or fallback_reason: "vector-outlined text" value confirms the page was rendered as geometric paths. You can also inspect the raw PDF—vector text pages show dense path operators (m, l, c, re) with few or no Tj/TJ operators.

Does the OCR fallback work for all languages?

pdf-inspector uses Tesseract OCR, so language support depends on your Tesseract installation and trained data files. Common languages (English, Spanish, French, German, Chinese, Japanese) work reliably with appropriate language packs installed. The library passes through Tesseract's configuration options for language selection.

Is the detection heuristic ever wrong?

The operator-ratio heuristic in src/detector.rs may misclassify pages with extremely complex vector graphics that aren't text. However, false positives trigger unnecessary OCR—preserving text accuracy—while false negatives (missing vector text) are rarer because path-dominant pages almost always represent outlined content. Tuning thresholds is possible via the detector's internal parameters.

Can I disable automatic OCR for vector text?

Currently, pdf-inspector does not expose a flag to disable the vector-text OCR fallback while keeping other auto-detections. Using --ocr-only forces OCR on all pages; for finer control, use the Rust API to implement custom detector logic before calling process_pdf.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →