# How pdf-inspector Handles Vector Text (Path-Operator Text) Versus Actual Text in PDFs

> Discover how pdf-inspector differentiates vector text from actual text by analyzing path and text operators. It automatically uses OCR for geometric path text.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: internals
- Published: 2026-08-06

---

**pdf-inspector detects vector-drawn text by counting path operators versus text operators in the PDF content stream, then automatically falls back to OCR when glyphs are rendered as geometric paths rather than actual text objects.**

PDF documents store text in two fundamentally different ways: as **actual text objects** using standard PDF text operators (`Tj`, `TJ`, `Td`, `Tm`), or as **vector-drawn glyphs** rendered through geometric path operators (`m`, `l`, `c`, `re`). The firecrawl/pdf-inspector library distinguishes between these approaches using a purpose-built detection heuristic, ensuring reliable text extraction regardless of how the content was originally encoded.

## How Vector-Drawn Text Differs from Actual Text

**Actual text** in PDFs uses the `Show` operator family to display character codes mapped to fonts. This allows direct extraction by reading the content stream and decoding the glyphs.

**Vector-drawn text** converts characters into pure geometric outlines—lines, curves, and rectangles—without any text operators. This technique appears frequently in:

- Scanned documents "cleaned" by converting text to outlines
- Design-focused PDFs from graphic applications
- Protected documents attempting to prevent simple copy-paste

Since no text operators exist, standard extraction returns empty results.

## The Two-Stage Detection and Recovery Strategy

pdf-inspector implements a deliberate pipeline to identify and handle vector-outlined content without unnecessary performance overhead.

### Stage 1: Vector-Text Detection

In [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) (approximately lines 530–560), the library scans each page's content stream and maintains counters for operator types. The heuristic, described in the comment "Whether the page has vector-outlined text (massive path ops, minimal text ops)", evaluates the ratio of **path operators** to **text operators**.

When path operations dominate and text operations are scarce, the detector flags the page via the `has_vector_text` return value. This check is computationally cheap—requiring only operator counting during the initial parse—avoiding expensive OCR on properly encoded text pages.

### Stage 2: OCR Fallback Activation

The orchestration logic in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) (around line 560, per the comment "Also covers vector-outlined text (glyphs drawn as paths, not…)") checks the detector's output. When `needs_ocr` evaluates true due to vector-text detection, the standard extraction pipeline is bypassed.

The library then invokes Tesseract OCR through the `pdf-ocr` crate, rasterizes the page, and performs optical character recognition. The extracted text merges into the `PdfProcessResult`, with `fallback_reason` recording that the "vector-outlined text" path was used.

### Stage 3: Unified Post-Processing

OCR-derived text flows through identical downstream processing as native text extraction: layout analysis, table detection, and markdown conversion in `src/extractor/*` and `src/markdown/*` modules. This ensures **output consistency**—the final Markdown or JSON structure remains uniform regardless of extraction source.

## Key Implementation Files

| File | Purpose |
|------|---------|
| [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) | Heuristic detection of vector-outlined text via operator counting |
| [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) | Pipeline orchestration and OCR fallback decision logic |
| [`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs) | Native text operator parsing (`Tj`, `TJ`, etc.) |
| [`src/extractor/fonts.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/fonts.rs) | Font decoding for genuine text objects |
| [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) | Final text-to-Markdown conversion, source-agnostic |

## Practical Usage Examples

### Command-Line Extraction with Auto-Detection

```bash

# Standard extraction—vector text detection runs automatically

pdf2md my_report.pdf > output.md

# Force OCR bypassing detection (useful for debugging)

pdf2md --ocr-only my_report.pdf > ocr_output.md

# Inspect detection results including vector-text flags

detect-pdf --json my_report.pdf

```

### Rust API Integration

```rust
use pdf_inspector::process_pdf;

fn main() -> Result<(), pdf_inspector::PdfError> {
    // process_pdf handles vector-text detection and OCR fallback transparently
    let result = process_pdf("my_report.pdf")?;
    println!("{}", result.markdown);
    Ok(())
}

```

## Why This Architecture Matters

- **Accuracy**: Prevents silent data loss on vector-outlined pages that would otherwise return no extractable text
- **Efficiency**: Operator-counting detection avoids OCR overhead on standard text-encoded PDFs
- **Consistency**: Single downstream pipeline guarantees uniform Markdown output structure

## Summary

- pdf-inspector detects vector-drawn text in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) by counting path operators versus text operators in the PDF content stream
- Flagged pages trigger automatic OCR fallback in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), using Tesseract via the `pdf-ocr` crate
- OCR results merge into the standard `PdfProcessResult` and flow through identical post-processing as native extraction
- The approach balances accuracy, performance, and output consistency for documents with mixed or problematic encoding

## Frequently Asked Questions

### How can I tell if my PDF contains vector-outlined text?

Run `detect-pdf --json your_file.pdf` to see the detection flags. A `has_vector_text: true` or `fallback_reason: "vector-outlined text"` value confirms the page was rendered as geometric paths. You can also inspect the raw PDF—vector text pages show dense path operators (`m`, `l`, `c`, `re`) with few or no `Tj`/`TJ` operators.

### Does the OCR fallback work for all languages?

pdf-inspector uses Tesseract OCR, so language support depends on your Tesseract installation and trained data files. Common languages (English, Spanish, French, German, Chinese, Japanese) work reliably with appropriate language packs installed. The library passes through Tesseract's configuration options for language selection.

### Is the detection heuristic ever wrong?

The operator-ratio heuristic in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) may misclassify pages with extremely complex vector graphics that aren't text. However, false positives trigger unnecessary OCR—preserving text accuracy—while false negatives (missing vector text) are rarer because path-dominant pages almost always represent outlined content. Tuning thresholds is possible via the detector's internal parameters.

### Can I disable automatic OCR for vector text?

Currently, pdf-inspector does not expose a flag to disable the vector-text OCR fallback while keeping other auto-detections. Using `--ocr-only` forces OCR on all pages; for finer control, use the Rust API to implement custom detector logic before calling `process_pdf`.