# When Does LiteParse Use Selective OCR Instead of Native PDF Text? 4 Heuristics Explained

> Discover when LiteParse uses selective OCR over native PDF text extraction. Learn the 4 key heuristics for optimizing your document parsing.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: deep-dive
- Published: 2026-06-06

---

**LiteParse runs OCR only on pages that are text-poor, have low text coverage, contain images, or contain garbled native text, preferring native PDF extraction everywhere else.**

LiteParse is an open-source PDF parser from the [run-llama/liteparse](https://github.com/run-llama/liteparse) repository that intelligently balances speed and accuracy by using native PDF text when possible and falling back to OCR only when necessary. Understanding when LiteParse decides to perform selective OCR versus using native PDF text is essential for tuning extraction pipelines and avoiding unnecessary computation. The decision logic lives entirely inside the Rust core, specifically in the [`ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/ocr_merge.rs) module, where four simple heuristics are evaluated for every page.

## How LiteParse Chooses Between Native PDF Text and Selective OCR

The selective OCR check happens during the OCR-pre-render step in `render_pages_for_ocr`. For each page, LiteParse computes four conditions. If **any** condition is true, the page is marked as `needs_ocr` and rendered to a bitmap for OCR processing.

- **Sparse native text** (`text_length < 20`) — The sum of the lengths of all native text items, ignoring garbled fragments. When the total is extremely short (for example, a scanned page with only a logo), LiteParse treats the page as text-poor and sends it to OCR.

- **Low text coverage** (`text_coverage < 0.15`) — The ratio of the total area covered by native-text bounding boxes to the overall page area. When fewer than 15% of the pixels are occupied by text, the page is flagged for OCR.

- **Presence of images** (`has_images`) — If PDFium reports any image bounds on the page, `needs_ocr` is set. This catches embedded bitmaps such as scanned figures or mixed raster/vector documents.

- **Garbled native text** (`page_is_garbled(page)`) — A dedicated heuristic that detects substitution-cipher or broken-cmap text. When the extracted native text is nonsensical, LiteParse abandons it and relies on OCR.

## Core Predicate in [`ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/ocr_merge.rs)

The gate that decides the fate of each page is implemented in [[`crates/liteparse/src/ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr_merge.rs)](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr_merge.rs#L54-L56):

```rust
let needs_ocr =
    text_length < 20 || text_coverage < 0.15 || has_images || page_is_garbled(page);

```

This single expression is evaluated after native text has been extracted but before any bitmap rendering occurs. When `needs_ocr` is true, the page is rasterized and passed to the configured OCR engine. When false, the native items flow through unchanged.

## Merging OCR Results Back into the Page Model

After the OCR engine processes the flagged pages, LiteParse merges the results back into the document model. The merge step is handled by [`ocr_and_merge_rendered`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr_merge.rs#L82-L84) in the same file. Native text that passed the heuristic stays intact, while OCR-derived text is inserted only for the pages that required it. This hybrid approach keeps output fidelity high without rasterizing entire documents unnecessarily.

## Practical Examples: Controlling OCR in LiteParse

You can control whether and how OCR is applied through the TypeScript API. Below are three common patterns.

### Enable Selective OCR (Default)

By default, setting `ocr_enabled: true` lets LiteParse apply the four heuristics and run OCR only on the pages that need it.

```typescript
import { LiteParse } from "liteparse";

const lp = new LiteParse({ ocr_enabled: true });
const result = await lp.parse("report.pdf");

// `result.text` contains native text where possible,
// and OCR-derived text on pages that were text-poor.
console.log(result.text);

```

### Force OCR on Every Page

To bypass the selective guard and run OCR on every page, supply a custom `OcrEngine`. The following example uses a dummy engine to demonstrate the hook, though in practice you would use a real implementation.

```typescript
import { LiteParse, OcrEngine } from "liteparse";

// A dummy engine that pretends to run OCR on every page
class ForceAllOcrEngine implements OcrEngine {
  async recognize(_bytes, _w, _h, _options) {
    return [{ text: "OCRed page", bbox: [0, 0, 100, 100], confidence: 1.0 }];
  }
}

const lp = new LiteParse({ ocr_enabled: true })
  .with_ocr_engine(new ForceAllOcrEngine());

const result = await lp.parse("scanned.pdf");
console.log(result.text);   // OCR text for *all* pages

```

### Disable OCR Completely

When you know a document contains only clean, native text, disabling OCR eliminates any heuristic overhead.

```typescript
import { LiteParse } from "liteparse";

const lp = new LiteParse({ ocr_enabled: false });
const result = await lp.parse("text-only.pdf");

// No OCR is performed; only native PDF text is returned.
console.log(result.text);

```

## Key Source Files Behind the Decision

Three files in the `run-llama/liteparse` repository coordinate the selective OCR behavior:

- [[`crates/liteparse/src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs)](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs) — Orchestrates parsing, checks whether OCR is enabled, instantiates the OCR engine, and invokes the OCR merge step.

- [[`crates/liteparse/src/ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr_merge.rs)](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr_merge.rs) — Contains `render_pages_for_ocr` with the four-heuristic predicate and the `ocr_and_merge_rendered` routine that blends native and OCR text.

- [[`crates/liteparse/src/config.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/config.rs)](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/config.rs) — Defines `LiteParseConfig` fields such as `ocr_enabled`, `ocr_server_url`, and `tessdata_path`, which govern whether the OCR pipeline is active at all.

## Summary

- LiteParse prefers native PDF text and resorts to **selective OCR** only when a page fails one of four heuristics.
- The decision gate is a single boolean expression in [`ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/ocr_merge.rs): sparse text (`< 20` chars), low text coverage (`< 0.15`), embedded images, or garbled glyphs.
- Flagged pages are rasterized and processed by the OCR engine; clean pages pass through untouched.
- The parser exposes `ocr_enabled` and custom engine hooks so developers can opt into selective OCR, force full OCR, or disable it entirely.
- Entry points for this logic are `render_pages_for_ocr` and `ocr_and_merge_rendered` inside [`crates/liteparse/src/ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr_merge.rs).

## Frequently Asked Questions

### Can LiteParse fall back to OCR on a per-page basis?

Yes. According to the `run-llama/liteparse` source code, the OCR-vs-native decision is made independently for each page during `render_pages_for_ocr` in [`ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/ocr_merge.rs). One page can use native text while the next is sent to OCR based on the four heuristics.

### What text coverage threshold triggers OCR in LiteParse?

LiteParse flags a page for OCR when `text_coverage` is below **0.15**, meaning the combined area of native-text bounding boxes covers less than 15% of the total page area. This threshold is hard-coded in the `needs_ocr` predicate inside [`ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/ocr_merge.rs).

### How do I force OCR on every page in LiteParse?

Provide a custom implementation of the `OcrEngine` interface and attach it with `.with_ocr_engine()`. Because LiteParse will invoke your engine for every rendered page, you effectively override the selective guard, as shown in the `ForceAllOcrEngine` example above.

### Where is the selective OCR logic implemented in the LiteParse source code?

The core logic lives in [[`crates/liteparse/src/ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr_merge.rs)](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr_merge.rs), specifically the `render_pages_for_ocr` function that evaluates `text_length`, `text_coverage`, `has_images`, and `page_is_garbled`. The orchestration code that calls this logic is found in [`crates/liteparse/src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs).