# How LiteParse Detects Complex Content with the is_complex Command

> Learn how LiteParse's is_complex command quickly analyzes PDFs for complex content using text density image coverage garbled text detection and vector outline analysis before OCR.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: how-to-guide
- Published: 2026-06-25

---

**The `is_complex` command in LiteParse performs a fast, pre-OCR analysis of PDF pages to determine if they require optical character recognition by checking text density, image coverage, garbled text detection, and vector outline analysis.**

LiteParse, developed by the [run-llama/liteparse](https://github.com/run-llama/liteparse) repository, provides an efficient way to analyze PDF documents without fully rendering them. The `is_complex` command serves as a critical gatekeeper that identifies which pages need expensive OCR processing versus those that contain usable native text. This lightweight detection system operates in three distinct phases to minimize computational overhead while maximizing accuracy.

## The Three-Phase Detection Architecture

The `is_complex` implementation follows a pipeline designed to avoid expensive operations until absolutely necessary.

### Phase 1: Lightweight Document Loading

The entry point in [`crates/liteparse/src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs) initiates the process by calling `extract_pages_and_images` with critical performance flags set to `false`:

```rust
let (pages, _) = extract::extract_pages_and_images(
    &document,
    target_pages.as_deref(),
    self.config.max_pages,
    false, // **render_images = false** – images are *not* rasterised
    false, // extract_links = false – hyperlinks are ignored
)?;

```

This design choice ensures **no image rasterization** occurs during the initial pass. The extractor only retrieves text items, page dimensions, and image bounds—keeping the operation memory-efficient and fast.

### Phase 2: Per-Page Complexity Calculation

Each extracted page is passed to `ocr_merge::calculate_page_complexity` in [`crates/liteparse/src/ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr_merge.rs). This function executes five distinct heuristics:

1. **Native text length** – Filters out unusable text items and sums the length of valid content
2. **Image analysis** – Counts raster images above `MIN_IMAGE_SIZE_PT`, excluding full-page backgrounds using `MAX_IMAGE_PAGE_COVERAGE`
3. **Coverage metrics** – Calculates text coverage ratio and image coverage relative to page area
4. **Garbled text detection** – Identifies broken font encodings via `page_is_garbled`
5. **Vector outline evaluation** – Conditionally checks for text drawn as filled paths using `filled_path_bounds` when cheaper checks return no results

### Phase 3: OCR Decision and Statistics

The function returns a `PageComplexityStats` struct (defined in [`crates/liteparse/src/types.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/types.rs)) containing a boolean `needs_ocr` field and a vector of `ComplexityReason` enums explaining the decision. If any heuristic triggers, `needs_ocr` is set to `true`.

## Core Heuristics in calculate_page_complexity

The `calculate_page_complexity` function implements specific thresholds to classify content:

**Insufficient Native Text** (`text_length < 20`)
- Triggers `ComplexityReason::NoText` or `ComplexityReason::Scanned` (if a full-page image exists)
- Indicates blank pages or image-only scans

**Sparse Text Detection** (`text_length < 2000 && text_coverage < 0.15`)
- Triggers `ComplexityReason::SparseText`
- Typical of scanned documents with minimal selectable text

**Embedded Raster Images**
- Triggers `ComplexityReason::EmbeddedImages`
- Images must exceed `MIN_IMAGE_SIZE_PT` and not cover the full page (determined by `MAX_IMAGE_PAGE_COVERAGE`)

**Garbled Text Detection**
- Triggers `ComplexityReason::Garbled`
- Uses `page_is_garbled` to detect broken font encodings producing unreadable strings

**Vector-Outline Text** (conditional)
- Triggers `ComplexityReason::VectorText`
- Only evaluates when `reasons.is_empty()` using `uncovered_path_area` with threshold `UNCOVERED_VECTOR_AREA_THRESHOLD`

```rust
pub(crate) fn calculate_page_complexity(
    page: &Page,
    page_obj: &pdfium::Page,
) -> Result<PageComplexityStats, LiteParseError> {
    // 1️⃣ Text length analysis
    let text_length: usize = page.text_items
        .iter()
        .filter(|item| !is_unusable_native(item))
        .map(|item| item.text.len())
        .sum();

    // 2️⃣ Image analysis with size filtering
    let all_images = page_obj.image_bounds(MIN_IMAGE_SIZE_PT, f32::INFINITY);
    let is_full_page = |b: &ImageBounds| {
        b.width > pw * MAX_IMAGE_PAGE_COVERAGE && b.height > ph * MAX_IMAGE_PAGE_COVERAGE
    };
    let full_page_image = all_images.iter().any(is_full_page);
    let image_bounds: Vec<&ImageBounds> = all_images.iter().filter(|b| !is_full_page(b)).collect();
    let has_images = !image_bounds.is_empty();

    // 3️⃣ Sparse and garbled text detection
    let sparse_text = text_length < 2000 && text_coverage < 0.15;
    let is_garbled = page_is_garbled(page);

    // Build complexity reasons
    let mut reasons = Vec::new();
    if text_length < 20 {
        reasons.push(if full_page_image {
            ComplexityReason::Scanned
        } else {
            ComplexityReason::NoText
        });
    } else if sparse_text {
        reasons.push(ComplexityReason::SparseText);
    }
    if has_images { reasons.push(ComplexityReason::EmbeddedImages); }
    if is_garbled { reasons.push(ComplexityReason::Garbled); }

    // 4️⃣ Expensive vector check only when necessary
    let uncovered_vector_area = if reasons.is_empty() {
        let path_bounds = page_obj.filled_path_bounds(3.0, 0.9);
        let uncovered = uncovered_path_area(&path_bounds, &page.text_items);
        if uncovered >= UNCOVERED_VECTOR_AREA_THRESHOLD {
            reasons.push(ComplexityReason::VectorText);
            Some(uncovered)
        } else {
            None
        }
    } else {
        None
    };

    let needs_ocr = !reasons.is_empty();

    Ok(PageComplexityStats {
        page_number: page.page_number,
        text_length,
        text_coverage,
        has_substantial_images: has_images,
        image_block_count: image_bounds.len(),
        image_coverage,
        largest_image_coverage,
        full_page_image,
        uncovered_vector_area,
        is_garbled,
        page_area,
        needs_ocr,
        reasons,
    })
}

```

## Language Bindings and Usage Examples

All language bindings expose the same Rust implementation through native wrappers.

### Rust

```rust
use liteparse::LiteParse;
use liteparse::PdfInput;

#[tokio::main]
async fn main() {
    let lp = LiteParse::default();
    let stats = lp.is_complex(PdfInput::Path("document.pdf".into()))
        .await
        .expect("failed to evaluate complexity");
    for s in stats {
        println!("Page {}: needs_ocr={} – reasons: {:?}",
                 s.page_number, s.needs_ocr, s.reasons);
    }
}

```

### Node.js

```javascript
const { LiteParse } = require('liteparse');

const lp = new LiteParse();
const stats = await lp.is_complex('document.pdf');
console.log(stats);
// Output: [{ pageNumber: 1, needsOcr: true, reasons: ['Scanned'] }, ...]

```

### Python

```python
from liteparse import LiteParse

lp = LiteParse()
stats = lp.is_complex('document.pdf')
for page in stats:
    print(f"Page {page.page_number}: needs_ocr={page.needs_ocr}")

```

### WASM

```javascript
import * as liteparse from 'liteparse-wasm';

const pdfBytes = new Uint8Array(await fetch('document.pdf').then(r => r.arrayBuffer()));
const stats = await liteparse.is_complex(pdfBytes);

```

## Summary

- **LiteParse** detects complex content through a lightweight, multi-heuristic analysis that avoids expensive OCR until necessary.
- The **`is_complex`** command in [`crates/liteparse/src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs) loads PDFs without rasterizing images, keeping the initial pass fast.
- **`calculate_page_complexity`** in [`crates/liteparse/src/ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr_merge.rs) evaluates text length, coverage ratios, image presence, garbled text, and vector outlines to determine if OCR is required.
- The system returns a **`PageComplexityStats`** struct with a boolean `needs_ocr` flag and specific `ComplexityReason` explanations for each page.
- Available across **Rust**, **Node.js**, **Python**, and **WASM** bindings with identical behavior and performance characteristics.

## Frequently Asked Questions

### What is the threshold for determining if a page needs OCR?

LiteParse classifies a page as complex requiring OCR when any of these conditions are met: fewer than 20 characters of usable text (triggering `NoText` or `Scanned`), sparse text under 2000 characters with less than 15% coverage (triggering `SparseText`), presence of substantial embedded images, garbled font encodings, or large uncovered vector path areas. The specific thresholds are defined as constants in [`crates/liteparse/src/ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr_merge.rs).

### Does is_complex render images or perform OCR itself?

No. The `is_complex` command explicitly passes `render_images = false` to `extract_pages_and_images`, ensuring no rasterization occurs. It only analyzes metadata about image bounds, text items, and vector paths. This design keeps the check fast while providing enough information to decide whether subsequent OCR processing is warranted.

### How does LiteParse distinguish between scanned documents and digital PDFs?

The system distinguishes these through the `is_full_page` check combined with text length analysis. If a page contains a full-page image covering more than `MAX_IMAGE_PAGE_COVERAGE` (typically 90%) of the page area and has fewer than 20 characters of text, it triggers `ComplexityReason::Scanned`. Digital PDFs with native text typically pass the text length threshold and lack the full-page image marker, resulting in `needs_ocr = false`.

### What are the performance characteristics of the is_complex check?

The `is_complex` operation is designed to be **sub-linear** relative to full document processing. Since it skips image rasterization and only examines text items, image bounds, and optionally vector paths (only when cheaper checks pass), it executes significantly faster than full OCR. The vector outline check is conditionally executed only when no other complexity reasons are found, ensuring expensive path calculations are avoided for obviously complex or simple pages.