How LiteParse Detects Complex Content with the is_complex Command

The is_complex command in LiteParse performs a fast, pre-OCR analysis of PDF pages to determine if they require optical character recognition by checking text density, image coverage, garbled text detection, and vector outline analysis.

LiteParse, developed by the run-llama/liteparse repository, provides an efficient way to analyze PDF documents without fully rendering them. The is_complex command serves as a critical gatekeeper that identifies which pages need expensive OCR processing versus those that contain usable native text. This lightweight detection system operates in three distinct phases to minimize computational overhead while maximizing accuracy.

The Three-Phase Detection Architecture

The is_complex implementation follows a pipeline designed to avoid expensive operations until absolutely necessary.

Phase 1: Lightweight Document Loading

The entry point in crates/liteparse/src/parser.rs initiates the process by calling extract_pages_and_images with critical performance flags set to false:

let (pages, _) = extract::extract_pages_and_images(
    &document,
    target_pages.as_deref(),
    self.config.max_pages,
    false, // **render_images = false** – images are *not* rasterised
    false, // extract_links = false – hyperlinks are ignored
)?;

This design choice ensures no image rasterization occurs during the initial pass. The extractor only retrieves text items, page dimensions, and image bounds—keeping the operation memory-efficient and fast.

Phase 2: Per-Page Complexity Calculation

Each extracted page is passed to ocr_merge::calculate_page_complexity in crates/liteparse/src/ocr_merge.rs. This function executes five distinct heuristics:

  1. Native text length – Filters out unusable text items and sums the length of valid content
  2. Image analysis – Counts raster images above MIN_IMAGE_SIZE_PT, excluding full-page backgrounds using MAX_IMAGE_PAGE_COVERAGE
  3. Coverage metrics – Calculates text coverage ratio and image coverage relative to page area
  4. Garbled text detection – Identifies broken font encodings via page_is_garbled
  5. Vector outline evaluation – Conditionally checks for text drawn as filled paths using filled_path_bounds when cheaper checks return no results

Phase 3: OCR Decision and Statistics

The function returns a PageComplexityStats struct (defined in crates/liteparse/src/types.rs) containing a boolean needs_ocr field and a vector of ComplexityReason enums explaining the decision. If any heuristic triggers, needs_ocr is set to true.

Core Heuristics in calculate_page_complexity

The calculate_page_complexity function implements specific thresholds to classify content:

Insufficient Native Text (text_length < 20)

  • Triggers ComplexityReason::NoText or ComplexityReason::Scanned (if a full-page image exists)
  • Indicates blank pages or image-only scans

Sparse Text Detection (text_length < 2000 && text_coverage < 0.15)

  • Triggers ComplexityReason::SparseText
  • Typical of scanned documents with minimal selectable text

Embedded Raster Images

  • Triggers ComplexityReason::EmbeddedImages
  • Images must exceed MIN_IMAGE_SIZE_PT and not cover the full page (determined by MAX_IMAGE_PAGE_COVERAGE)

Garbled Text Detection

  • Triggers ComplexityReason::Garbled
  • Uses page_is_garbled to detect broken font encodings producing unreadable strings

Vector-Outline Text (conditional)

  • Triggers ComplexityReason::VectorText
  • Only evaluates when reasons.is_empty() using uncovered_path_area with threshold UNCOVERED_VECTOR_AREA_THRESHOLD
pub(crate) fn calculate_page_complexity(
    page: &Page,
    page_obj: &pdfium::Page,
) -> Result<PageComplexityStats, LiteParseError> {
    // 1️⃣ Text length analysis
    let text_length: usize = page.text_items
        .iter()
        .filter(|item| !is_unusable_native(item))
        .map(|item| item.text.len())
        .sum();

    // 2️⃣ Image analysis with size filtering
    let all_images = page_obj.image_bounds(MIN_IMAGE_SIZE_PT, f32::INFINITY);
    let is_full_page = |b: &ImageBounds| {
        b.width > pw * MAX_IMAGE_PAGE_COVERAGE && b.height > ph * MAX_IMAGE_PAGE_COVERAGE
    };
    let full_page_image = all_images.iter().any(is_full_page);
    let image_bounds: Vec<&ImageBounds> = all_images.iter().filter(|b| !is_full_page(b)).collect();
    let has_images = !image_bounds.is_empty();

    // 3️⃣ Sparse and garbled text detection
    let sparse_text = text_length < 2000 && text_coverage < 0.15;
    let is_garbled = page_is_garbled(page);

    // Build complexity reasons
    let mut reasons = Vec::new();
    if text_length < 20 {
        reasons.push(if full_page_image {
            ComplexityReason::Scanned
        } else {
            ComplexityReason::NoText
        });
    } else if sparse_text {
        reasons.push(ComplexityReason::SparseText);
    }
    if has_images { reasons.push(ComplexityReason::EmbeddedImages); }
    if is_garbled { reasons.push(ComplexityReason::Garbled); }

    // 4️⃣ Expensive vector check only when necessary
    let uncovered_vector_area = if reasons.is_empty() {
        let path_bounds = page_obj.filled_path_bounds(3.0, 0.9);
        let uncovered = uncovered_path_area(&path_bounds, &page.text_items);
        if uncovered >= UNCOVERED_VECTOR_AREA_THRESHOLD {
            reasons.push(ComplexityReason::VectorText);
            Some(uncovered)
        } else {
            None
        }
    } else {
        None
    };

    let needs_ocr = !reasons.is_empty();

    Ok(PageComplexityStats {
        page_number: page.page_number,
        text_length,
        text_coverage,
        has_substantial_images: has_images,
        image_block_count: image_bounds.len(),
        image_coverage,
        largest_image_coverage,
        full_page_image,
        uncovered_vector_area,
        is_garbled,
        page_area,
        needs_ocr,
        reasons,
    })
}

Language Bindings and Usage Examples

All language bindings expose the same Rust implementation through native wrappers.

Rust

use liteparse::LiteParse;
use liteparse::PdfInput;

#[tokio::main]
async fn main() {
    let lp = LiteParse::default();
    let stats = lp.is_complex(PdfInput::Path("document.pdf".into()))
        .await
        .expect("failed to evaluate complexity");
    for s in stats {
        println!("Page {}: needs_ocr={} – reasons: {:?}",
                 s.page_number, s.needs_ocr, s.reasons);
    }
}

Node.js

const { LiteParse } = require('liteparse');

const lp = new LiteParse();
const stats = await lp.is_complex('document.pdf');
console.log(stats);
// Output: [{ pageNumber: 1, needsOcr: true, reasons: ['Scanned'] }, ...]

Python

from liteparse import LiteParse

lp = LiteParse()
stats = lp.is_complex('document.pdf')
for page in stats:
    print(f"Page {page.page_number}: needs_ocr={page.needs_ocr}")

WASM

import * as liteparse from 'liteparse-wasm';

const pdfBytes = new Uint8Array(await fetch('document.pdf').then(r => r.arrayBuffer()));
const stats = await liteparse.is_complex(pdfBytes);

Summary

  • LiteParse detects complex content through a lightweight, multi-heuristic analysis that avoids expensive OCR until necessary.
  • The is_complex command in crates/liteparse/src/parser.rs loads PDFs without rasterizing images, keeping the initial pass fast.
  • calculate_page_complexity in crates/liteparse/src/ocr_merge.rs evaluates text length, coverage ratios, image presence, garbled text, and vector outlines to determine if OCR is required.
  • The system returns a PageComplexityStats struct with a boolean needs_ocr flag and specific ComplexityReason explanations for each page.
  • Available across Rust, Node.js, Python, and WASM bindings with identical behavior and performance characteristics.

Frequently Asked Questions

What is the threshold for determining if a page needs OCR?

LiteParse classifies a page as complex requiring OCR when any of these conditions are met: fewer than 20 characters of usable text (triggering NoText or Scanned), sparse text under 2000 characters with less than 15% coverage (triggering SparseText), presence of substantial embedded images, garbled font encodings, or large uncovered vector path areas. The specific thresholds are defined as constants in crates/liteparse/src/ocr_merge.rs.

Does is_complex render images or perform OCR itself?

No. The is_complex command explicitly passes render_images = false to extract_pages_and_images, ensuring no rasterization occurs. It only analyzes metadata about image bounds, text items, and vector paths. This design keeps the check fast while providing enough information to decide whether subsequent OCR processing is warranted.

How does LiteParse distinguish between scanned documents and digital PDFs?

The system distinguishes these through the is_full_page check combined with text length analysis. If a page contains a full-page image covering more than MAX_IMAGE_PAGE_COVERAGE (typically 90%) of the page area and has fewer than 20 characters of text, it triggers ComplexityReason::Scanned. Digital PDFs with native text typically pass the text length threshold and lack the full-page image marker, resulting in needs_ocr = false.

What are the performance characteristics of the is_complex check?

The is_complex operation is designed to be sub-linear relative to full document processing. Since it skips image rasterization and only examines text items, image bounds, and optionally vector paths (only when cheaper checks pass), it executes significantly faster than full OCR. The vector outline check is conditionally executed only when no other complexity reasons are found, ensuring expensive path calculations are avoided for obviously complex or simple pages.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →