How to Handle Large PDF Documents with LiteParse's max_pages Limit

LiteParse caps PDF parsing at 1000 pages by default to prevent memory exhaustion, but you can override this limit via the --max-pages CLI flag, the LiteParseConfig struct in Rust, or equivalent options in the Node.js and Python bindings.

LiteParse, the high-performance PDF parsing library from run-llama, includes a built-in safeguard that limits extraction to 1000 pages by default. This max_pages ceiling prevents uncontrolled memory usage when processing massive documents, but understanding how to adjust it is essential for handling large PDF documents with LiteParse's max_pages limit in production workflows. Whether you are extracting text from thousand-page reports or archival books, the configuration options in crates/liteparse/src/config.rs provide the flexibility to scale beyond the default boundary.

Understanding the Default 1000-Page Limit

In crates/liteparse/src/config.rs, the LiteParseConfig struct defines max_pages with a default value of 1000. This guard is enforced in the core extraction routine extract_pages_and_images located in crates/liteparse/src/parser.rs (lines 28-33), where the parser stops extracting additional pages once the internal counter reaches the configured limit.

Three Methods to Increase the Page Limit

You can raise the max_pages threshold through three interfaces depending on your integration approach.

Command-Line Interface (--max-pages)

The CLI exposes the limit via the --max-pages flag defined in crates/liteparse/src/main.rs (lines 73-75).

liteparse parse my-big-file.pdf --max-pages 5000

Rust Programmatic Configuration

When constructing the parser directly in Rust, modify the LiteParseConfig before instantiation.

use liteparse::LiteParse;
use liteparse::LiteParseConfig;

let mut cfg = LiteParseConfig::default();
cfg.max_pages = 5_000;  // Lift the limit to 5,000 pages
let parser = LiteParse::new(cfg);
let result = parser.parse("my-big-file.pdf").await?;

Language Bindings (Node.js and Python)

The same option propagates through the Node (NAPI) and Python wrappers.

Node.js (TypeScript):

import { LiteParse } from "liteparse";

const lp = new LiteParse({ max_pages: 5000 });
const res = await lp.parse("my-big-file.pdf");

Python:

from liteparse import LiteParse

lp = LiteParse(max_pages=5000)
result = lp.parse("my-big-file.pdf")

Memory Considerations and Performance Impact

Raising max_pages directly increases RAM consumption because LiteParse loads each page’s text items, raster images (when OCR is active), and metadata into memory before grid projection. While the default 1000-page bound is deliberately generous, setting values above 10,000 may cause out-of-memory (OOM) errors on modest hardware. Note that this limit operates independently of the OCR worker pool (num_workers) and DPI settings, which remain separate performance tuning knobs.

Production Code Examples

Rust Implementation with 8,000-Page Limit

use liteparse::{LiteParse, LiteParseConfig};

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    let mut cfg = LiteParseConfig::default();
    cfg.max_pages = 8_000;      // Allow up to 8,000 pages
    cfg.ocr_enabled = false;    // Optional: disable OCR to conserve RAM

    let parser = LiteParse::new(cfg);
    let result = parser.parse("huge-document.pdf").await?;

    println!("Parsed {} pages", result.pages.len());
    Ok(())
}

Node.js TypeScript for Large Documents

import { LiteParse } from "liteparse";

(async () => {
  const lp = new LiteParse({ 
    max_pages: 10_000, 
    ocr_enabled: false 
  });
  const parsed = await lp.parse("huge-document.pdf");
  console.log(`Extracted ${parsed.pages.length} pages`);
})();

Python Command-Line Shortcut

python -m liteparse parse huge-document.pdf --max-pages 12000 --no-ocr

Summary

  • Default protection: LiteParse limits extraction to 1000 pages via LiteParseConfig::max_pages in crates/liteparse/src/config.rs to prevent memory exhaustion.
  • Override methods: Use the --max-pages CLI flag, modify LiteParseConfig in Rust, or pass max_pages to the Node.js or Python constructors.
  • Memory awareness: Increasing the limit consumes more RAM for text items and images; disable OCR (ocr_enabled = false) to reduce footprint when processing massive documents.
  • Source enforcement: The limit is applied in crates/liteparse/src/parser.rs and crates/liteparse/src/extract.rs during the page vector construction.

Frequently Asked Questions

What is the default max_pages limit in LiteParse?

LiteParse defaults to 1000 pages as defined in crates/liteparse/src/config.rs. This prevents uncontrolled memory usage when parsing unexpectedly large documents.

How do I parse a PDF with more than 1000 pages using the CLI?

Use the --max-pages flag followed by your desired limit. For example: liteparse parse document.pdf --max-pages 5000. This flag is defined in crates/liteparse/src/main.rs (lines 73-75).

Does increasing max_pages affect OCR performance?

No, the max_pages limit is independent of the OCR worker pool (num_workers) and DPI settings. However, enabling OCR with a high max_pages value significantly increases memory consumption because raster images are loaded into memory alongside text items.

Where is the max_pages limit enforced in the source code?

The limit is enforced in crates/liteparse/src/parser.rs (lines 28-33) where the configuration passes the value to the extraction routine, and in crates/liteparse/src/extract.rs where the parser stops building the Vec<Page> once the counter reaches the threshold.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →