How LiteParse target_pages Works for Parsing Specific Page Ranges

LiteParse's target_pages option accepts a string such as "1-5,10,15-20", expands it into a sorted, deduplicated Vec<u32> via parse_target_pages, and passes that list to the core extractor so only the requested pages enter the processing pipeline.

The target_pages feature in the run-llama/liteparse repository lets you isolate specific pages from large documents without loading or parsing the entire file. By supplying a human-readable range string through the configuration, you cut memory use and processing time while keeping the output format identical.

Where target_pages Is Defined in LiteParseConfig

The option lives inside LiteParseConfig in crates/liteparse/src/config.rs as an optional String:

pub struct LiteParseConfig {
    // ...
    /// Specific pages to parse (e.g., "1-5,10,15-20").
    pub target_pages: Option<String>,
    // ...
}

When the parser initializes, it reads this field to decide whether to filter pages or process the whole document.

How parse_target_pages Converts the Range String

The heavy lifting is done by the parse_target_pages helper in the same file. The function signature is:

pub fn parse_target_pages(s: &str) -> Result<Vec<u32>, String> {
    // ...split, trim, parse, expand ranges...
}

According to the source, the implementation follows four steps:

  1. Split and trim. The input string is divided on commas, and whitespace is trimmed from each segment.
  2. Detect ranges. A segment containing "-" is treated as an inclusive range; otherwise it is treated as a single page number.
  3. Parse and validate. Each token is converted to a u32. For ranges, the start must be less than or equal to the end.
  4. Expand, sort, and deduplicate. Ranges are expanded into individual page numbers, and the final list is sorted and deduplicated before being returned.

If any token is non-numeric or a range is reversed (for example, "5-3"), the function immediately returns a descriptive Err.

How the Page Filter Is Applied During Extraction

In crates/liteparse/src/parser.rs, the optional string is transformed into the page vector and forwarded to the low-level extractor:

let target_pages = self.target_pages.map(|s| parse_target_pages(s))?;
// ...
self.extract(&document, target_pages.as_deref(), …).await?;

The extract function in crates/liteparse/src/extract.rs receives this optional slice. If target_pages is Some(&[...]), the extractor only processes those pages; otherwise it walks the entire document. All downstream stages—such as rendering, OCR, and output formatting—then operate exclusively on the selected pages.

Using target_pages in CLI and Language Bindings

The same configuration option is exposed consistently across every LiteParse interface.

Command-Line Interface

The CLI flag --target-pages maps directly to LiteParseConfig::target_pages and is documented in docs/src/content/docs/liteparse/cli-reference.md.


# Parse pages 1-5, 10, and 15-20 only

lit parse invoice.pdf --target-pages "1-5,10,15-20" --no-ocr

Node.js / TypeScript

The TypeScript wrapper in packages/node/src/lib.ts defines targetPages as an optional string:

import { LiteParse } from "liteparse";

const parser = new LiteParse({
  ocrEnabled: false,
  targetPages: "1-5,10,15-20",
});

await parser.parse("invoice.pdf");

Python

The Python wrapper in packages/python/liteparse/parser.py exposes the parameter as target_pages:

from liteparse import LiteParse

parser = LiteParse(
    ocr_enabled=False,
    target_pages="1-5,10,15-20"
)

result = parser.parse("invoice.pdf")

Direct Rust Library Call

For Rust developers, set the field directly on LiteParseConfig before instantiating the parser:

use liteparse::{LiteParse, LiteParseConfig};

let cfg = LiteParseConfig {
    target_pages: Some("1-5,10,15-20".to_string()),
    ocr_enabled: false,
    ..Default::default()
};

let parser = LiteParse::new(cfg);
let result = parser.parse("invoice.pdf", None).await?;

Every interface ultimately invokes the same parse_target_pages routine in config.rs, ensuring identical behavior whether you call LiteParse from the shell, JavaScript, Python, or native Rust.

Summary

  • LiteParseConfig in crates/liteparse/src/config.rs declares target_pages as an optional range string.
  • parse_target_pages splits the string on commas, expands hyphenated ranges, validates bounds, and returns a sorted, deduplicated Vec<u32>.
  • parser.rs forwards the parsed list into extract.rs, which gates page processing before any heavy work begins.
  • The feature is available in the CLI (--target-pages), Node.js (targetPages), and Python (target_pages) APIs with the same underlying logic.

Frequently Asked Questions

What string format does target_pages accept?

It accepts comma-separated page numbers and inclusive hyphenated ranges, such as "1-5,10,15-20". The parser trims whitespace around each segment, so "1-5, 10" is also valid.

Does target_pages support overlapping or out-of-order ranges?

Yes. The parse_target_pages function expands all ranges and single pages into a single list, then sorts and deduplicates the result automatically. Overlaps and out-of-order declarations are normalized in the final Vec<u32>.

What happens if I pass an invalid page range to target_pages?

The parser returns a descriptive Err from parse_target_pages. Reversed ranges like "5-3" or non-numeric tokens trigger an error instead of causing unexpected behavior.

Is target_pages available in all LiteParse language bindings?

Yes. The option is exposed in the CLI (--target-pages), the Node.js wrapper in packages/node/src/lib.ts as targetPages, and the Python wrapper in packages/python/liteparse/parser.py as target_pages, all invoking the same Rust core logic.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →