How to Parse Specific Page Ranges in PDF Documents Using LiteParse
LiteParse parses specific page ranges in PDF documents by accepting a comma-separated string of single pages and ranges (e.g., "1-5,10,15-20") via the --target-pages CLI flag or targetPages/target_pages options in its language bindings, which internally expands into a filtered vector to process only selected pages.
Parsing large PDF documents can be resource-intensive when you only need specific sections. The run-llama/liteparse library provides a built-in mechanism to parse specific page ranges in PDF documents, allowing you to extract content from only the pages you need while skipping the rest.
How Page Range Parsing Works Internally
LiteParse implements page range filtering through a three-stage pipeline that converts user input into optimized extraction instructions.
Configuration Parsing in config.rs
The process begins in crates/liteparse/src/config.rs, where the LiteParseConfig::parse_target_pages function (lines 9-53) handles the initial string parsing. This function accepts a comma-separated list where each entry is either a single page number or a range expressed as start-end. It expands these ranges into a Vec<u32> and stores them in the configuration struct for later use.
Range Resolution in parser.rs
When the parser executes, LiteParse::resolve_target_pages in crates/liteparse/src/parser.rs (lines 81-110) reads the parsed page vector from the configuration. This method validates the requested pages against the document's actual page count and prepares the final list for extraction.
Low-Level Extraction in extract.rs
The actual filtering occurs in crates/liteparse/src/extract.rs (lines 30-77) within the extract_pages_from_document function. This low-level component receives the optional page slice and loads only the specified pages from the PDFium document, ignoring all others to save memory and CPU time.
Syntax for Specifying Page Ranges
LiteParse uses a consistent string format across all bindings:
- Single pages: Specify individual page numbers as integers (e.g.,
3,7,12) - Inclusive ranges: Use hyphen-separated start and end values (e.g.,
1-5,15-20) - Mixed specifications: Combine singles and ranges in a comma-separated list (e.g.,
1-5,10,15-20)
Page numbers are 1-indexed and the ranges are inclusive of both endpoints.
Implementation Examples Across Languages
Command Line Interface
The Rust CLI exposes the functionality via the --target-pages argument in crates/liteparse/src/main.rs (lines 239-358):
# Parse pages 1-5, page 10, and pages 15-20
lit parse document.pdf --target-pages "1-5,10,15-20" --format json -o subset.json
Node.js and TypeScript
The Node.js binding exposes targetPages as a string option in packages/node/src/lib.ts (lines 48-60):
import { LiteParse } from '@llamaindex/liteparse';
const parser = new LiteParse({
outputFormat: 'markdown',
targetPages: '2-4,7,9-12'
});
const result = await parser.parse('report.pdf');
console.log(result.text);
Python
The Python binding accepts target_pages as an optional string argument in packages/python/liteparse/parser.py (lines 44-58):
from liteparse import LiteParse
parser = LiteParse(
output_format='json',
target_pages='3,5-8'
)
result = parser.parse('invoice.pdf')
print(result['text'])
WebAssembly (Browser)
The WASM binding exposes target_pages as a configuration field in crates/liteparse-wasm/src/lib.rs (lines 45-78):
import init, { LiteParse } from '@llamaindex/liteparse-wasm';
await init();
const parser = new LiteParse({
target_pages: '1-3,6'
});
const pdfBytes = await fetch('file.pdf').then(r => r.arrayBuffer());
const result = await parser.parse(pdfBytes);
console.log(result.text);
Safety Limits and Validation
To prevent accidental memory exhaustion attacks, LiteParse enforces a hard-coded safety limit defined as MAX_TARGET_PAGES = 100_000 in crates/liteparse/src/config.rs (lines 100-107). If your range string expands to more than 100,000 individual pages, the parser will reject the configuration before attempting extraction.
Summary
- LiteParse filters PDF pages by accepting comma-separated range strings via the
--target-pagesCLI flag ortargetPages/target_pagesoptions in language bindings. - The string parser expands ranges in
config.rs, resolves them inparser.rs, and executes selective extraction inextract.rs. - Supported syntax includes single pages (
5), ranges (1-10), and mixed combinations (1-5,10,15-20). - A safety limit of 100,000 pages prevents accidental resource exhaustion.
- All bindings (CLI, Node.js, Python, and WASM) implement identical range specification semantics.
Frequently Asked Questions
What is the correct format for specifying page ranges in LiteParse?
LiteParse accepts a comma-separated string where each element is either a single 1-indexed page number or an inclusive range using a hyphen (e.g., "1-5,10,15-20"). This format is consistent across the CLI (--target-pages), Node.js (targetPages), Python (target_pages), and WASM (target_pages) interfaces.
Does LiteParse support parsing specific page ranges in all language bindings?
Yes, according to the source code in run-llama/liteparse, the target_pages functionality is exposed in every official binding: the Rust CLI (main.rs), Node.js/TypeScript (lib.ts), Python (parser.py), and WebAssembly (lib.rs). Each binding passes the range string to the same underlying Rust parser for consistent behavior.
What happens if I specify a page number that exceeds the PDF's total pages?
The LiteParse::resolve_target_pages function in parser.rs validates the requested pages against the document's actual page count. If you specify pages beyond the document's length, the parser will handle the error appropriately, typically by ignoring out-of-range pages or raising a validation error depending on the specific binding's error handling implementation.
Is there a performance benefit to using page range filtering?
Yes, by specifying target_pages, you prevent LiteParse from loading the entire PDF into memory. The extract_pages_from_document function in extract.rs only requests the specific pages from the underlying PDFium library, significantly reducing memory usage and processing time for large documents when you only need a subset of pages.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →