How to Parse Specific Page Ranges from Documents Using LiteParse
LiteParse enables efficient parsing of specific page ranges from documents through the target_pages configuration option, which accepts compact range strings like "1-5,10,15-20" and restricts processing to only those pages across CLI, Node.js, Python, and WASM interfaces.
Parsing entire multi-page documents when you only need a subset wastes computational resources and time. The run-llama/liteparse repository provides a targeted solution through its page-range filtering system, allowing you to parse specific page ranges from documents using LiteParse without loading unnecessary content. This capability is implemented consistently across all language bindings through the core LiteParseConfig structure.
How Page Range Filtering Works in LiteParse
The architecture centers on three core components that transform a range string into optimized page extraction.
Configuration and Range Parsing
The LiteParseConfig struct defined in crates/liteparse/src/config.rs (lines 14-18) stores the target_pages field as an optional string. When parsing begins, the parse_target_pages function (lines 66-96) processes this string by splitting on commas, expanding hyphenated ranges into individual page numbers, trimming whitespace, and returning a sorted, deduplicated Vec<u32>.
Core Parser Integration
In crates/liteparse/src/parser.rs (lines 91-100), the LiteParse::parse method reads the config.target_pages value, invokes parse_target_pages, and forwards the resulting slice to the extraction layer. This ensures that only the specified page indices reach the document processing pipeline.
Selective Extraction Layer
The extraction functions extract_pages_from_document and extract_pages_from_input in crates/liteparse/src/extract.rs receive the validated page list as Option<&[u32]>. When provided, these functions load only the requested pages from the PDFium document, drastically reducing I/O operations and OCR workload.
Parsing Page Ranges via CLI
The CLI implements this through the ParseCommand struct in crates/liteparse/src/main.rs (lines 65-71), which defines target_pages as Option<String>.
liteparse parse report.pdf \
--target-pages "1-3,7,10-12" \
--max-pages 10 \
--format json \
--output selected_pages.json
The --target-pages flag accepts comma-separated values and hyphenated ranges. The optional --max-pages parameter acts as a hard ceiling that prevents processing beyond the specified limit, even if the range string includes more pages.
Programmatic Usage in Node.js and Python
Both the Node.js and Python bindings serialize the configuration to the same LiteParseConfig structure used by the Rust core.
Node.js / TypeScript Implementation
import { LiteParse } from "liteparse";
const parser = new LiteParse({
target_pages: "2-4,8",
max_pages: 5,
ocr_enabled: false,
});
await parser.parse("report.pdf", { output: "out.json", format: "json" });
Python Implementation
from liteparse import LiteParse
parser = LiteParse(
target_pages="5,9-11",
max_pages=7,
ocr_enabled=False,
)
result = parser.parse("report.pdf", format="json")
Optimizing Performance with Page Ranges
To maximize efficiency when you parse specific page ranges from documents using LiteParse, combine the target_pages option with these strategies:
- Enable
max_pagesas a safety guard: Set this to a reasonable upper bound to prevent accidental resource exhaustion from oversized ranges. - Disable OCR when unnecessary: Set
ocr_enabledtofalse(or use--no-ocrin CLI) to skip optical character recognition on non-scanned documents. - Use contiguous range syntax: Prefer
"1-100"over listing individual pages, as the parser normalizes ranges internally but compact strings parse faster. - Leverage screenshot targeting: The screenshot command respects the same range logic via
parse_target_pages(seemain.rslines 42-48):
liteparse screenshot report.pdf --target-pages "1,3,5" --output-dir pages/
Summary
- LiteParse restricts parsing to specific subsets via the
target_pagesconfiguration option, accepting strings in the format"1-5,10,15-20". - The
parse_target_pagesfunction incrates/liteparse/src/config.rsvalidates, expands, and deduplicates range strings into sortedVec<u32>vectors. - The extraction layer in
crates/liteparse/src/extract.rsloads only requested pages, minimizing I/O and processing overhead. - CLI, Node.js, and Python interfaces all map to the same
LiteParseConfigstructure for consistent behavior. - Combining
target_pageswithmax_pagesand disabling OCR when unneeded provides maximum performance.
Frequently Asked Questions
What syntax does LiteParse use for specifying page ranges?
LiteParse accepts a compact comma-separated string where hyphens denote inclusive ranges. For example, "1-3,7,10-12" selects pages 1, 2, 3, 7, 10, 11, and 12. The parse_target_pages function in crates/liteparse/src/config.rs handles the expansion and validation.
Does LiteParse support parsing non-contiguous page ranges?
Yes. You can specify disconnected ranges like "1-5,10,15-20" in the target_pages string. The parser expands these into a deduplicated, sorted vector before extraction, ensuring each requested page is processed exactly once regardless of how the ranges overlap in your input string.
How does LiteParse handle page ranges that exceed the document length?
If the target_pages string includes page numbers beyond the document's actual page count, LiteParse processes only the valid pages that exist. The max_pages configuration option provides an additional ceiling to prevent processing more than a specified number of pages, acting as a safeguard against pathological inputs.
Can I use page range filtering when generating screenshots?
Yes. The screenshot command reuses the same parse_target_pages logic defined in crates/liteparse/src/main.rs (lines 42-48). Pass the --target-pages flag to render only specific pages: liteparse screenshot doc.pdf --target-pages "1,3,5".
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →