How to Specify a Range of Pages for Extraction Using the `pages` Option
Use a comma-separated string combining individual pages (e.g., 1,3) and hyphenated ranges (e.g., 5-7) with the --pages CLI flag or pages API property to control which PDF pages are processed.
The opendataloader-pdf library provides granular control over PDF processing through its flexible pages option. Whether you need to extract content from specific chapters or skip unnecessary front matter, you can specify a range of pages for extraction using a simple comma-separated syntax. This functionality is available across the CLI, Node.js, and Java APIs, with consistent parsing logic implemented in the core configuration class.
Understanding the pages Syntax
The pages option accepts a comma-separated string that supports mixed notation for individual pages and continuous ranges. When provided, this string is parsed into a list of integers that determines exactly which pages are handed to the PDF extractors. The CLI option is documented in the auto-generated reference table at content/docs/cli-options-reference.mdx.
Supported Format Patterns
- Single pages: Specify individual pages using 1-based indexing (e.g.,
1,2,10). - Continuous ranges: Use hyphenated notation to define inclusive ranges (e.g.,
5-7extracts pages 5, 6, and 7). - Mixed specifications: Combine singles and ranges in one string (e.g.,
1,3,5-7extracts pages 1, 3, 5, 6, and 7).
If the pages option is omitted or set to an empty string, the processing pipeline extracts all pages from the document.
How Page Ranges Are Parsed
According to the opendataloader-pdf source code, the parsing logic resides in Config.java at java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/Config.java. The parsePageRanges method (lines 648-690) implements the following validation and expansion algorithm:
- Tokenization: The input string is split on commas.
- Range detection: Each token is trimmed and checked for the presence of
-. - Range expansion: If a hyphen is found,
parseRangevalidates that both numbers are positive and that the start value is less than or equal to the end value, then adds every integer in that interval to the result list. - Single page validation: Tokens without hyphens are processed by
parseSinglePage, which validates the token as a positive integer. - Caching: The final list is stored in
Config.cachedPageNumbersfor use by the processing pipeline.
Invalid formats trigger an IllegalArgumentException with a descriptive message indicating the problematic format.
Usage Examples
CLI Usage
Pass the --pages flag with a quoted string to extract specific pages:
# Extract pages 1, 3, and 5 through 7 as JSON and Markdown
opendataloader-pdf document.pdf \
--format json,markdown \
--pages "1,3,5-7" \
-o ./output
Node.js API
When using the TypeScript/JavaScript wrapper, provide the pages property in the options object:
import { convert } from 'opendataloader-pdf';
await convert('document.pdf', {
format: ['json', 'markdown'],
pages: '2,4-6',
outputDir: './output'
});
The TypeScript definition in node/opendataloader-pdf/src/convert-options.generated.ts declares this property as pages?: string;, mirroring the Java implementation.
Java API
Programmatically configure page ranges using the Config class:
import org.opendataloader.pdf.api.Config;
import org.opendataloader.pdf.core.PdfProcessor;
Config config = new Config();
config.setPages("1,3,5-7");
PdfProcessor.process("document.pdf", config);
Error Handling and Validation
The parser strictly validates input to prevent processing errors. Providing malformed ranges or non-numeric values results in immediate feedback.
try {
config.setPages("1,abc,5-3"); // Invalid: non-numeric and reversed range
} catch (IllegalArgumentException e) {
System.err.println("Invalid page spec: " + e.getMessage());
}
As implemented in Config.java, the validator ensures all page numbers are positive integers and that range start values do not exceed their end values.
Summary
- The
pagesoption accepts a comma-separated string supporting individual pages (1,3) and inclusive ranges (5-7). - Syntax validation and expansion occur in
Config.parsePageRangeswithinjava/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/Config.java. - The parser uses 1-based indexing and throws
IllegalArgumentExceptionfor invalid formats or reversed ranges. - Omitting the option causes the pipeline to process all pages by default.
- The same syntax works across CLI (
--pages), Node.js (pagesproperty), and Java (setPagesmethod) interfaces.
Frequently Asked Questions
What happens if I omit the pages option?
If you omit the pages option or provide an empty string, opendataloader-pdf processes all pages in the PDF document. The Config.cachedPageNumbers list remains unset, causing the pipeline to skip page filtering and extract the entire document.
Can I use zero-based indexing for page numbers?
No. The opendataloader-pdf parser uses 1-based indexing as implemented in Config.parseSinglePage and Config.parseRange. Specifying page 0 or negative integers triggers an IllegalArgumentException during validation.
How do I extract non-contiguous pages?
Use comma-separated values to specify individual pages and ranges in any order. For example, "1,3,5-7,10" extracts pages 1, 3, 5, 6, 7, and 10, skipping all others. The parser expands these into a deduplicated list stored in Config.cachedPageNumbers.
What error occurs with invalid page range formats?
Invalid formats throw an IllegalArgumentException with a message identifying the specific error, such as non-numeric tokens (e.g., "abc"), reversed ranges (e.g., "5-3"), or malformed syntax. This validation occurs immediately when calling setPages() or equivalent API methods, preventing invalid configurations from reaching the extraction pipeline.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →