# How to Specify a Range of Pages for Extraction Using the `pages` Option

> Specify a page range for PDF extraction using the pages option. Use comma-separated values and hyphenated ranges with the --pages CLI flag or pages API property to control page processing.

- Repository: [opendataloader-project/opendataloader-pdf](https://github.com/opendataloader-project/opendataloader-pdf)
- Tags: how-to-guide
- Published: 2026-03-20

---

**Use a comma-separated string combining individual pages (e.g., `1,3`) and hyphenated ranges (e.g., `5-7`) with the `--pages` CLI flag or `pages` API property to control which PDF pages are processed.**

The **opendataloader-pdf** library provides granular control over PDF processing through its flexible `pages` option. Whether you need to extract content from specific chapters or skip unnecessary front matter, you can **specify a range of pages for extraction** using a simple comma-separated syntax. This functionality is available across the CLI, Node.js, and Java APIs, with consistent parsing logic implemented in the core configuration class.

## Understanding the `pages` Syntax

The `pages` option accepts a **comma-separated string** that supports mixed notation for individual pages and continuous ranges. When provided, this string is parsed into a list of integers that determines exactly which pages are handed to the PDF extractors. The CLI option is documented in the auto-generated reference table at `content/docs/cli-options-reference.mdx`.

### Supported Format Patterns

- **Single pages**: Specify individual pages using 1-based indexing (e.g., `1`, `2`, `10`).
- **Continuous ranges**: Use hyphenated notation to define inclusive ranges (e.g., `5-7` extracts pages 5, 6, and 7).
- **Mixed specifications**: Combine singles and ranges in one string (e.g., `1,3,5-7` extracts pages 1, 3, 5, 6, and 7).

If the `pages` option is omitted or set to an empty string, the processing pipeline extracts **all pages** from the document.

## How Page Ranges Are Parsed

According to the opendataloader-pdf source code, the parsing logic resides in [`Config.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/Config.java) at [`java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/Config.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/Config.java). The `parsePageRanges` method (lines 648-690) implements the following validation and expansion algorithm:

1. **Tokenization**: The input string is split on commas.
2. **Range detection**: Each token is trimmed and checked for the presence of `-`.
3. **Range expansion**: If a hyphen is found, `parseRange` validates that both numbers are positive and that the start value is less than or equal to the end value, then adds every integer in that interval to the result list.
4. **Single page validation**: Tokens without hyphens are processed by `parseSinglePage`, which validates the token as a positive integer.
5. **Caching**: The final list is stored in `Config.cachedPageNumbers` for use by the processing pipeline.

Invalid formats trigger an `IllegalArgumentException` with a descriptive message indicating the problematic format.

## Usage Examples

### CLI Usage

Pass the `--pages` flag with a quoted string to extract specific pages:

```bash

# Extract pages 1, 3, and 5 through 7 as JSON and Markdown

opendataloader-pdf document.pdf \
  --format json,markdown \
  --pages "1,3,5-7" \
  -o ./output

```

### Node.js API

When using the TypeScript/JavaScript wrapper, provide the `pages` property in the options object:

```javascript
import { convert } from 'opendataloader-pdf';

await convert('document.pdf', {
  format: ['json', 'markdown'],
  pages: '2,4-6',
  outputDir: './output'
});

```

The TypeScript definition in [`node/opendataloader-pdf/src/convert-options.generated.ts`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/node/opendataloader-pdf/src/convert-options.generated.ts) declares this property as `pages?: string;`, mirroring the Java implementation.

### Java API

Programmatically configure page ranges using the `Config` class:

```java
import org.opendataloader.pdf.api.Config;
import org.opendataloader.pdf.core.PdfProcessor;

Config config = new Config();
config.setPages("1,3,5-7");
PdfProcessor.process("document.pdf", config);

```

## Error Handling and Validation

The parser strictly validates input to prevent processing errors. Providing malformed ranges or non-numeric values results in immediate feedback.

```java
try {
    config.setPages("1,abc,5-3");  // Invalid: non-numeric and reversed range
} catch (IllegalArgumentException e) {
    System.err.println("Invalid page spec: " + e.getMessage());
}

```

As implemented in [`Config.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/Config.java), the validator ensures all page numbers are positive integers and that range start values do not exceed their end values.

## Summary

- The `pages` option accepts a **comma-separated string** supporting individual pages (`1`, `3`) and inclusive ranges (`5-7`).
- Syntax validation and expansion occur in `Config.parsePageRanges` within [`java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/Config.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/Config.java).
- The parser uses **1-based indexing** and throws `IllegalArgumentException` for invalid formats or reversed ranges.
- Omitting the option causes the pipeline to process **all pages** by default.
- The same syntax works across CLI (`--pages`), Node.js (`pages` property), and Java (`setPages` method) interfaces.

## Frequently Asked Questions

### What happens if I omit the `pages` option?

If you omit the `pages` option or provide an empty string, opendataloader-pdf processes **all pages** in the PDF document. The `Config.cachedPageNumbers` list remains unset, causing the pipeline to skip page filtering and extract the entire document.

### Can I use zero-based indexing for page numbers?

No. The opendataloader-pdf parser uses **1-based indexing** as implemented in `Config.parseSinglePage` and `Config.parseRange`. Specifying page `0` or negative integers triggers an `IllegalArgumentException` during validation.

### How do I extract non-contiguous pages?

Use comma-separated values to specify individual pages and ranges in any order. For example, `"1,3,5-7,10"` extracts pages 1, 3, 5, 6, 7, and 10, skipping all others. The parser expands these into a deduplicated list stored in `Config.cachedPageNumbers`.

### What error occurs with invalid page range formats?

Invalid formats throw an `IllegalArgumentException` with a message identifying the specific error, such as non-numeric tokens (e.g., `"abc"`), reversed ranges (e.g., `"5-3"`), or malformed syntax. This validation occurs immediately when calling `setPages()` or equivalent API methods, preventing invalid configurations from reaching the extraction pipeline.