How to Define Custom Page Separators for HTML Output in OpenDataLoader-PDF

OpenDataLoader-PDF allows you to inject custom HTML between converted pages using the --html-page-separator CLI flag or the htmlPageSeparator configuration property, with support for the %page-number% placeholder for dynamic page indexing.

When converting PDF documents to HTML format, the OpenDataLoader-PDF library provides precise control over how individual pages are delimited in the output stream. You can define custom page separators for HTML output through a dedicated configuration option that supports dynamic placeholder substitution. This feature is implemented in the core Java engine and exposed consistently across the CLI, Python, and Java APIs.

Configuring the HTML Page Separator

You can specify the separator string through either command-line arguments or programmatic configuration objects. The value is stored in the Config.htmlPageSeparator field (defined in java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/Config.java at lines 73-75) and consumed by the HtmlGenerator class during the rendering phase.

Command-Line Interface

Pass the desired HTML fragment using the --html-page-separator flag (documented in content/docs/cli-options-reference.mdx at lines 27-29) when invoking the converter:

opendataloader-pdf document.pdf -f html \
  --html-page-separator "<hr class=\"page-break\"/> <!-- Page %page-number% -->"

This inserts the specified markup between each page's content in the generated HTML file, rendering the page number as an HTML comment.

Python API

When using the Python bindings, provide the separator via the html_page_separator parameter:

from opendataloader_pdf import OpenDataLoaderPdf

loader = OpenDataLoaderPdf(
    input_path="document.pdf",
    output_folder="out",
    html_page_separator='<div class="separator">Page %page-number%</div>'
)
loader.run()

The html_page_separator argument maps directly to Config.htmlPageSeparator and undergoes the same placeholder substitution as the CLI.

Java API

For direct Java integration, configure the Config instance before instantiating the processor:

import org.opendataloader.pdf.api.Config;
import org.opendataloader.pdf.core.OpenDataLoaderPdf;

Config cfg = new Config();
cfg.setHtmlPageSeparator("<section class=\"page-break\" data-page=\"%page-number%\"></section>");
cfg.setGenerateHtml(true);

OpenDataLoaderPdf loader = new OpenDataLoaderPdf(cfg);
loader.process("document.pdf");

Internal Implementation and Placeholder Substitution

The separator injection logic resides in java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/html/HtmlGenerator.java, specifically within the writePageSeparator method (lines 30-34). During the conversion pipeline:

  1. Configuration Storage – When parsing command-line arguments, the value is stored in Config.htmlPageSeparator (also exposed via Node.js in node/opendataloader-pdf/src/cli-options.generated.ts at lines 21-23).
  2. Generator Initialization – HtmlGenerator receives the Config instance and copies htmlPageSeparator to its own field via config.getHtmlPageSeparator().
  3. Page Loop Execution – For each page, writePageSeparator(pageNumber) is invoked before the page's content is emitted.
  4. Placeholder Handling – If the separator contains Config.PAGE_NUMBER_STRING (%page-number%), the method replaces it with pageNumber + 1 (1-based indexing); otherwise, the literal string is written unmodified.
  5. Output Generation – The processed string is written to the HTML output stream, appearing between consecutive page blocks.

This architecture treats the separator as raw HTML, allowing any valid markup including comments, horizontal rules, or semantic container elements.

Generated Output Structure

When processing a multi-page PDF with a custom separator configured, the resulting HTML structure resembles:

<body>
  <!-- page 1 -->
  <p>First page content …</p>
  <section class="page-break" data-page="1"></section>

  <!-- page 2 -->
  <p>Second page content …</p>
  <section class="page-break" data-page="2"></section>
</body>

The separator appears between each page block but not after the final page, maintaining clean document boundaries.

Summary

  • CLI Configuration – Use --html-page-separator followed by an HTML string to inject markup between pages from the command line.
  • Placeholder Support – Include %page-number% in your separator to dynamically insert the current 1-based page index at runtime.
  • Implementation Location – The feature is defined in Config.java (lines 73-75) and executed in HtmlGenerator.java (lines 30-34).
  • API Consistency – The same functionality is available across Node.js (via cli-options.generated.ts), Python, and native Java interfaces.
  • Raw HTML Injection – The separator value is written unescaped, supporting any valid HTML markup for styling or semantic separation.

Frequently Asked Questions

What is the default value for the HTML page separator?

By default, Config.htmlPageSeparator is initialized as an empty string (""), meaning no separator is inserted between pages unless explicitly configured via the CLI flag or programmatic API.

Can I use multiple placeholders or custom variables in the separator?

The current implementation only replaces the specific %page-number% placeholder (defined as the constant Config.PAGE_NUMBER_STRING). No other dynamic variables are supported; any additional text in the separator string is rendered as static HTML.

How does the page numbering start in the placeholder?

The placeholder substitution uses 1-based indexing. When HtmlGenerator.writePageSeparator processes page index 0, it outputs 1 for the first occurrence of %page-number%, incrementing accordingly for subsequent pages.

Is the separator included after the last page?

No. The separator is inserted between pages only. The HtmlGenerator writes the separator before emitting each subsequent page's content, so the final page in the PDF does not have a trailing separator appended to the HTML output.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →