# Preserve Original Line Breaks in OpenDataLoader-PDF: Impact on Text Output Formatting

> Discover how `preserve original line breaks` in OpenDataLoader-PDF controls newline characters, impacting table cell formatting and ensuring accurate text output in Markdown or HTML.

- Repository: [opendataloader-project/opendataloader-pdf](https://github.com/opendataloader-project/opendataloader-pdf)
- Tags: deep-dive
- Published: 2026-03-20

---

**When enabled via the `keepLineBreaks` configuration flag, preserve original line breaks retains newline characters (`\n`) inside PDF table cells, rendering them as actual line breaks in Markdown or `<br/>` tags in HTML; outside of tables, line breaks are always normalized to spaces to ensure proper paragraph flow.**

The `opendataloader-project/opendataloader-pdf` library provides granular control over text extraction through the **preserve original line breaks** option. This setting determines whether line-break characters inside PDF text blocks are maintained or collapsed during conversion to Markdown or HTML. Understanding this flag is essential when processing tabular data that relies on manual line breaks for visual layout.

## How the `keepLineBreaks` Flag Controls Line Break Preservation

The preserve original line breaks behavior is governed by the `keepLineBreaks` boolean property in the `Config` class, located at [`java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/Config.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/Config.java). The flag defaults to `false`, meaning line breaks are automatically converted to single spaces during extraction.

When processing begins, the `DocumentProcessor.updateStaticContainers` method (in [`java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/processors/DocumentProcessor.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/processors/DocumentProcessor.java)) propagates this flag into the static containers used throughout the rendering pipeline. This makes the configuration available to both the Markdown and HTML generators during text node processing.

## Impact on Markdown Output Formatting

Inside the `MarkdownGenerator.writeSemanticTextNode` method ([`java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/markdown/MarkdownGenerator.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/markdown/MarkdownGenerator.java)), the library checks the flag to determine how to handle newline characters.

### Table Cells

When `keepLineBreaks` is set to `true` and the text resides inside a table cell (`isInsideTable()` returns true), newline characters are preserved as actual `\n` line breaks. This maintains the visual structure of multi-line cell content. When the flag is `false` (default), newlines are replaced with spaces, compressing the content into a single logical line.

### Headings and Regular Paragraphs

For Markdown headings, line breaks are **always** converted to spaces regardless of the flag setting. This prevents accidental fragmentation of heading syntax (e.g., breaking a `# Header` across lines). Outside of tables, regular paragraphs follow the same rule: line breaks are normalized to spaces to produce natural paragraph flow suitable for downstream Markdown consumers.

## Impact on HTML Output Formatting

The HTML generator applies similar logic within the `HtmlGenerator.writeParagraph` method ([`java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/html/HtmlGenerator.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/html/HtmlGenerator.java)).

### Table Cell Rendering

When `keepLineBreaks` is enabled, newline characters inside table cells are transformed into explicit `<br/>` tags (using `HtmlSyntax.HTML_LINE_BREAK_TAG`). This produces visual line breaks when the HTML is rendered in a browser. When disabled, the newlines are collapsed into spaces, resulting in continuous text within the `<td>` element.

Note that outside of table cells, this flag has no effect on HTML output; paragraph breaks are handled through standard block-level elements rather than preserved newlines.

## Configuration Examples

Enable preserve original line breaks through the Java API or Node.js wrapper.

Java configuration:

```java
import org.opendataloader.pdf.api.Config;

Config cfg = new Config();
cfg.setKeepLineBreaks(true);  // Preserve \n inside table cells
// Pass cfg to your DocumentProcessor

```

Node.js wrapper:

```javascript
import { convert } from "opendataloader-pdf";

await convert({
  input: "sample.pdf",
  output: "out.md",
  keepLineBreaks: true  // Maps to the Java flag
});

```

## Output Comparison: With and Without Preserved Line Breaks

Consider a PDF table cell containing the text "First line" followed by a newline and "second line".

With `keepLineBreaks: false` (default), Markdown output appears as:

```markdown
| Name | Description |
|------|-------------|
| Item | First line second line |

```

With `keepLineBreaks: true`:

```markdown
| Name | Description |
|------|-------------|
| Item | First line  
second line |

```

In HTML, the same flag produces different `<td>` content. Disabled:

```html
<td>First line second line</td>

```

Enabled:

```html
<td>First line<br/>second line</td>

```

## Summary

- The **preserve original line breaks** option (`keepLineBreaks`) affects only table cell content in both Markdown and HTML output.
- When enabled, `\n` characters are preserved as actual line breaks in Markdown or `<br/>` tags in HTML.
- When disabled (default), line breaks collapse into single spaces to create continuous text flow.
- The flag has no effect on regular paragraphs outside tables or on Markdown headings, where line breaks are always normalized to spaces.
- Configuration occurs through `Config.setKeepLineBreaks()` in Java or the `keepLineBreaks` option in the Node.js wrapper.

## Frequently Asked Questions

### Does preserve original line breaks affect regular paragraphs outside tables?

No. According to the source code in [`MarkdownGenerator.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/MarkdownGenerator.java) and [`HtmlGenerator.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/HtmlGenerator.java), the flag only impacts text processing when `isInsideTable()` returns true. Regular paragraphs always have their line breaks converted to spaces to ensure proper paragraph flow in downstream processing.

### Why are line breaks always converted to spaces in Markdown headings?

To prevent syntax breakage. The `MarkdownGenerator.writeSemanticTextNode` method explicitly replaces newlines with spaces inside headings to avoid fragmenting the Markdown heading structure (e.g., splitting `# Title` across multiple lines), which would break proper Markdown parsing.

### How do I enable preserve original line breaks in the Node.js wrapper?

Set the `keepLineBreaks` property to `true` in the options object passed to the `convert()` function. This maps directly to the Java `Config.setKeepLineBreaks()` method through the binding defined in [`node/opendataloader-pdf/src/convert-options.generated.ts`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/node/opendataloader-pdf/src/convert-options.generated.ts).

### What HTML tag is used when preserving line breaks in table cells?

When enabled, the system inserts a `<br/>` tag (referenced as `HtmlSyntax.HTML_LINE_BREAK_TAG` in the source) at the position of each preserved newline character within HTML table cells. This tag creates the visual line break when the HTML is rendered, preserving the original PDF layout.