# Embedded vs External `image_output` Modes in OpenDataLoader-PDF: Key Implications and Trade-offs

> Explore embedded vs external image_output modes in OpenDataLoader-PDF. Understand trade-offs in file size, memory, and portability for your data.

- Repository: [opendataloader-project/opendataloader-pdf](https://github.com/opendataloader-project/opendataloader-pdf)
- Tags: deep-dive
- Published: 2026-03-20

---

**Choosing `embedded` mode encodes images as Base64 data-URIs directly into your JSON, HTML, or Markdown output, while `external` mode saves images as separate files and references them by relative paths, significantly impacting file size, memory usage, and content portability.**

The `image_output` configuration option in **opendataloader-pdf** controls how raster images extracted from PDFs are handled during document conversion. Understanding the distinction between these modes is critical for optimizing performance and output portability in production workflows.

## What Are the `image_output` Modes?

The `image_output` setting accepts three distinct values that determine image extraction behavior:

- **`off`** – Disables image extraction entirely; no image data is processed or written.
- **`embedded`** – Images are encoded as **Base64 data-URIs** and embedded directly into the generated output files.
- **`external`** – Images are saved as separate binary files in a designated directory, with the output containing only **relative file-path references** (this is the default behavior).

In [`Config.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/Config.java), the `isEmbedImages()` method returns `true` when the mode is set to `embedded` and `false` for `external` or `off` modes, as shown in lines 514-520 of the source file.

## How the Mode Propagates Through the Codebase

Understanding the data flow helps explain why the mode choice affects both memory and I/O performance.

### Configuration Layer

The `Config` class defines the constants `IMAGE_OUTPUT_EMBEDDED` and `IMAGE_OUTPUT_EXTERNAL`, providing the `setImageOutput()` and `isEmbedImages()` methods. When `isEmbedImages()` returns `true`, downstream components know to prepare Base64 strings rather than file handles.

### CLI Parsing

Command-line arguments are parsed in [`CLIOptions.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/CLIOptions.java) (lines 97-105), where the `--image-output` flag is validated and stored in the `Config` instance. The CLI defaults to `external` unless explicitly overridden.

### Runtime Propagation

During document processing, [`DocumentProcessor.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/DocumentProcessor.java) (lines 283-286) copies the embedding flag into `StaticLayoutContainers`, a static holder that makes the configuration accessible to all serializers without passing references through every method call.

### Serialization Logic

Finally, [`ImageSerializer.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/ImageSerializer.java) (lines 46-51) checks `StaticLayoutContainers.isEmbedImages()` at write time. If **embedded**, it writes a `data:image/png;base64,` URI containing the encoded bytes. If **external**, it writes the filename and relies on the image having already been saved to disk by `DocumentProcessor`.

## Practical Implications: Embedded vs External

The choice between modes creates distinct trade-offs across several operational dimensions:

**Self-contained output**
- **Embedded**: Produces a single, self-contained file that includes all image data. Ideal for email attachments or single-file API responses.
- **External**: Requires managing both the output file and the accompanying image directory; losing the directory breaks image references.

**File size and storage**
- **Embedded**: Increases output size by approximately **33%** due to Base64 encoding overhead. Large PDFs with many high-resolution images generate massive JSON/HTML files.
- **External**: Keeps the output document lightweight; images remain as optimized binary files.

**Performance characteristics**
- **Embedded**: Consumes significantly more **memory and CPU** during encoding and subsequent parsing, as the entire Base64 string must be held in memory and processed as text.
- **External**: Enables faster streaming with lower memory pressure; images can be served via CDN or cached independently without parsing large text blocks.

**Downstream consumption**
- **Embedded**: Allows simple JSON viewers or web browsers to display images immediately without resolving external paths.
- **External**: Requires consumers to handle file-path resolution; consumers that cannot access the filesystem or resolve relative URLs will fail to render images.

## Code Examples

### Programmatic Configuration

Configure the mode directly in Java using the `Config` API:

```java
import org.opendataloader.pdf.api.Config;

// Embedded mode: Base64 data-URIs
Config embeddedConfig = new Config();
embeddedConfig.setImageOutput(Config.IMAGE_OUTPUT_EMBEDDED);
System.out.println("Embedding enabled: " + embeddedConfig.isEmbedImages()); // true

// External mode: Separate files (default behavior)
Config externalConfig = new Config();
externalConfig.setImageOutput(Config.IMAGE_OUTPUT_EXTERNAL);
System.out.println("Embedding enabled: " + externalConfig.isEmbedImages()); // false

```

### Command-Line Usage

Set the mode via the CLI when processing PDFs:

```bash

# Embedded: Single output file with all images encoded

java -jar opendataloader-pdf-cli.jar input.pdf --image-output embedded > output.json

# External: Images saved to ./images directory with path references

java -jar opendataloader-pdf-cli.jar input.pdf --image-output external --image-dir ./images > output.json

```

### Serializer Branching Logic

The actual serialization logic in [`ImageSerializer.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/ImageSerializer.java) follows this pattern:

```java
if (StaticLayoutContainers.isEmbedImages()) {
    // Encode and embed
    String base64 = Base64.getEncoder().encodeToString(imageBytes);
    jsonWriter.value("data:image/png;base64," + base64);
} else {
    // Reference external file already written to disk
    jsonWriter.value(imageFileName);
}

```

## Summary

- **Embedded mode** produces self-contained, portable files at the cost of ~33% size inflation and higher memory/CPU usage.
- **External mode** (default) minimizes output size and memory pressure by storing images separately, but requires managing accompanying file directories.
- The mode propagates from `CLIOptions` through `Config` into `StaticLayoutContainers`, where `ImageSerializer` branches on `isEmbedImages()` to determine output format.
- Choose **embedded** for single-file transports and simple web embedding; choose **external** for large document batches, CDN delivery, or memory-constrained environments.

## Frequently Asked Questions

### What is the default image_output mode in opendataloader-pdf?

The library defaults to **`external`** mode. According to the source code in [`CLIOptions.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/CLIOptions.java), when no `--image-output` flag is specified, the system initializes with external file references to maintain lightweight output and minimize memory overhead.

### How does embedded mode affect JSON file size?

Embedded mode typically increases JSON output size by approximately **33%** compared to the raw binary image data, due to Base64 encoding overhead. For PDFs containing many high-resolution images, this can result in multi-megabyte or gigabyte-scale text files that are slow to parse and transmit.

### Can I switch modes without modifying the source code?

Yes. The mode is fully configurable via the **`--image-output`** CLI flag or programmatically through the `Config.setImageOutput()` method. No source code changes are required; simply pass `embedded` or `external` as the argument value.

### Why does external mode require an image directory?

When using external mode, `DocumentProcessor` extracts image bytes and writes them as separate files (e.g., PNG or JPEG) to the filesystem. The output JSON/HTML contains only relative paths pointing to these files. If the image directory is moved or deleted, the references become broken and images will fail to load in the rendered output.