How markdown-with-images Handles Embedded Data in OpenDataLoader-PDF

The markdown-with-images format in opendataloader-pdf embeds images as Base-64 data-URIs directly into the Markdown output, eliminating the need for external image files while keeping the document self-contained.

The opendataloader-pdf library supports multiple output formats for converting PDF documents, with markdown-with-images specifically designed to produce Markdown files where visual content remains intact without requiring separate asset folders. This format processes embedded data through a configurable pipeline that transforms binary image data into inline Base-64 strings. Understanding this mechanism helps developers optimize for portability versus payload size.

How markdown-with-images Processes Embedded Data

CLI and API Configuration Options

According to the source code in CLIOptions.java, users activate the format through three primary interfaces:

  • Command-line interface: Pass --format markdown-with-images to the CLI jar
  • Python/Node bindings: Specify format=['markdown-with-images'] in the conversion call
  • Java API: Set addImageToMarkdown to true in the CLIOptions configuration

These entry points propagate the format request to the core convert() function, initializing the embedded data pipeline.

Core Configuration Flags

The Config.java class maintains two critical boolean flags that govern embedded data handling:

  1. addImageToMarkdown – Enables the markdown-with-images format mode
  2. embedImages – Determines whether images are written as Base-64 data-URIs or stored as external file paths

When the option parser processes user input, it invokes the setters for these flags, storing the configuration for downstream processors.

Thread-Local State Propagation

Before document processing begins, the DocumentProcessor copies the embedImages setting into a thread-local variable via StaticLayoutContainers.setEmbedImages(...). This mechanism, implemented in StaticLayoutContainers.java, makes the embedding configuration available to all downstream generators without requiring explicit Config object passing through every method call.

Markdown Generation and Base-64 Conversion

The MarkdownGenerator.java (lines 141-155 and 171-186) handles the actual embedded data generation:

  • When encountering SemanticImage or SemanticPicture elements, it constructs standard Markdown image syntax: ![alt](src)
  • If embedImages is true, the generator invokes [Base64ImageUtils.java] to convert the image file to a data-URI string
  • The utility reads the binary image data, checks size limits, and returns a data:image/<type>;base64,… string
  • If embedding is disabled, the generator references the relative path to the extracted image file in the images/ subdirectory

Embedded vs. External Image Storage

The distinction between embedded and external storage modes affects document portability and performance:

Embedded Base-64 Data-URIs create a single self-contained Markdown file suitable for GitHub rendering, Jupyter notebooks, or email attachments where external file references would break.

External File References store images separately in the images/ subdirectory, producing smaller Markdown files but requiring the entire directory structure to remain intact for proper rendering.

Users control this behavior through the embedImages configuration flag, which defaults to true for the markdown-with-images format but can be disabled via CLI flags or API parameters.

Practical Implementation Examples

Command-Line Usage

Generate Markdown with embedded images (default behavior):

java -jar opendataloader-pdf-cli.jar -i sample.pdf -f markdown-with-images

Generate Markdown with external image files:

java -jar opendataloader-pdf-cli.jar -i sample.pdf -f markdown-with-images --no-embed-images

The --no-embed-images flag explicitly sets Config.embedImages to false, forcing the extractor to write files to the images/ directory instead of embedding data-URIs.

Python Integration

from opendataloader_pdf import convert

# Default: embed images as Base-64 data

convert(
    input_path="sample.pdf",
    output_dir="out",
    format=["markdown-with-images"],
    embed_images=True
)

# External file mode

convert(
    input_path="sample.pdf",
    output_dir="out",
    format=["markdown-with-images"],
    embed_images=False
)

The embed_images argument maps directly to Config.setEmbedImages() in the Java core.

Node.js Implementation

import { convert } from "opendataloader-pdf";

await convert("sample.pdf", {
  outputDir: "out",
  format: ["markdown-with-images"],
  embedImages: true  // Optional; defaults to true
});

Direct Java API

import org.opendataloader.pdf.api.Config;
import org.opendataloader.pdf.Converter;
import java.util.List;

Config cfg = new Config();
cfg.setAddImageToMarkdown(true);   // Enable markdown-with-images format
cfg.setEmbedImages(true);          // Embed as data-URI (default)

List<String> formats = List.of("markdown-with-images");

new Converter().convert("sample.pdf", null, cfg, formats);

Summary

  • The markdown-with-images format embeds PDF images as Base-64 data-URIs directly within Markdown syntax, producing self-contained documents
  • Configuration flows from [CLIOptions.java] through [Config.java] and propagates via thread-local storage in [StaticLayoutContainers.java]
  • The [MarkdownGenerator.java] coordinates with [Base64ImageUtils.java] to convert binary image data to data:image/<type>;base64,… strings
  • Users toggle embedding behavior through the embedImages flag (CLI: --no-embed-images, APIs: embed_images/embedImages parameters)
  • External file mode extracts images to an images/ subdirectory, reducing Markdown payload size but requiring directory persistence

Frequently Asked Questions

What is the maximum image size for Base-64 embedding in markdown-with-images?

The opendataloader-pdf source code does not enforce a strict size limit within the embedding logic itself, though practical constraints apply based on Markdown renderer capabilities. The [Base64ImageUtils.java] class checks image dimensions during processing, but the Base-64 encoding overhead increases file size by approximately 33%. For very large PDFs with high-resolution images, disabling embedding via embedImages: false produces more manageable output.

Can I use markdown-with-images without embedding the actual image data?

Yes. Set the embedImages configuration flag to false using --no-embed-images on the CLI, embed_images=False in Python, or setEmbedImages(false) in Java. This mode stores extracted images in a separate images/ folder and references them via relative paths like ![image 1](images/001_image_1.png), keeping the Markdown file lightweight while maintaining image accessibility.

Does the embedded data format affect image quality or type?

No. The [Base64ImageUtils.java] utility preserves the original image format (PNG, JPEG, etc.) and binary integrity during the embedding process. The conversion to Base-64 is a reversible encoding transformation that does not re-encode or compress the image data, ensuring identical visual output whether images are embedded or stored externally.

How does thread-local storage improve embedded data handling in multi-threaded conversions?

The StaticLayoutContainers.java implementation uses thread-local variables to propagate the embedImages flag to all document processors without passing configuration objects through every method signature. This design allows concurrent PDF conversions to maintain independent embedding settings per thread, preventing race conditions when processing multiple documents with different output format requirements simultaneously.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →