How to Configure Image Output Mode (Placeholder, Embed, Off) in LiteParse Markdown

LiteParse provides three image output modes—placeholder (default), off, and embed—that control how raster images appear when converting documents to Markdown, configured via the image_mode parameter in code or the --image-mode CLI flag.

LiteParse is a Rust-based document parsing library developed by LlamaIndex (run-llama/liteparse) that converts PDFs and other formats to structured Markdown. When you configure image output mode in LiteParse Markdown generation, you determine whether images appear as lightweight placeholders, embedded files, or are removed entirely from the output.

Understanding the Three Image Output Modes

LiteParse defines image handling behavior through the ImageMode enum in crates/liteparse/src/config.rs. Each mode affects how the markdown.rs formatter processes raster images during document conversion.

Placeholder Mode (Default)

Placeholder is the default setting when initializing LiteParseConfig. In this mode, the parser emits Markdown image syntax such as ![](image_pN_K.png) at the image's vertical position without including actual image bytes in the response. This produces lightweight Markdown text that preserves the document structure while omitting binary data.

Off Mode

When configured to Off, LiteParse strips all image references from the Markdown output. According to the implementation in crates/liteparse/src/markdown_layout/classify.rs, the page classifier filters out any Block::Figure entries before they reach the formatter. This ensures no image markup appears in the final text, useful when you need pure text extraction without visual element references.

Embed Mode

The Embed variant extracts each detected image as a PNG file and references it in the generated Markdown. When using this mode, the parser populates ParseResult.images with the image bytes, and the CLI writes these files to the directory specified by --image-output-dir. This mode requires both the mode setting and an output directory to persist the extracted images.

How to Configure Image Output Mode

You can configure the image mode through the command line interface or programmatically via the Rust API and language bindings.

Command Line Interface

The CLI accepts the --image-mode flag, parsed by the parse_image_mode function in crates/liteparse/src/main.rs. This function accepts "off", "none", "placeholder", or "embed" (case-insensitive) and maps them to the corresponding enum variants.


# Default placeholder mode

lit parse doc.pdf --format markdown

# Completely remove images

lit parse doc.pdf --format markdown --image-mode off

# Extract and embed images to a directory

lit parse doc.pdf --format markdown \
    --image-mode embed --image-output-dir ./images

Rust Library API

When constructing a LiteParseConfig struct, set the image_mode field to the desired ImageMode enum variant. The enum is defined in crates/liteparse/src/config.rs with serde attributes for lowercase serialization.

use liteparse::config::{ImageMode, LiteParseConfig, OutputFormat};
use liteparse::LiteParse;

let config = LiteParseConfig {
    output_format: OutputFormat::Markdown,
    image_mode: ImageMode::Embed,  // or ImageMode::Off / ImageMode::Placeholder
    ..Default::default()
};

let parser = LiteParse::new(config);
let result = parser.parse("doc.pdf").await?;

// result.text contains Markdown with image references
// result.images holds PNG bytes only when ImageMode::Embed is used

Python and TypeScript Bindings

The configuration passes through to the Python and Node.js bindings via the image_mode string parameter.

Python:

from liteparse import LiteParse

parser = LiteParse(
    output_format="markdown",
    image_mode="embed",          # "placeholder" | "off" | "embed"

    image_output_dir="./images"  # Required for embed mode

)

result = parser.parse("doc.pdf")
print(result.text)  # Markdown with embedded image references

TypeScript:

import { LiteParse } from "@llamaindex/liteparse";

const parser = new LiteParse({
  outputFormat: "markdown",
  imageMode: "off",  // "placeholder" (default) | "off" | "embed"
});

const result = await parser.parse("doc.pdf");
console.log(result.text);  // No image references appear

Technical Implementation Details

The ImageMode enum in crates/liteparse/src/config.rs uses #[serde(rename_all = "lowercase")] to accept lowercase string inputs in JSON and YAML configurations. The default value is ImageMode::Placeholder.

The Markdown formatter in crates/liteparse/src/output/markdown.rs receives the selected image_mode and passes it to classify_page_with_filters in the classifier module. This routine determines whether to keep, drop, or placeholder-replace images based on the active mode.

When ImageMode::Embed is active, the CLI handling code in crates/liteparse/src/main.rs (lines 24-32) writes the collected images to disk only if image_output_dir is specified. The write step is skipped entirely for placeholder and off modes, optimizing performance when image extraction is unnecessary.

Summary

  • Three modes available: placeholder (default), off, and embed control Markdown image representation.
  • Configuration locations: Set via --image-mode CLI flag or image_mode field in LiteParseConfig.
  • File paths: Enum defined in crates/liteparse/src/config.rs, CLI parsing in main.rs, and filtering logic in markdown_layout/classify.rs.
  • Embed requirements: Requires --image-output-dir to persist extracted PNG files; other modes ignore this parameter.
  • Performance impact: off and placeholder modes skip image file writing, making them faster for text-only extraction.

Frequently Asked Questions

What is the default image mode in LiteParse?

The default image mode is placeholder, defined in crates/liteparse/src/config.rs as the default value for the ImageMode enum. This mode inserts Markdown image syntax without including the actual binary data, producing lightweight text output while preserving document structure.

How do I completely remove images from Markdown output?

Set the image mode to off using the CLI flag --image-mode off or the configuration parameter image_mode: ImageMode::Off. This triggers the classifier in crates/liteparse/src/markdown_layout/classify.rs to filter out all Block::Figure entries before they reach the Markdown formatter.

Where are embedded images saved when using the embed mode?

When using ImageMode::Embed, images are written to the directory specified by the --image-output-dir CLI argument (or image_output_dir in library configurations). The CLI code in crates/liteparse/src/main.rs creates this directory if it doesn't exist and writes files as image_{id}.png format. The image bytes are also returned in ParseResult.images for programmatic access.

Can I use different image modes for different documents in the same application?

Yes, image mode is configured per-parser instance. Create separate LiteParse instances with different LiteParseConfig structs, or invoke the CLI with different --image-mode values for each document. Each parser maintains its own configuration state, allowing you to process one document with embed mode while another uses off mode within the same application.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →