What Document Formats Does LiteParse Support Besides PDF?

LiteParse supports DOCX, XLSX, PPTX, PNG, JPG, JPEG, GIF, SVG, and other raster images by automatically converting them to PDF through LibreOffice or ImageMagick before applying its spatial text extraction engine.

The run-llama/liteparse library is built around a fast PDFium-based core, yet it accepts a wide range of non-PDF document formats through a transparent conversion pipeline. Understanding what document formats LiteParse supports beyond PDF helps developers integrate a single parser for Office files, presentations, and images without managing external converters manually.

How LiteParse Routes Non-PDF Files Through Conversion

Internally, any input that is not already a PDF is forwarded to a conversion stage defined in crates/liteparse/src/conversion.rs. According to the run-llama/liteparse source code, two backend tools power this layer: LibreOffice handles Microsoft Office and OpenDocument files, while ImageMagick processes raster and vector images. Once converted, the resulting PDF flows into the same parser used for native PDFs, ensuring identical spatial-text extraction and OCR-merge logic across all inputs.

Office Documents: DOCX, XLSX, and PPTX

LiteParse treats Word documents, Excel spreadsheets, and PowerPoint presentations as first-class inputs. When parse() receives a .docx, .xlsx, or .pptx file, the engine invokes LibreOffice in headless mode to generate a PDF intermediary. This path is documented in the project README and the docs/src/content/docs/liteparse/index.md page, which highlights "Parse Office files and images with support for DOCX, XLSX, PPTX, PNG, JPG, and more via automatic conversion."

Image Formats: PNG, JPG, JPEG, GIF, SVG, and Raster Variants

For visual inputs, LiteParse leverages ImageMagick to convert PNG, JPG, JPEG, GIF, SVG, and additional raster types such as BMP and TIFF into PDF pages. After conversion, the library extracts text and bounding boxes exactly as it would from a scanned PDF. The CLI flag --render-pages also works for these image-derived PDFs because the pipeline is shared.

Extension Validation in conversion.rs

Before any conversion begins, LiteParse validates the input extension. The function is_supported_extension, located on line 128 of crates/liteparse/src/conversion.rs, enumerates the allowed suffixes: pdf, docx, xlsx, pptx, png, jpg, jpeg, gif, and svg. Files outside this set—such as executables—are rejected immediately. The CLI entry point in crates/liteparse/src/main.rs performs similar validation before invoking the conversion routine.

Parsing Non-PDF Files in Node.js, Python, and Rust

The public API remains identical regardless of input format because the conversion step is transparent to callers.

TypeScript Example: Parsing a DOCX File

import { LiteParse } from '@liteparse/node';

async function parseDocx() {
  const parser = new LiteParse({ ocr: false });
  const result = await parser.parse('example.docx');
  console.log(result.text);        // extracted text
  console.log(result.boxes);       // spatial layout data
}
parseDocx();

The parse method in packages/node/src/lib.ts automatically converts the DOCX to PDF before feeding it to the core engine.

Python Example: Parsing an XLSX File

from liteparse import LiteParse

parser = LiteParse(ocr=False)
result = parser.parse("budget.xlsx")
print(result.text)      # plain-text output

print(result.json())    # structured JSON with bounding boxes

The Python wrapper in packages/python/liteparse/parser.py forwards the call to the native binary, which handles the LibreOffice conversion internally.

Rust CLI Example: Rendering Pages from a PPTX

liteparse --input presentation.pptx --render-pages

Because the Rust CLI in crates/liteparse/src/main.rs delegates to the same conversion routine, the --render-pages flag works for any supported input once it has been normalized to PDF.

Key Source Files Behind Multi-Format Support

Summary

  • LiteParse supports DOCX, XLSX, PPTX, PNG, JPG, JPEG, GIF, SVG, and additional raster images beyond native PDF.
  • Non-PDF files pass through an automatic conversion layer in crates/liteparse/src/conversion.rs that uses LibreOffice for Office documents and ImageMagick for images.
  • The is_supported_extension function on line 128 of conversion.rs gates all inputs by checking against an explicit allow-list of extensions.
  • Wrappers for Node.js, Python, and the Rust CLI expose the same unified parse API regardless of original format.

Frequently Asked Questions

Does LiteParse support Microsoft Word documents?

Yes. LiteParse accepts DOCX files and routes them through LibreOffice to generate a PDF intermediary before extracting text and layout boxes in the same manner as a native PDF.

Can LiteParse extract text from images like PNG or JPG?

Yes. Raster images including PNG, JPG, JPEG, GIF, and SVG are converted to PDF via ImageMagick. Additional raster formats such as BMP and TIFF are also handled through the same pipeline.

Is the conversion step manual or automatic?

The conversion is fully automatic and transparent. Users call the same parse method across all supported document formats; the library manages LibreOffice or ImageMagick execution internally.

Where does LiteParse validate file extensions?

The is_supported_extension function on line 128 of crates/liteparse/src/conversion.rs validates inputs. It allows pdf, docx, xlsx, pptx, png, jpg, jpeg, gif, and svg, rejecting any unsupported extension before conversion begins.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →