What Document Formats Does LiteParse Support Besides PDF?
LiteParse supports DOCX, XLSX, PPTX, PNG, JPG, JPEG, GIF, SVG, and other raster images by automatically converting them to PDF through LibreOffice or ImageMagick before applying its spatial text extraction engine.
The run-llama/liteparse library is built around a fast PDFium-based core, yet it accepts a wide range of non-PDF document formats through a transparent conversion pipeline. Understanding what document formats LiteParse supports beyond PDF helps developers integrate a single parser for Office files, presentations, and images without managing external converters manually.
How LiteParse Routes Non-PDF Files Through Conversion
Internally, any input that is not already a PDF is forwarded to a conversion stage defined in crates/liteparse/src/conversion.rs. According to the run-llama/liteparse source code, two backend tools power this layer: LibreOffice handles Microsoft Office and OpenDocument files, while ImageMagick processes raster and vector images. Once converted, the resulting PDF flows into the same parser used for native PDFs, ensuring identical spatial-text extraction and OCR-merge logic across all inputs.
Office Documents: DOCX, XLSX, and PPTX
LiteParse treats Word documents, Excel spreadsheets, and PowerPoint presentations as first-class inputs. When parse() receives a .docx, .xlsx, or .pptx file, the engine invokes LibreOffice in headless mode to generate a PDF intermediary. This path is documented in the project README and the docs/src/content/docs/liteparse/index.md page, which highlights "Parse Office files and images with support for DOCX, XLSX, PPTX, PNG, JPG, and more via automatic conversion."
Image Formats: PNG, JPG, JPEG, GIF, SVG, and Raster Variants
For visual inputs, LiteParse leverages ImageMagick to convert PNG, JPG, JPEG, GIF, SVG, and additional raster types such as BMP and TIFF into PDF pages. After conversion, the library extracts text and bounding boxes exactly as it would from a scanned PDF. The CLI flag --render-pages also works for these image-derived PDFs because the pipeline is shared.
Extension Validation in conversion.rs
Before any conversion begins, LiteParse validates the input extension. The function is_supported_extension, located on line 128 of crates/liteparse/src/conversion.rs, enumerates the allowed suffixes: pdf, docx, xlsx, pptx, png, jpg, jpeg, gif, and svg. Files outside this set—such as executables—are rejected immediately. The CLI entry point in crates/liteparse/src/main.rs performs similar validation before invoking the conversion routine.
Parsing Non-PDF Files in Node.js, Python, and Rust
The public API remains identical regardless of input format because the conversion step is transparent to callers.
TypeScript Example: Parsing a DOCX File
import { LiteParse } from '@liteparse/node';
async function parseDocx() {
const parser = new LiteParse({ ocr: false });
const result = await parser.parse('example.docx');
console.log(result.text); // extracted text
console.log(result.boxes); // spatial layout data
}
parseDocx();
The parse method in packages/node/src/lib.ts automatically converts the DOCX to PDF before feeding it to the core engine.
Python Example: Parsing an XLSX File
from liteparse import LiteParse
parser = LiteParse(ocr=False)
result = parser.parse("budget.xlsx")
print(result.text) # plain-text output
print(result.json()) # structured JSON with bounding boxes
The Python wrapper in packages/python/liteparse/parser.py forwards the call to the native binary, which handles the LibreOffice conversion internally.
Rust CLI Example: Rendering Pages from a PPTX
liteparse --input presentation.pptx --render-pages
Because the Rust CLI in crates/liteparse/src/main.rs delegates to the same conversion routine, the --render-pages flag works for any supported input once it has been normalized to PDF.
Key Source Files Behind Multi-Format Support
crates/liteparse/src/conversion.rs— Defines extension checks and orchestrates LibreOffice and ImageMagick to produce PDFs.crates/liteparse/src/main.rs— Serves as the CLI entry point; validates input extensions before invoking conversion.packages/node/src/lib.ts— Exposes theLiteParseclass to JavaScript and TypeScript callers.packages/python/liteparse/parser.py— Provides the Python wrapper that forwards calls to the native binary.README.md— Lists the high-level feature set and supported formats.docs/src/content/docs/liteparse/index.md— Documents the multi-format support page for end users.
Summary
- LiteParse supports DOCX, XLSX, PPTX, PNG, JPG, JPEG, GIF, SVG, and additional raster images beyond native PDF.
- Non-PDF files pass through an automatic conversion layer in
crates/liteparse/src/conversion.rsthat uses LibreOffice for Office documents and ImageMagick for images. - The
is_supported_extensionfunction on line 128 ofconversion.rsgates all inputs by checking against an explicit allow-list of extensions. - Wrappers for Node.js, Python, and the Rust CLI expose the same unified
parseAPI regardless of original format.
Frequently Asked Questions
Does LiteParse support Microsoft Word documents?
Yes. LiteParse accepts DOCX files and routes them through LibreOffice to generate a PDF intermediary before extracting text and layout boxes in the same manner as a native PDF.
Can LiteParse extract text from images like PNG or JPG?
Yes. Raster images including PNG, JPG, JPEG, GIF, and SVG are converted to PDF via ImageMagick. Additional raster formats such as BMP and TIFF are also handled through the same pipeline.
Is the conversion step manual or automatic?
The conversion is fully automatic and transparent. Users call the same parse method across all supported document formats; the library manages LibreOffice or ImageMagick execution internally.
Where does LiteParse validate file extensions?
The is_supported_extension function on line 128 of crates/liteparse/src/conversion.rs validates inputs. It allows pdf, docx, xlsx, pptx, png, jpg, jpeg, gif, and svg, rejecting any unsupported extension before conversion begins.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →