# What Document Formats Does LiteParse Support Besides PDF?

> Discover document formats LiteParse supports beyond PDF including DOCX, XLSX, PPTX, and various image types. Learn how LiteParse efficiently extracts text from diverse files.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: getting-started
- Published: 2026-06-07

---

**LiteParse supports DOCX, XLSX, PPTX, PNG, JPG, JPEG, GIF, SVG, and other raster images by automatically converting them to PDF through LibreOffice or ImageMagick before applying its spatial text extraction engine.**

The `run-llama/liteparse` library is built around a fast PDFium-based core, yet it accepts a wide range of non-PDF document formats through a transparent conversion pipeline. Understanding what document formats LiteParse supports beyond PDF helps developers integrate a single parser for Office files, presentations, and images without managing external converters manually.

## How LiteParse Routes Non-PDF Files Through Conversion

Internally, any input that is not already a PDF is forwarded to a conversion stage defined in [`crates/liteparse/src/conversion.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/conversion.rs). According to the run-llama/liteparse source code, two backend tools power this layer: **LibreOffice** handles Microsoft Office and OpenDocument files, while **ImageMagick** processes raster and vector images. Once converted, the resulting PDF flows into the same parser used for native PDFs, ensuring identical spatial-text extraction and OCR-merge logic across all inputs.

### Office Documents: DOCX, XLSX, and PPTX

LiteParse treats Word documents, Excel spreadsheets, and PowerPoint presentations as first-class inputs. When `parse()` receives a `.docx`, `.xlsx`, or `.pptx` file, the engine invokes LibreOffice in headless mode to generate a PDF intermediary. This path is documented in the project README and the [`docs/src/content/docs/liteparse/index.md`](https://github.com/run-llama/liteparse/blob/main/docs/src/content/docs/liteparse/index.md) page, which highlights "Parse Office files and images with support for DOCX, XLSX, PPTX, PNG, JPG, and more via automatic conversion."

### Image Formats: PNG, JPG, JPEG, GIF, SVG, and Raster Variants

For visual inputs, LiteParse leverages ImageMagick to convert **PNG**, **JPG**, **JPEG**, **GIF**, **SVG**, and additional raster types such as BMP and TIFF into PDF pages. After conversion, the library extracts text and bounding boxes exactly as it would from a scanned PDF. The CLI flag `--render-pages` also works for these image-derived PDFs because the pipeline is shared.

## Extension Validation in [`conversion.rs`](https://github.com/run-llama/liteparse/blob/main/conversion.rs)

Before any conversion begins, LiteParse validates the input extension. The function `is_supported_extension`, located on line 128 of [`crates/liteparse/src/conversion.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/conversion.rs), enumerates the allowed suffixes: `pdf`, `docx`, `xlsx`, `pptx`, `png`, `jpg`, `jpeg`, `gif`, and `svg`. Files outside this set—such as executables—are rejected immediately. The CLI entry point in [`crates/liteparse/src/main.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/main.rs) performs similar validation before invoking the conversion routine.

## Parsing Non-PDF Files in Node.js, Python, and Rust

The public API remains identical regardless of input format because the conversion step is transparent to callers.

### TypeScript Example: Parsing a DOCX File

```typescript
import { LiteParse } from '@liteparse/node';

async function parseDocx() {
  const parser = new LiteParse({ ocr: false });
  const result = await parser.parse('example.docx');
  console.log(result.text);        // extracted text
  console.log(result.boxes);       // spatial layout data
}
parseDocx();

```

The `parse` method in [`packages/node/src/lib.ts`](https://github.com/run-llama/liteparse/blob/main/packages/node/src/lib.ts) automatically converts the DOCX to PDF before feeding it to the core engine.

### Python Example: Parsing an XLSX File

```python
from liteparse import LiteParse

parser = LiteParse(ocr=False)
result = parser.parse("budget.xlsx")
print(result.text)      # plain-text output

print(result.json())    # structured JSON with bounding boxes

```

The Python wrapper in [`packages/python/liteparse/parser.py`](https://github.com/run-llama/liteparse/blob/main/packages/python/liteparse/parser.py) forwards the call to the native binary, which handles the LibreOffice conversion internally.

### Rust CLI Example: Rendering Pages from a PPTX

```bash
liteparse --input presentation.pptx --render-pages

```

Because the Rust CLI in [`crates/liteparse/src/main.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/main.rs) delegates to the same conversion routine, the `--render-pages` flag works for any supported input once it has been normalized to PDF.

## Key Source Files Behind Multi-Format Support

- **[`crates/liteparse/src/conversion.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/conversion.rs)** — Defines extension checks and orchestrates LibreOffice and ImageMagick to produce PDFs.
- **[`crates/liteparse/src/main.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/main.rs)** — Serves as the CLI entry point; validates input extensions before invoking conversion.
- **[`packages/node/src/lib.ts`](https://github.com/run-llama/liteparse/blob/main/packages/node/src/lib.ts)** — Exposes the `LiteParse` class to JavaScript and TypeScript callers.
- **[`packages/python/liteparse/parser.py`](https://github.com/run-llama/liteparse/blob/main/packages/python/liteparse/parser.py)** — Provides the Python wrapper that forwards calls to the native binary.
- **[`README.md`](https://github.com/run-llama/liteparse/blob/main/README.md)** — Lists the high-level feature set and supported formats.
- **[`docs/src/content/docs/liteparse/index.md`](https://github.com/run-llama/liteparse/blob/main/docs/src/content/docs/liteparse/index.md)** — Documents the multi-format support page for end users.

## Summary

- LiteParse supports DOCX, XLSX, PPTX, PNG, JPG, JPEG, GIF, SVG, and additional raster images beyond native PDF.
- Non-PDF files pass through an automatic conversion layer in [`crates/liteparse/src/conversion.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/conversion.rs) that uses LibreOffice for Office documents and ImageMagick for images.
- The `is_supported_extension` function on line 128 of [`conversion.rs`](https://github.com/run-llama/liteparse/blob/main/conversion.rs) gates all inputs by checking against an explicit allow-list of extensions.
- Wrappers for Node.js, Python, and the Rust CLI expose the same unified `parse` API regardless of original format.

## Frequently Asked Questions

### Does LiteParse support Microsoft Word documents?

Yes. LiteParse accepts DOCX files and routes them through LibreOffice to generate a PDF intermediary before extracting text and layout boxes in the same manner as a native PDF.

### Can LiteParse extract text from images like PNG or JPG?

Yes. Raster images including PNG, JPG, JPEG, GIF, and SVG are converted to PDF via ImageMagick. Additional raster formats such as BMP and TIFF are also handled through the same pipeline.

### Is the conversion step manual or automatic?

The conversion is fully automatic and transparent. Users call the same `parse` method across all supported document formats; the library manages LibreOffice or ImageMagick execution internally.

### Where does LiteParse validate file extensions?

The `is_supported_extension` function on line 128 of [`crates/liteparse/src/conversion.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/conversion.rs) validates inputs. It allows `pdf`, `docx`, `xlsx`, `pptx`, `png`, `jpg`, `jpeg`, `gif`, and `svg`, rejecting any unsupported extension before conversion begins.