# How to Handle Non-PDF Formats via LibreOffice Conversion with LiteParse

> Easily parse DOCX, XLSX, and PPTX files with LiteParse. Learn how it automatically converts Office documents to PDF using LibreOffice for seamless parsing without manual steps.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: how-to-guide
- Published: 2026-06-24

---

**LiteParse automatically converts Office documents (DOCX, XLSX, PPTX) to PDF using LibreOffice before parsing, requiring no manual conversion steps—just pass the file path to `LiteParse::parse`.**

LiteParse is a Rust-based document parsing library that handles complex Office formats by leveraging LibreOffice's headless conversion capabilities. Whether you're processing Word documents, Excel spreadsheets, or PowerPoint presentations, the library transparently handles non-PDF formats via LibreOffice conversion without requiring intermediate files or manual intervention.

## How LibreOffice Conversion Works in LiteParse

The conversion pipeline is orchestrated through the `conversion` module in [`crates/liteparse/src/conversion.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/conversion.rs). When you invoke `LiteParse::parse` (defined in `parser.rs#L83-L101`), the library first routes the input through `resolve_pdf_input` (`conversion.rs#L86-L115`) to determine if conversion is necessary.

### Entry Point and Routing

The function `resolve_pdf_input` checks file extensions against internal whitelists including `OFFICE_EXTENSIONS`, `PRESENTATION_EXTENSIONS`, and `SPREADSHEET_EXTENSIONS`. If the input is not a PDF, it calls `convert_to_pdf` (`conversion.rs#L44-L81`), which selects `ConversionTool::LibreOffice` for Office documents.

### Tool Detection and Execution

Before spawning the process, `find_libre_office_command` (`conversion.rs#L60-L71`) locates the binary across platforms—checking `$PATH` for `libreoffice` or `soffice`, probing macOS bundles at `/Applications/LibreOffice.app/`, and checking Windows defaults at `C:\Program Files\Libreoffice\program\soffice.exe`.

The `convert_office_document` function (`conversion.rs#L102-L149`) executes LibreOffice in headless mode with `--headless --convert-to pdf`. It creates a unique temporary user profile using `tempfile::Builder` to avoid the built-in profile lock that typically causes contention during concurrent conversions.

### PDF Retrieval

After LibreOffice writes the output, `find_pdf_in_dir` (`conversion.rs#L150-L176`) scans the temporary directory to locate the generated PDF. LibreOffice may sanitize filenames during conversion, so this function handles name mapping before returning the PDF path to the parser for standard PDFium extraction.

## Supported File Formats

LiteParse handles the following non-PDF formats through LibreOffice conversion:

- **Word documents**: DOCX, DOC, ODT
- **Excel spreadsheets**: XLSX, XLS, ODS  
- **PowerPoint presentations**: PPTX, PPT, ODP

Image-only inputs route to ImageMagick instead, while plain-text files bypass conversion entirely.

## Implementation Examples

### Rust

Pass a DOCX path directly to `parse`. The conversion happens asynchronously before PDFium extraction:

```rust
use liteparse::LiteParse;
use liteparse::config::LiteParseConfig;

#[tokio::main]
async fn main() -> Result<(), liteparse::error::LiteParseError> {
    let cfg = LiteParseConfig::default();
    let lp = LiteParse::new(cfg);
    
    // Automatically triggers LibreOffice conversion for DOCX
    let result = lp.parse("report.docx").await?;
    println!("Extracted {} pages", result.pages.len());
    Ok(())
}

```

### Python

The Python bindings expose the same functionality through `parse_file`:

```python
from liteparse import LiteParse

lp = LiteParse()
result = lp.parse_file("data.xlsx")  # Converts XLSX to PDF first

print(f"Pages: {len(result.pages)}")
print(result.text)

```

### Node.js

The TypeScript wrapper forwards paths to the native Rust implementation:

```typescript
import { LiteParse } from "liteparse";

const lp = new LiteParse({});
const { text } = await lp.parseFile("presentation.pptx");
console.log(text);

```

### CLI

Process Office documents directly from the command line:

```bash
liteparse parse meeting_notes.docx

```

## Error Handling and Platform Support

If LibreOffice is not found, `convert_office_document` returns a `LiteParseError::Conversion` with a clear installation message. The conversion runs within a Tokio async task (`execute_command`) but blocks only for the external LibreOffice process; subsequent OCR and layout reconstruction proceed concurrently after the PDF is produced.

On macOS, the library automatically detects the application bundle path. On Windows, it checks standard installation directories. Linux systems require `libreoffice` or `soffice` in `$PATH`.

## Summary

- **LiteParse** handles DOCX, XLSX, and PPTX via transparent LibreOffice conversion in [`crates/liteparse/src/conversion.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/conversion.rs)
- The `resolve_pdf_input` function routes non-PDFs to `convert_to_pdf`, which spawns LibreOffice with `--headless --convert-to pdf`
- **Temporary user profiles** prevent lock contention during concurrent conversions
- **Cross-platform detection** automatically finds LibreOffice binaries on macOS, Windows, and Linux
- All conversion steps are hidden from the API consumer—simply pass the original Office file path to `LiteParse::parse`

## Frequently Asked Questions

### What happens if LibreOffice is not installed?

If `find_libre_office_command` fails to locate the binary, `convert_office_document` returns a `LiteParseError::Conversion` error with a message instructing you to install LibreOffice. The library checks for `libreoffice`, `soffice`, and platform-specific default paths before failing.

### Is the conversion process thread-safe?

Yes. Each conversion creates an isolated temporary user profile directory using `tempfile::Builder`, preventing the profile locks that typically prevent concurrent LibreOffice instances. This allows multiple `LiteParse::parse` calls to run in parallel without collision.

### Can I disable automatic conversion for specific file types?

No. Conversion is intrinsic to the parsing pipeline for non-PDF formats. If you need to handle Office files without conversion, you would need to convert them externally before passing them to LiteParse, though this bypasses the integrated error handling in [`crates/liteparse/src/conversion.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/conversion.rs).

### Why does the converted PDF sometimes have a different filename?

LibreOffice sanitizes filenames during conversion to ensurePDF compatibility. The `find_pdf_in_dir` function (`conversion.rs#L150-L176`) accounts for this by scanning the output directory for any `.pdf` file rather than expecting a specific name match.