How to Handle Non-PDF Formats via LibreOffice Conversion with LiteParse
LiteParse automatically converts Office documents (DOCX, XLSX, PPTX) to PDF using LibreOffice before parsing, requiring no manual conversion steps—just pass the file path to LiteParse::parse.
LiteParse is a Rust-based document parsing library that handles complex Office formats by leveraging LibreOffice's headless conversion capabilities. Whether you're processing Word documents, Excel spreadsheets, or PowerPoint presentations, the library transparently handles non-PDF formats via LibreOffice conversion without requiring intermediate files or manual intervention.
How LibreOffice Conversion Works in LiteParse
The conversion pipeline is orchestrated through the conversion module in crates/liteparse/src/conversion.rs. When you invoke LiteParse::parse (defined in parser.rs#L83-L101), the library first routes the input through resolve_pdf_input (conversion.rs#L86-L115) to determine if conversion is necessary.
Entry Point and Routing
The function resolve_pdf_input checks file extensions against internal whitelists including OFFICE_EXTENSIONS, PRESENTATION_EXTENSIONS, and SPREADSHEET_EXTENSIONS. If the input is not a PDF, it calls convert_to_pdf (conversion.rs#L44-L81), which selects ConversionTool::LibreOffice for Office documents.
Tool Detection and Execution
Before spawning the process, find_libre_office_command (conversion.rs#L60-L71) locates the binary across platforms—checking $PATH for libreoffice or soffice, probing macOS bundles at /Applications/LibreOffice.app/, and checking Windows defaults at C:\Program Files\Libreoffice\program\soffice.exe.
The convert_office_document function (conversion.rs#L102-L149) executes LibreOffice in headless mode with --headless --convert-to pdf. It creates a unique temporary user profile using tempfile::Builder to avoid the built-in profile lock that typically causes contention during concurrent conversions.
PDF Retrieval
After LibreOffice writes the output, find_pdf_in_dir (conversion.rs#L150-L176) scans the temporary directory to locate the generated PDF. LibreOffice may sanitize filenames during conversion, so this function handles name mapping before returning the PDF path to the parser for standard PDFium extraction.
Supported File Formats
LiteParse handles the following non-PDF formats through LibreOffice conversion:
- Word documents: DOCX, DOC, ODT
- Excel spreadsheets: XLSX, XLS, ODS
- PowerPoint presentations: PPTX, PPT, ODP
Image-only inputs route to ImageMagick instead, while plain-text files bypass conversion entirely.
Implementation Examples
Rust
Pass a DOCX path directly to parse. The conversion happens asynchronously before PDFium extraction:
use liteparse::LiteParse;
use liteparse::config::LiteParseConfig;
#[tokio::main]
async fn main() -> Result<(), liteparse::error::LiteParseError> {
let cfg = LiteParseConfig::default();
let lp = LiteParse::new(cfg);
// Automatically triggers LibreOffice conversion for DOCX
let result = lp.parse("report.docx").await?;
println!("Extracted {} pages", result.pages.len());
Ok(())
}
Python
The Python bindings expose the same functionality through parse_file:
from liteparse import LiteParse
lp = LiteParse()
result = lp.parse_file("data.xlsx") # Converts XLSX to PDF first
print(f"Pages: {len(result.pages)}")
print(result.text)
Node.js
The TypeScript wrapper forwards paths to the native Rust implementation:
import { LiteParse } from "liteparse";
const lp = new LiteParse({});
const { text } = await lp.parseFile("presentation.pptx");
console.log(text);
CLI
Process Office documents directly from the command line:
liteparse parse meeting_notes.docx
Error Handling and Platform Support
If LibreOffice is not found, convert_office_document returns a LiteParseError::Conversion with a clear installation message. The conversion runs within a Tokio async task (execute_command) but blocks only for the external LibreOffice process; subsequent OCR and layout reconstruction proceed concurrently after the PDF is produced.
On macOS, the library automatically detects the application bundle path. On Windows, it checks standard installation directories. Linux systems require libreoffice or soffice in $PATH.
Summary
- LiteParse handles DOCX, XLSX, and PPTX via transparent LibreOffice conversion in
crates/liteparse/src/conversion.rs - The
resolve_pdf_inputfunction routes non-PDFs toconvert_to_pdf, which spawns LibreOffice with--headless --convert-to pdf - Temporary user profiles prevent lock contention during concurrent conversions
- Cross-platform detection automatically finds LibreOffice binaries on macOS, Windows, and Linux
- All conversion steps are hidden from the API consumer—simply pass the original Office file path to
LiteParse::parse
Frequently Asked Questions
What happens if LibreOffice is not installed?
If find_libre_office_command fails to locate the binary, convert_office_document returns a LiteParseError::Conversion error with a message instructing you to install LibreOffice. The library checks for libreoffice, soffice, and platform-specific default paths before failing.
Is the conversion process thread-safe?
Yes. Each conversion creates an isolated temporary user profile directory using tempfile::Builder, preventing the profile locks that typically prevent concurrent LibreOffice instances. This allows multiple LiteParse::parse calls to run in parallel without collision.
Can I disable automatic conversion for specific file types?
No. Conversion is intrinsic to the parsing pipeline for non-PDF formats. If you need to handle Office files without conversion, you would need to convert them externally before passing them to LiteParse, though this bypasses the integrated error handling in crates/liteparse/src/conversion.rs.
Why does the converted PDF sometimes have a different filename?
LibreOffice sanitizes filenames during conversion to ensurePDF compatibility. The find_pdf_in_dir function (conversion.rs#L150-L176) accounts for this by scanning the output directory for any .pdf file rather than expecting a specific name match.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →