How LiteParse Converts DOCX, XLSX, and PPTX to PDF Before Parsing
LiteParse converts DOCX, XLSX, and PPTX files to PDF using LibreOffice in headless mode, creating temporary directories for isolated processing before extracting the generated PDF for parsing.
The run-llama/liteparse library handles Microsoft Office documents by first transforming them into PDF format. This conversion step is mandatory because LiteParse's core extraction engine is optimized for PDF structure analysis. All conversion logic is centralized in crates/liteparse/src/conversion.rs, which orchestrates the transformation without requiring external cloud services.
The Conversion Pipeline for Office Documents
When you pass a DOCX, XLSX, or PPTX file to LiteParse, the library executes a six-step conversion pipeline. Each step ensures reliable PDF generation while maintaining system isolation.
File Type Detection
LiteParse maintains three static extension lists in crates/liteparse/src/conversion.rs to categorize office documents. The OFFICE_EXTENSIONS slice defines DOCX eligibility, while PRESENTATION_EXTENSIONS covers PPTX and SPREADSHEET_EXTENSIONS handles XLSX. When convert_to_pdf receives a file path, it compares the extension against these lists to determine if LibreOffice conversion is required.
Conversion Tool Selection
If the extension matches any office category, LiteParse selects the ConversionTool::LibreOffice variant. This selection occurs at lines 68-73 of the conversion module, where the library branches between different conversion backends based on document type.
Temporary Workspace Creation
Before invoking LibreOffice, LiteParse allocates a fresh temporary directory using tempfile::Builder::new(). This creates a TempDir guard that automatically cleans up the generated PDF and intermediate files when dropped. The temporary directory serves as both the workspace and output destination for the conversion process.
Headless LibreOffice Execution
The convert_office_document function constructs a command-line invocation with strict isolation parameters. The library executes LibreOffice with the following flags:
--headless --invisible- Runs without GUI components--convert-to pdf- Specifies PDF output format--outdir <temp_dir>- Directs output to the temporary directory-env:UserInstallation=file://<profile_path>- Creates a per-process profile to isolate concurrent invocations
For password-protected documents, LiteParse appends the --infilter= argument to pass decryption credentials. The command construction and execution logic resides at lines 102-138 of crates/liteparse/src/conversion.rs.
PDF Discovery and Return
LibreOffice may sanitize filenames during conversion. To handle this, LiteParse calls find_pdf_in_dir (lines 150-170), which scans the temporary output directory and returns the first file with a .pdf extension. The function then packages this path into a ConversionResult containing both the pdf_path and original extension metadata.
The resolve_pdf_input function coordinates this entire flow, propagating the PDF path back to the main parsing pipeline for text extraction. This ensures that downstream processors receive a standardized PDF input regardless of the original Office format.
Implementation Code Examples
You can trigger this conversion automatically through multiple interfaces. The conversion happens transparently whenever you provide a non-PDF document path.
Rust API
use liteparse::parser::LiteParse;
use liteparse::types::PdfInput;
#[tokio::main]
async fn main() -> Result<(), liteparse::error::LiteParseError> {
// Provide the DOCX path; LiteParse will convert it to a PDF internally.
let input = PdfInput::Path("examples/sample.docx".into());
// Parse the PDF (conversion happens automatically).
let parser = LiteParse::new().await?;
let result = parser.parse(input).await?;
println!("Extracted {} text items", result.items.len());
Ok(())
}
Command Line Interface
# The CLI automatically runs conversion when a non-PDF is supplied.
liteparse parse examples/presentation.pptx --output json > out.json
Node.js Integration
import { LiteParse } from "liteparse";
(async () => {
const parser = await LiteParse.create();
const result = await parser.parse("report.xlsx"); // XLSX → PDF internally
console.log(result.text); // plain-text extraction
})();
Core Source Files
The conversion system spans several key files in the repository:
crates/liteparse/src/conversion.rs- Contains the core conversion logic, includingconvert_office_document,find_pdf_in_dir, and the extension detection slices.crates/liteparse/src/parser.rs- Houses the high-levelLiteParsestruct that callsresolve_pdf_inputbefore PDF extraction.crates/liteparse/src/types.rs- Defines thePdfInputenum for accepting file paths or raw bytes.packages/node/src/lib.ts- Node.js bindings that forward file paths to the Rust core.packages/python/liteparse/parser.py- Python wrapper exposing automatic conversion capabilities.
Summary
- LiteParse converts DOCX, XLSX, and PPTX to PDF using LibreOffice in headless mode before parsing.
- The conversion pipeline uses temporary directories (
TempDir) for automatic cleanup and isolation. - File type detection relies on static extension lists (
OFFICE_EXTENSIONS,PRESENTATION_EXTENSIONS,SPREADSHEET_EXTENSIONS) inconversion.rs. - The
convert_office_documentfunction executes LibreOffice with--headless --invisible --convert-to pdfflags and per-process user profiles. - If LibreOffice is not installed, the library returns
LiteParseError::Conversionwith installation instructions. - The conversion happens automatically across all LiteParse interfaces: Rust API, CLI, Node.js, and Python.
Frequently Asked Questions
Does LiteParse require an internet connection to convert documents?
No. LiteParse performs all conversions locally using an installed LibreOffice binary. The convert_office_document function in crates/liteparse/src/conversion.rs executes the soffice command entirely offline, making the conversion process suitable for air-gapped environments.
What happens if LibreOffice is not installed on the system?
LiteParse returns a LiteParseError::Conversion error with a message prompting you to install LibreOffice. The library checks for the binary availability before attempting execution, allowing you to handle the dependency error gracefully in your application code.
Can LiteParse convert password-protected Office documents?
Yes. When parsing encrypted DOCX, XLSX, or PPTX files, LiteParse passes the password to LibreOffice using the --infilter= command-line argument within the convert_office_document function. You must provide the password when initializing the parse operation.
Does the conversion preserve the original document formatting?
LiteParse relies on LibreOffice's PDF export capabilities, which generally preserve formatting, fonts, and layouts. However, the library's primary goal is text extraction accuracy rather than visual fidelity, as the generated PDF serves only as an intermediate format for the parsing pipeline.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →