How LiteParse Converts DOCX, XLSX, and PPTX to PDF Before Parsing

LiteParse converts DOCX, XLSX, and PPTX files to PDF using LibreOffice in headless mode, creating temporary directories for isolated processing before extracting the generated PDF for parsing.

The run-llama/liteparse library handles Microsoft Office documents by first transforming them into PDF format. This conversion step is mandatory because LiteParse's core extraction engine is optimized for PDF structure analysis. All conversion logic is centralized in crates/liteparse/src/conversion.rs, which orchestrates the transformation without requiring external cloud services.

The Conversion Pipeline for Office Documents

When you pass a DOCX, XLSX, or PPTX file to LiteParse, the library executes a six-step conversion pipeline. Each step ensures reliable PDF generation while maintaining system isolation.

File Type Detection

LiteParse maintains three static extension lists in crates/liteparse/src/conversion.rs to categorize office documents. The OFFICE_EXTENSIONS slice defines DOCX eligibility, while PRESENTATION_EXTENSIONS covers PPTX and SPREADSHEET_EXTENSIONS handles XLSX. When convert_to_pdf receives a file path, it compares the extension against these lists to determine if LibreOffice conversion is required.

Conversion Tool Selection

If the extension matches any office category, LiteParse selects the ConversionTool::LibreOffice variant. This selection occurs at lines 68-73 of the conversion module, where the library branches between different conversion backends based on document type.

Temporary Workspace Creation

Before invoking LibreOffice, LiteParse allocates a fresh temporary directory using tempfile::Builder::new(). This creates a TempDir guard that automatically cleans up the generated PDF and intermediate files when dropped. The temporary directory serves as both the workspace and output destination for the conversion process.

Headless LibreOffice Execution

The convert_office_document function constructs a command-line invocation with strict isolation parameters. The library executes LibreOffice with the following flags:

  • --headless --invisible - Runs without GUI components
  • --convert-to pdf - Specifies PDF output format
  • --outdir <temp_dir> - Directs output to the temporary directory
  • -env:UserInstallation=file://<profile_path> - Creates a per-process profile to isolate concurrent invocations

For password-protected documents, LiteParse appends the --infilter= argument to pass decryption credentials. The command construction and execution logic resides at lines 102-138 of crates/liteparse/src/conversion.rs.

PDF Discovery and Return

LibreOffice may sanitize filenames during conversion. To handle this, LiteParse calls find_pdf_in_dir (lines 150-170), which scans the temporary output directory and returns the first file with a .pdf extension. The function then packages this path into a ConversionResult containing both the pdf_path and original extension metadata.

The resolve_pdf_input function coordinates this entire flow, propagating the PDF path back to the main parsing pipeline for text extraction. This ensures that downstream processors receive a standardized PDF input regardless of the original Office format.

Implementation Code Examples

You can trigger this conversion automatically through multiple interfaces. The conversion happens transparently whenever you provide a non-PDF document path.

Rust API

use liteparse::parser::LiteParse;
use liteparse::types::PdfInput;

#[tokio::main]
async fn main() -> Result<(), liteparse::error::LiteParseError> {
    // Provide the DOCX path; LiteParse will convert it to a PDF internally.
    let input = PdfInput::Path("examples/sample.docx".into());

    // Parse the PDF (conversion happens automatically).
    let parser = LiteParse::new().await?;
    let result = parser.parse(input).await?;

    println!("Extracted {} text items", result.items.len());
    Ok(())
}

Command Line Interface


# The CLI automatically runs conversion when a non-PDF is supplied.

liteparse parse examples/presentation.pptx --output json > out.json

Node.js Integration

import { LiteParse } from "liteparse";

(async () => {
  const parser = await LiteParse.create();
  const result = await parser.parse("report.xlsx"); // XLSX → PDF internally
  console.log(result.text); // plain-text extraction
})();

Core Source Files

The conversion system spans several key files in the repository:

Summary

  • LiteParse converts DOCX, XLSX, and PPTX to PDF using LibreOffice in headless mode before parsing.
  • The conversion pipeline uses temporary directories (TempDir) for automatic cleanup and isolation.
  • File type detection relies on static extension lists (OFFICE_EXTENSIONS, PRESENTATION_EXTENSIONS, SPREADSHEET_EXTENSIONS) in conversion.rs.
  • The convert_office_document function executes LibreOffice with --headless --invisible --convert-to pdf flags and per-process user profiles.
  • If LibreOffice is not installed, the library returns LiteParseError::Conversion with installation instructions.
  • The conversion happens automatically across all LiteParse interfaces: Rust API, CLI, Node.js, and Python.

Frequently Asked Questions

Does LiteParse require an internet connection to convert documents?

No. LiteParse performs all conversions locally using an installed LibreOffice binary. The convert_office_document function in crates/liteparse/src/conversion.rs executes the soffice command entirely offline, making the conversion process suitable for air-gapped environments.

What happens if LibreOffice is not installed on the system?

LiteParse returns a LiteParseError::Conversion error with a message prompting you to install LibreOffice. The library checks for the binary availability before attempting execution, allowing you to handle the dependency error gracefully in your application code.

Can LiteParse convert password-protected Office documents?

Yes. When parsing encrypted DOCX, XLSX, or PPTX files, LiteParse passes the password to LibreOffice using the --infilter= command-line argument within the convert_office_document function. You must provide the password when initializing the parse operation.

Does the conversion preserve the original document formatting?

LiteParse relies on LibreOffice's PDF export capabilities, which generally preserve formatting, fonts, and layouts. However, the library's primary goal is text extraction accuracy rather than visual fidelity, as the generated PDF serves only as an intermediate format for the parsing pipeline.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →