# How Macro's Document Text Extractor Processes PDF and DOCX Files

> Learn how Macro's Rust-based text extractor processes PDF and DOCX files using pdfium-render and XML parsing for efficient text extraction from documents.

- Repository: [Macro/macro](https://github.com/macro-inc/macro)
- Tags: how-to-guide
- Published: 2026-08-18

---

**Macro's document text extractor is a Rust-based AWS Lambda service that extracts plain text from PDF files using the pdfium-render library and from DOCX files by parsing the underlying XML from the ZIP archive.**

The document text extractor in the [macro-inc/macro](https://github.com/macro-inc/macro) repository is a serverless component designed to convert uploaded documents into searchable plain text. It handles both **PDF** and **DOCX** formats through distinct processing pipelines, storing the extracted content for downstream search indexing. This article examines the implementation details of how this **document text extractor** processes these file formats using native libraries and custom XML parsers.

## PDF Processing with pdfium-render

The **PDF** pipeline relies on the `pdfium-render` library, which is compiled as a native binary and bundled with the Lambda deployment package.

### Library Initialization and Binding

When the Lambda cold-starts, the system initializes the PDFium library using a compile-time constant. In [`services/document_text_extractor/src/main.rs`](https://github.com/macro-inc/macro/blob/main/services/document_text_extractor/src/main.rs) (lines 12-20), the code sets up the library path and binds to the native binary:

- The [`build.rs`](https://github.com/macro-inc/macro/blob/main/build.rs) script verifies that `pdfium-lib/linux/libpdfium.so` exists at compile time
- The constant `PDFIUM_LIB_PATH` stores the path to the native library
- `Pdfium::bind_to_library` loads `libpdfium.so` (Linux) or `libpdfium.dylib` (macOS) at runtime

This initialization ensures the `pdfium-render` crate can parse PDF byte streams without external dependencies.

### Page-by-Page Text Extraction

Once initialized, the extractor processes raw PDF bytes from S3. The logic resides in [`services/search_processing_service/src/parsers/pdf.rs`](https://github.com/macro-inc/macro/blob/main/services/search_processing_service/src/parsers/pdf.rs):

1. `Pdfium::load_pdf_from_byte_vec` parses the byte vector into a `PdfDocument` object
2. The code iterates over every page in the document
3. Each page's text is extracted via `page.get_text()`
4. Results are concatenated into a single UTF-8 string

Errors during extraction are logged using `tracing::error!` (see lines 57-62 of the PDF parser), ensuring visibility into malformed documents or library failures.

## DOCX Processing via XML Parsing

Unlike PDFs, **DOCX** files are ZIP archives containing XML markup. The extractor treats them as compressed file systems rather than binary documents.

### Archive Extraction

The DOCX pipeline begins in [`services/search_processing_service/src/parsers/docx.rs`](https://github.com/macro-inc/macro/blob/main/services/search_processing_service/src/parsers/docx.rs) (lines 15-22). The implementation uses the `zip` crate to:

- Unzip the incoming byte stream
- Locate [`word/document.xml`](https://github.com/macro-inc/macro/blob/main/word/document.xml) within the archive structure
- Read the XML content into a string buffer

### Text Node Parsing

The core parsing logic walks the XML DOM and extracts text from `<w:t>` elements, which contain the actual character data in Word documents:

- The parser ignores formatting tags and styling attributes
- It concatenates the inner text of each `<w:t>` node sequentially
- The function returns an `anyhow::Result<String>` containing the full document text

This approach avoids heavy office-suite dependencies, keeping the Lambda package lightweight.

## Request Dispatch and Service Integration

The **document text extractor** operates as an S3-triggered Lambda function. The dispatcher in [`services/document_text_extractor/src/handler/extract_text_citations.rs`](https://github.com/macro-inc/macro/blob/main/services/document_text_extractor/src/handler/extract_text_citations.rs) (lines 145-150) selects the appropriate pipeline based on MIME type or stored metadata:

```rust
match file_type.as_str() {
    "pdf" => extract_pdf_text(&bytes).await,
    "docx" => extract_docx_text(&bytes).await,
    _ => Err(anyhow::anyhow!("unsupported file type")),
}?;

```

The handler in [`services/document_text_extractor/src/handler/mod.rs`](https://github.com/macro-inc/macro/blob/main/services/document_text_extractor/src/handler/mod.rs) (lines 20-30) receives S3 events, downloads the object bytes, and invokes the extractor. After processing, the extracted text is persisted to the **document-search** table for full-text indexing.

## Implementation Examples

**Extracting PDF text with pdfium-render:**

```rust
use pdfium_render::prelude::*;

let pdfium_binary = env!("PDFIUM_LIB_PATH");
let pdfium = Pdfium::new(
    Pdfium::bind_to_library(Pdfium::pdfium_platform_library_name_at_path(
        pdfium_binary,
    ))
    .expect("unable to bind to pdfium library"),
);
let document = pdfium.load_pdf_from_byte_vec(pdf_bytes, None)?;

// Collect text from each page
let mut full_text = String::new();
for page in document.pages() {
    full_text.push_str(&page.get_text()?);
}

```

**Extracting DOCX text from ZIP archive:**

```rust
use zip::ZipArchive;
use std::io::Read;

fn extract_docx_text(bytes: &[u8]) -> anyhow::Result<String> {
    let mut archive = ZipArchive::new(std::io::Cursor::new(bytes))?;
    let mut xml = String::new();
    archive.by_name("word/document.xml")?.read_to_string(&mut xml)?;
    Ok(parse_docx(&xml)?) // parse_docx lives in docx.rs
}

```

**Lambda handler wiring:**

```rust
// services/document_text_extractor/src/main.rs
// Initializes the Lambda runtime with pdfium bindings
let pdfium = initialize_pdfium()?;
lambda_runtime::run(service_fn(|event: S3Event| {
    handler::process_event(event, pdfium.clone())
})).await?;

```

The deployment configuration in [`infra/stacks/document-text-extractor/document-text-extractor.ts`](https://github.com/macro-inc/macro/blob/main/infra/stacks/document-text-extractor/document-text-extractor.ts) ensures the `libpdfium.so` binary is included in the Lambda deployment package, while the [`build.rs`](https://github.com/macro-inc/macro/blob/main/build.rs) script validates the library presence during compilation.

## Summary

- Macro's document text extractor is a **Rust Lambda** that processes PDF and DOCX files through specialized pipelines
- **PDFs** are handled by the `pdfium-render` library, which binds to a native `libpdfium.so` binary and extracts text page-by-page
- **DOCX** files are treated as ZIP archives; the extractor reads [`word/document.xml`](https://github.com/macro-inc/macro/blob/main/word/document.xml) and parses `<w:t>` nodes for text content
- The dispatcher in [`extract_text_citations.rs`](https://github.com/macro-inc/macro/blob/main/extract_text_citations.rs) routes files based on MIME type, storing results in the document-search table
- All error handling uses `anyhow::Result` with `tracing` for observability

## Frequently Asked Questions

### What library does Macro use to extract text from PDFs?

Macro uses the **pdfium-render** crate, which provides Rust bindings to the PDFium library (Google's open-source PDF rendering engine). The Lambda bundles `libpdfium.so` for Linux environments and initializes it via `Pdfium::bind_to_library` at startup.

### How does the extractor handle corrupted or password-protected DOCX files?

The DOCX parser in [`services/search_processing_service/src/parsers/docx.rs`](https://github.com/macro-inc/macro/blob/main/services/search_processing_service/src/parsers/docx.rs) returns `anyhow::Result<String>`. If the ZIP archive is malformed or [`word/document.xml`](https://github.com/macro-inc/macro/blob/main/word/document.xml) is missing, the `zip` crate or XML parser will return an error that propagates up to the handler, which logs the failure via `tracing::error!` and returns an error response without crashing the Lambda.

### Why does Macro use different extraction strategies for PDF versus DOCX?

PDFs are binary formats with complex rendering engines, requiring a native library like PDFium to handle text positioning, fonts, and encodings correctly. DOCX files are essentially ZIPped XML, making them straightforward to parse with standard compression and XML libraries without heavy dependencies, resulting in faster cold starts and smaller bundle sizes for the Lambda function.

### Where is the extracted text stored after processing?

After extraction, the plain text is persisted to the **document-search** table (referenced in the handler logic) and indexed for full-text search capabilities within the Macro platform. The S3 event trigger provides the object key, and the Lambda writes the extracted content back to the database for downstream consumption.