How Macro's Document Text Extractor Processes PDF and DOCX Files

Macro's document text extractor is a Rust-based AWS Lambda service that extracts plain text from PDF files using the pdfium-render library and from DOCX files by parsing the underlying XML from the ZIP archive.

The document text extractor in the macro-inc/macro repository is a serverless component designed to convert uploaded documents into searchable plain text. It handles both PDF and DOCX formats through distinct processing pipelines, storing the extracted content for downstream search indexing. This article examines the implementation details of how this document text extractor processes these file formats using native libraries and custom XML parsers.

PDF Processing with pdfium-render

The PDF pipeline relies on the pdfium-render library, which is compiled as a native binary and bundled with the Lambda deployment package.

Library Initialization and Binding

When the Lambda cold-starts, the system initializes the PDFium library using a compile-time constant. In services/document_text_extractor/src/main.rs (lines 12-20), the code sets up the library path and binds to the native binary:

  • The build.rs script verifies that pdfium-lib/linux/libpdfium.so exists at compile time
  • The constant PDFIUM_LIB_PATH stores the path to the native library
  • Pdfium::bind_to_library loads libpdfium.so (Linux) or libpdfium.dylib (macOS) at runtime

This initialization ensures the pdfium-render crate can parse PDF byte streams without external dependencies.

Page-by-Page Text Extraction

Once initialized, the extractor processes raw PDF bytes from S3. The logic resides in services/search_processing_service/src/parsers/pdf.rs:

  1. Pdfium::load_pdf_from_byte_vec parses the byte vector into a PdfDocument object
  2. The code iterates over every page in the document
  3. Each page's text is extracted via page.get_text()
  4. Results are concatenated into a single UTF-8 string

Errors during extraction are logged using tracing::error! (see lines 57-62 of the PDF parser), ensuring visibility into malformed documents or library failures.

DOCX Processing via XML Parsing

Unlike PDFs, DOCX files are ZIP archives containing XML markup. The extractor treats them as compressed file systems rather than binary documents.

Archive Extraction

The DOCX pipeline begins in services/search_processing_service/src/parsers/docx.rs (lines 15-22). The implementation uses the zip crate to:

  • Unzip the incoming byte stream
  • Locate word/document.xml within the archive structure
  • Read the XML content into a string buffer

Text Node Parsing

The core parsing logic walks the XML DOM and extracts text from <w:t> elements, which contain the actual character data in Word documents:

  • The parser ignores formatting tags and styling attributes
  • It concatenates the inner text of each <w:t> node sequentially
  • The function returns an anyhow::Result<String> containing the full document text

This approach avoids heavy office-suite dependencies, keeping the Lambda package lightweight.

Request Dispatch and Service Integration

The document text extractor operates as an S3-triggered Lambda function. The dispatcher in services/document_text_extractor/src/handler/extract_text_citations.rs (lines 145-150) selects the appropriate pipeline based on MIME type or stored metadata:

match file_type.as_str() {
    "pdf" => extract_pdf_text(&bytes).await,
    "docx" => extract_docx_text(&bytes).await,
    _ => Err(anyhow::anyhow!("unsupported file type")),
}?;

The handler in services/document_text_extractor/src/handler/mod.rs (lines 20-30) receives S3 events, downloads the object bytes, and invokes the extractor. After processing, the extracted text is persisted to the document-search table for full-text indexing.

Implementation Examples

Extracting PDF text with pdfium-render:

use pdfium_render::prelude::*;

let pdfium_binary = env!("PDFIUM_LIB_PATH");
let pdfium = Pdfium::new(
    Pdfium::bind_to_library(Pdfium::pdfium_platform_library_name_at_path(
        pdfium_binary,
    ))
    .expect("unable to bind to pdfium library"),
);
let document = pdfium.load_pdf_from_byte_vec(pdf_bytes, None)?;

// Collect text from each page
let mut full_text = String::new();
for page in document.pages() {
    full_text.push_str(&page.get_text()?);
}

Extracting DOCX text from ZIP archive:

use zip::ZipArchive;
use std::io::Read;

fn extract_docx_text(bytes: &[u8]) -> anyhow::Result<String> {
    let mut archive = ZipArchive::new(std::io::Cursor::new(bytes))?;
    let mut xml = String::new();
    archive.by_name("word/document.xml")?.read_to_string(&mut xml)?;
    Ok(parse_docx(&xml)?) // parse_docx lives in docx.rs
}

Lambda handler wiring:

// services/document_text_extractor/src/main.rs
// Initializes the Lambda runtime with pdfium bindings
let pdfium = initialize_pdfium()?;
lambda_runtime::run(service_fn(|event: S3Event| {
    handler::process_event(event, pdfium.clone())
})).await?;

The deployment configuration in infra/stacks/document-text-extractor/document-text-extractor.ts ensures the libpdfium.so binary is included in the Lambda deployment package, while the build.rs script validates the library presence during compilation.

Summary

  • Macro's document text extractor is a Rust Lambda that processes PDF and DOCX files through specialized pipelines
  • PDFs are handled by the pdfium-render library, which binds to a native libpdfium.so binary and extracts text page-by-page
  • DOCX files are treated as ZIP archives; the extractor reads word/document.xml and parses <w:t> nodes for text content
  • The dispatcher in extract_text_citations.rs routes files based on MIME type, storing results in the document-search table
  • All error handling uses anyhow::Result with tracing for observability

Frequently Asked Questions

What library does Macro use to extract text from PDFs?

Macro uses the pdfium-render crate, which provides Rust bindings to the PDFium library (Google's open-source PDF rendering engine). The Lambda bundles libpdfium.so for Linux environments and initializes it via Pdfium::bind_to_library at startup.

How does the extractor handle corrupted or password-protected DOCX files?

The DOCX parser in services/search_processing_service/src/parsers/docx.rs returns anyhow::Result<String>. If the ZIP archive is malformed or word/document.xml is missing, the zip crate or XML parser will return an error that propagates up to the handler, which logs the failure via tracing::error! and returns an error response without crashing the Lambda.

Why does Macro use different extraction strategies for PDF versus DOCX?

PDFs are binary formats with complex rendering engines, requiring a native library like PDFium to handle text positioning, fonts, and encodings correctly. DOCX files are essentially ZIPped XML, making them straightforward to parse with standard compression and XML libraries without heavy dependencies, resulting in faster cold starts and smaller bundle sizes for the Lambda function.

Where is the extracted text stored after processing?

After extraction, the plain text is persisted to the document-search table (referenced in the handler logic) and indexed for full-text search capabilities within the Macro platform. The S3 event trigger provides the object key, and the Lambda writes the extracted content back to the database for downstream consumption.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →