How Macro's Document Text Extractor Processes PDF and DOCX Files
Macro's document text extractor is a Rust-based AWS Lambda service that extracts plain text from PDF files using the pdfium-render library and from DOCX files by parsing the underlying XML from the ZIP archive.
The document text extractor in the macro-inc/macro repository is a serverless component designed to convert uploaded documents into searchable plain text. It handles both PDF and DOCX formats through distinct processing pipelines, storing the extracted content for downstream search indexing. This article examines the implementation details of how this document text extractor processes these file formats using native libraries and custom XML parsers.
PDF Processing with pdfium-render
The PDF pipeline relies on the pdfium-render library, which is compiled as a native binary and bundled with the Lambda deployment package.
Library Initialization and Binding
When the Lambda cold-starts, the system initializes the PDFium library using a compile-time constant. In services/document_text_extractor/src/main.rs (lines 12-20), the code sets up the library path and binds to the native binary:
- The
build.rsscript verifies thatpdfium-lib/linux/libpdfium.soexists at compile time - The constant
PDFIUM_LIB_PATHstores the path to the native library Pdfium::bind_to_libraryloadslibpdfium.so(Linux) orlibpdfium.dylib(macOS) at runtime
This initialization ensures the pdfium-render crate can parse PDF byte streams without external dependencies.
Page-by-Page Text Extraction
Once initialized, the extractor processes raw PDF bytes from S3. The logic resides in services/search_processing_service/src/parsers/pdf.rs:
Pdfium::load_pdf_from_byte_vecparses the byte vector into aPdfDocumentobject- The code iterates over every page in the document
- Each page's text is extracted via
page.get_text() - Results are concatenated into a single UTF-8 string
Errors during extraction are logged using tracing::error! (see lines 57-62 of the PDF parser), ensuring visibility into malformed documents or library failures.
DOCX Processing via XML Parsing
Unlike PDFs, DOCX files are ZIP archives containing XML markup. The extractor treats them as compressed file systems rather than binary documents.
Archive Extraction
The DOCX pipeline begins in services/search_processing_service/src/parsers/docx.rs (lines 15-22). The implementation uses the zip crate to:
- Unzip the incoming byte stream
- Locate
word/document.xmlwithin the archive structure - Read the XML content into a string buffer
Text Node Parsing
The core parsing logic walks the XML DOM and extracts text from <w:t> elements, which contain the actual character data in Word documents:
- The parser ignores formatting tags and styling attributes
- It concatenates the inner text of each
<w:t>node sequentially - The function returns an
anyhow::Result<String>containing the full document text
This approach avoids heavy office-suite dependencies, keeping the Lambda package lightweight.
Request Dispatch and Service Integration
The document text extractor operates as an S3-triggered Lambda function. The dispatcher in services/document_text_extractor/src/handler/extract_text_citations.rs (lines 145-150) selects the appropriate pipeline based on MIME type or stored metadata:
match file_type.as_str() {
"pdf" => extract_pdf_text(&bytes).await,
"docx" => extract_docx_text(&bytes).await,
_ => Err(anyhow::anyhow!("unsupported file type")),
}?;
The handler in services/document_text_extractor/src/handler/mod.rs (lines 20-30) receives S3 events, downloads the object bytes, and invokes the extractor. After processing, the extracted text is persisted to the document-search table for full-text indexing.
Implementation Examples
Extracting PDF text with pdfium-render:
use pdfium_render::prelude::*;
let pdfium_binary = env!("PDFIUM_LIB_PATH");
let pdfium = Pdfium::new(
Pdfium::bind_to_library(Pdfium::pdfium_platform_library_name_at_path(
pdfium_binary,
))
.expect("unable to bind to pdfium library"),
);
let document = pdfium.load_pdf_from_byte_vec(pdf_bytes, None)?;
// Collect text from each page
let mut full_text = String::new();
for page in document.pages() {
full_text.push_str(&page.get_text()?);
}
Extracting DOCX text from ZIP archive:
use zip::ZipArchive;
use std::io::Read;
fn extract_docx_text(bytes: &[u8]) -> anyhow::Result<String> {
let mut archive = ZipArchive::new(std::io::Cursor::new(bytes))?;
let mut xml = String::new();
archive.by_name("word/document.xml")?.read_to_string(&mut xml)?;
Ok(parse_docx(&xml)?) // parse_docx lives in docx.rs
}
Lambda handler wiring:
// services/document_text_extractor/src/main.rs
// Initializes the Lambda runtime with pdfium bindings
let pdfium = initialize_pdfium()?;
lambda_runtime::run(service_fn(|event: S3Event| {
handler::process_event(event, pdfium.clone())
})).await?;
The deployment configuration in infra/stacks/document-text-extractor/document-text-extractor.ts ensures the libpdfium.so binary is included in the Lambda deployment package, while the build.rs script validates the library presence during compilation.
Summary
- Macro's document text extractor is a Rust Lambda that processes PDF and DOCX files through specialized pipelines
- PDFs are handled by the
pdfium-renderlibrary, which binds to a nativelibpdfium.sobinary and extracts text page-by-page - DOCX files are treated as ZIP archives; the extractor reads
word/document.xmland parses<w:t>nodes for text content - The dispatcher in
extract_text_citations.rsroutes files based on MIME type, storing results in the document-search table - All error handling uses
anyhow::Resultwithtracingfor observability
Frequently Asked Questions
What library does Macro use to extract text from PDFs?
Macro uses the pdfium-render crate, which provides Rust bindings to the PDFium library (Google's open-source PDF rendering engine). The Lambda bundles libpdfium.so for Linux environments and initializes it via Pdfium::bind_to_library at startup.
How does the extractor handle corrupted or password-protected DOCX files?
The DOCX parser in services/search_processing_service/src/parsers/docx.rs returns anyhow::Result<String>. If the ZIP archive is malformed or word/document.xml is missing, the zip crate or XML parser will return an error that propagates up to the handler, which logs the failure via tracing::error! and returns an error response without crashing the Lambda.
Why does Macro use different extraction strategies for PDF versus DOCX?
PDFs are binary formats with complex rendering engines, requiring a native library like PDFium to handle text positioning, fonts, and encodings correctly. DOCX files are essentially ZIPped XML, making them straightforward to parse with standard compression and XML libraries without heavy dependencies, resulting in faster cold starts and smaller bundle sizes for the Lambda function.
Where is the extracted text stored after processing?
After extraction, the plain text is persisted to the document-search table (referenced in the handler logic) and indexed for full-text search capabilities within the Macro platform. The S3 event trigger provides the object key, and the Lambda writes the extracted content back to the database for downstream consumption.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →