Document Text Extraction Pipeline for PDF and DOCX Processing in Macro
Macro converts uploaded PDF and DOCX files into searchable text through a multi-service AWS pipeline that uses pdfium-render for extraction, tiktoken-rs for tokenization, and persists sentence-level references with bounding box metadata to PostgreSQL before indexing in OpenSearch.
The macro-inc/macro repository implements this robust document text extraction pipeline using Rust-based Lambda services. The architecture handles both native PDFs and Microsoft Word documents through a unified nine-stage workflow deployed on AWS infrastructure.
Architecture Overview
The pipeline consists of nine distinct stages orchestrated across multiple Lambda services. All components share the DocumentKey abstraction defined in crates/model/src/key/document_key.rs, which standardizes how the system identifies original uploads, converted PDFs, and metadata objects in S3. The entry point logic resides in extract_text_from_document (lines 65-82), which uses the DocumentKey::is_converted_pdf() helper to determine processing paths.
Step 1: Document Ingestion and S3 Storage
When a user uploads a file, the document_upload_finalizer_handler receives the request and writes the raw bytes to the documents S3 bucket. This service acts as the entry point for all supported file types, persisting the original binary regardless of format.
Implementation: services/document_upload_finalizer_handler/src/inbound/s3_notification.rs
Step 2: DOCX-to-PDF Conversion
For DOCX uploads, the docx_unzip_handler Lambda triggers next. This service treats the DOCX as a zip archive, extracts its contents, and invokes the convert_service to generate a temporary PDF representation using LibreOffice or docx2pdf tooling. The system writes this converted file to S3 with a converted.pdf suffix, identifiable by downstream stages.
Key files:
services/docx_unzip_handler/src/service/document/process_document.rsservices/convert_service/src/process/convert.rs
The DocumentKey::is_converted_pdf() helper in crates/model/src/key/document_key.rs distinguishes these converted keys from original uploads.
Step 3: PDF Text Extraction
The document_text_extractor service processes both native PDFs and converted DOCX-derived PDFs. Using the pdfium-render library, the extract_text_from_document function loads each page and extracts raw text alongside precise bounding rectangles for spatial indexing.
Source locations:
- Main handler:
services/document_text_extractor/src/handler/extract_text_citations.rs - PDF parsing utilities:
services/search_processing_service/src/parsers/pdf.rs
Step 4: Sentence Segmentation and Reference Generation
Within ref_id_extract_text, the pipeline groups extracted characters into logical sentences. Each sentence receives a UUID and an expanded bounding rectangle covering the complete text block. The function returns a tuple of (references, extracted_text) where references contain spatial metadata including page numbers and PDF coordinates.
Step 5: Token Counting
Before persistence, the system uses tiktoken-rs with the o200k_base encoding to calculate token counts for cost estimation. This occurs in extract_text_from_document at lines 13-15 of the handler file.
Step 6: Database Persistence
The extractor persists data through two database operations defined in services/document_text_extractor/src/service/db/document_text_parts.rs:
db.create_document_text: Writes the full extracted string to thedocument_textstabledb.insert_references: Bulk-insertsTextReferencerecords containing UUIDs, page numbers, and bounding rectangles into thedocument_referencestable
Step 7: Search Indexing
Once text and references are stored, a Kafka event triggers the search_processing_service. This service runs the OpenSearch indexing pipeline defined in services/search_processing_service/src/process/mod.rs, making the document fully searchable with citation-aware metadata.
Implementation Examples
The following Rust patterns demonstrate how to invoke extraction utilities manually:
// Manually invoke the extractor for a given S3 key
let pdfium = Rc::new(Pdfium::new(
Pdfium::bind_to_library(Pdfium::pdfium_platform_library_name_at_path(env!("PDFIUM_LIB_PATH")))
.expect("load pdfium lib"),
));
let s3 = Arc::new(service::s3::S3::new(...));
let db = Arc::new(service::db::DB::new(...));
let result = extract_text_from_document(
"documents/12345/converted.pdf", // can be original.pdf or converted.pdf
"macro-documents",
pdfium,
s3,
db,
).await?;
// Convert a DOCX key to its PDF equivalent (simplified)
let docx_key = DocumentKey::from_s3_key("documents/12345/file.docx")?;
let pdf_key = build_docx_to_pdf_converted_document_key(&docx_key);
s3.copy_object(bucket, docx_key.to_key(), pdf_key.to_key()).await?;
Summary
- Macro's pipeline uses pdfium-render for PDF parsing and tiktoken-rs with
o200k_basefor tokenization - DOCX files undergo conversion to PDF via
docx_unzip_handlerbefore extraction begins - The
document_text_extractorgenerates sentence-level UUIDs with bounding box metadata for precise citation tracking - References persist to PostgreSQL via
db.insert_referenceswhile full text stores in thedocument_textstable - Kafka events trigger final indexing in OpenSearch through the
search_processing_service
Frequently Asked Questions
How does Macro handle DOCX files differently from PDFs?
DOCX uploads trigger the docx_unzip_handler service, which converts the document to PDF using LibreOffice or docx2pdf before the extraction pipeline begins. The original DOCX remains in S3 for downloads, but only the converted PDF version undergoes text extraction and indexing. Native PDFs skip this conversion step entirely and proceed directly to the document_text_extractor.
What library does Macro use for PDF text extraction?
The system uses pdfium-render to parse PDF documents within the document_text_extractor service. This library enables extraction of both text content and precise bounding rectangle coordinates for every sentence on each page, as implemented in services/search_processing_service/src/parsers/pdf.rs.
How are text references structured and stored?
Each extracted sentence becomes a TextReference containing a UUID, page number, and bounding box coordinates. The ref_id_extract_text function generates these objects, and db.insert_references bulk-inserts them into the document_references table alongside the full text stored via db.create_document_text in services/document_text_extractor/src/service/db/document_text_parts.rs.
What triggers the final search indexing step?
Once the document_text_extractor successfully persists text and references to PostgreSQL, it emits a Kafka event that triggers the search_processing_service. This service executes the OpenSearch indexing pipeline defined in services/search_processing_service/src/process/mod.rs, completing the document processing workflow.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →