# Document Text Extraction Pipeline for PDF and DOCX Processing in Macro

> Discover Macro's document text extraction pipeline for PDFs and DOCX. Learn how it uses AWS services like pdfium-render and tiktoken-rs to convert files into searchable text with metadata.

- Repository: [Macro/macro](https://github.com/macro-inc/macro)
- Tags: how-to-guide
- Published: 2026-08-17

---

**Macro converts uploaded PDF and DOCX files into searchable text through a multi-service AWS pipeline that uses pdfium-render for extraction, tiktoken-rs for tokenization, and persists sentence-level references with bounding box metadata to PostgreSQL before indexing in OpenSearch.**

The macro-inc/macro repository implements this robust document text extraction pipeline using Rust-based Lambda services. The architecture handles both native PDFs and Microsoft Word documents through a unified nine-stage workflow deployed on AWS infrastructure.

## Architecture Overview

The pipeline consists of nine distinct stages orchestrated across multiple Lambda services. All components share the **`DocumentKey`** abstraction defined in [`crates/model/src/key/document_key.rs`](https://github.com/macro-inc/macro/blob/main/crates/model/src/key/document_key.rs), which standardizes how the system identifies original uploads, converted PDFs, and metadata objects in S3. The entry point logic resides in `extract_text_from_document` (lines 65-82), which uses the `DocumentKey::is_converted_pdf()` helper to determine processing paths.

## Step 1: Document Ingestion and S3 Storage

When a user uploads a file, the **`document_upload_finalizer_handler`** receives the request and writes the raw bytes to the *documents* S3 bucket. This service acts as the entry point for all supported file types, persisting the original binary regardless of format.

Implementation: [`services/document_upload_finalizer_handler/src/inbound/s3_notification.rs`](https://github.com/macro-inc/macro/blob/main/services/document_upload_finalizer_handler/src/inbound/s3_notification.rs)

## Step 2: DOCX-to-PDF Conversion

For DOCX uploads, the **`docx_unzip_handler`** Lambda triggers next. This service treats the DOCX as a zip archive, extracts its contents, and invokes the **`convert_service`** to generate a temporary PDF representation using **LibreOffice** or **docx2pdf** tooling. The system writes this converted file to S3 with a `converted.pdf` suffix, identifiable by downstream stages.

Key files:
- [`services/docx_unzip_handler/src/service/document/process_document.rs`](https://github.com/macro-inc/macro/blob/main/services/docx_unzip_handler/src/service/document/process_document.rs)
- [`services/convert_service/src/process/convert.rs`](https://github.com/macro-inc/macro/blob/main/services/convert_service/src/process/convert.rs)

The `DocumentKey::is_converted_pdf()` helper in [`crates/model/src/key/document_key.rs`](https://github.com/macro-inc/macro/blob/main/crates/model/src/key/document_key.rs) distinguishes these converted keys from original uploads.

## Step 3: PDF Text Extraction

The **`document_text_extractor`** service processes both native PDFs and converted DOCX-derived PDFs. Using the **pdfium-render** library, the `extract_text_from_document` function loads each page and extracts raw text alongside precise bounding rectangles for spatial indexing.

Source locations:
- Main handler: [`services/document_text_extractor/src/handler/extract_text_citations.rs`](https://github.com/macro-inc/macro/blob/main/services/document_text_extractor/src/handler/extract_text_citations.rs)
- PDF parsing utilities: [`services/search_processing_service/src/parsers/pdf.rs`](https://github.com/macro-inc/macro/blob/main/services/search_processing_service/src/parsers/pdf.rs)

## Step 4: Sentence Segmentation and Reference Generation

Within `ref_id_extract_text`, the pipeline groups extracted characters into logical sentences. Each sentence receives a UUID and an expanded bounding rectangle covering the complete text block. The function returns a tuple of **(references, extracted_text)** where references contain spatial metadata including page numbers and PDF coordinates.

## Step 5: Token Counting

Before persistence, the system uses **tiktoken-rs** with the `o200k_base` encoding to calculate token counts for cost estimation. This occurs in `extract_text_from_document` at lines 13-15 of the handler file.

## Step 6: Database Persistence

The extractor persists data through two database operations defined in [`services/document_text_extractor/src/service/db/document_text_parts.rs`](https://github.com/macro-inc/macro/blob/main/services/document_text_extractor/src/service/db/document_text_parts.rs):

- **`db.create_document_text`**: Writes the full extracted string to the `document_texts` table
- **`db.insert_references`**: Bulk-inserts `TextReference` records containing UUIDs, page numbers, and bounding rectangles into the `document_references` table

## Step 7: Search Indexing

Once text and references are stored, a **Kafka** event triggers the **`search_processing_service`**. This service runs the **OpenSearch** indexing pipeline defined in [`services/search_processing_service/src/process/mod.rs`](https://github.com/macro-inc/macro/blob/main/services/search_processing_service/src/process/mod.rs), making the document fully searchable with citation-aware metadata.

## Implementation Examples

The following Rust patterns demonstrate how to invoke extraction utilities manually:

```rust
// Manually invoke the extractor for a given S3 key
let pdfium = Rc::new(Pdfium::new(
    Pdfium::bind_to_library(Pdfium::pdfium_platform_library_name_at_path(env!("PDFIUM_LIB_PATH")))
        .expect("load pdfium lib"),
));
let s3 = Arc::new(service::s3::S3::new(...));
let db = Arc::new(service::db::DB::new(...));

let result = extract_text_from_document(
    "documents/12345/converted.pdf",   // can be original.pdf or converted.pdf
    "macro-documents",
    pdfium,
    s3,
    db,
).await?;

```

```rust
// Convert a DOCX key to its PDF equivalent (simplified)
let docx_key = DocumentKey::from_s3_key("documents/12345/file.docx")?;
let pdf_key = build_docx_to_pdf_converted_document_key(&docx_key);
s3.copy_object(bucket, docx_key.to_key(), pdf_key.to_key()).await?;

```

## Summary

- Macro's pipeline uses **pdfium-render** for PDF parsing and **tiktoken-rs** with `o200k_base` for tokenization
- DOCX files undergo conversion to PDF via **`docx_unzip_handler`** before extraction begins
- The **`document_text_extractor`** generates sentence-level UUIDs with bounding box metadata for precise citation tracking
- References persist to PostgreSQL via `db.insert_references` while full text stores in the `document_texts` table
- **Kafka** events trigger final indexing in **OpenSearch** through the **`search_processing_service`**

## Frequently Asked Questions

### How does Macro handle DOCX files differently from PDFs?

DOCX uploads trigger the **`docx_unzip_handler`** service, which converts the document to PDF using LibreOffice or docx2pdf before the extraction pipeline begins. The original DOCX remains in S3 for downloads, but only the converted PDF version undergoes text extraction and indexing. Native PDFs skip this conversion step entirely and proceed directly to the **`document_text_extractor`**.

### What library does Macro use for PDF text extraction?

The system uses **pdfium-render** to parse PDF documents within the **`document_text_extractor`** service. This library enables extraction of both text content and precise bounding rectangle coordinates for every sentence on each page, as implemented in [`services/search_processing_service/src/parsers/pdf.rs`](https://github.com/macro-inc/macro/blob/main/services/search_processing_service/src/parsers/pdf.rs).

### How are text references structured and stored?

Each extracted sentence becomes a `TextReference` containing a UUID, page number, and bounding box coordinates. The `ref_id_extract_text` function generates these objects, and `db.insert_references` bulk-inserts them into the `document_references` table alongside the full text stored via `db.create_document_text` in [`services/document_text_extractor/src/service/db/document_text_parts.rs`](https://github.com/macro-inc/macro/blob/main/services/document_text_extractor/src/service/db/document_text_parts.rs).

### What triggers the final search indexing step?

Once the **`document_text_extractor`** successfully persists text and references to PostgreSQL, it emits a Kafka event that triggers the **`search_processing_service`**. This service executes the OpenSearch indexing pipeline defined in [`services/search_processing_service/src/process/mod.rs`](https://github.com/macro-inc/macro/blob/main/services/search_processing_service/src/process/mod.rs), completing the document processing workflow.