# How Macro's Document Processing Pipeline Works: From Upload to Search Indexing

> Discover Macro's document processing pipeline. Learn how binary data becomes searchable text through S3, Lambda workers, and OpenSearch indexing.

- Repository: [Macro/macro](https://github.com/macro-inc/macro)
- Tags: internals
- Published: 2026-08-16

---

**Macro converts uploaded files into searchable text through an event‑driven, micro‑services pipeline that streams binary data to S3, extracts content via Lambda workers, and indexes results into OpenSearch.**

The document processing pipeline in the [macro‑inc/macro](https://github.com/macro-inc/macro) repository handles everything from the moment a user drops a file into the web UI to the point where that document appears in search results. This article breaks down each step using actual source paths and implementation details from the codebase.

## Step 1: Client Upload via Presigned S3 URL

The pipeline begins in the browser. When a user selects a file, the **web client** requests a presigned URL from the backend and streams the binary directly to S3.

In [`apps/web/src/lib/service-clients/service-storage/util/upload.ts`](https://github.com/macro-inc/macro/blob/main/apps/web/src/lib/service-clients/service-storage/util/upload.ts), the `uploadFile` function orchestrates this:

```typescript
// apps/web/src/lib/service-clients/service-storage/util/upload.ts
import { uploadToPresignedUrl } from './uploadToPresignedUrl';

export async function uploadFile(file: File, folderId: string) {
  // 1️⃣ Ask the backend for a presigned URL
  const { url, key } = await fetch(
    `/api/storage/${folderId}/presigned-upload`,
    { method: 'POST', body: JSON.stringify({ name: file.name }) }
  ).then(r => r.json());

  // 2️⃣ Stream the file to S3
  await uploadToPresignedUrl(url, file);

  // 3️⃣ Return the S3 key so the UI can display the new document
  return key;
}

```

This approach keeps the Macro application servers out of the data path—handling large files without consuming application memory or bandwidth.

## Step 2: Metadata Persistence in MacroDB

Once the file lands in the **document‑storage‑service** S3 bucket, the API records its metadata. The [`simple_save.rs`](https://github.com/macro-inc/macro/blob/main/simple_save.rs) module in the document storage service creates the database entry:

```rust
// services/document_storage_service/src/api/documents/simple_save.rs
// Inserts a row into the Document table with owner, type, and S3 key

```

This write to **MacroDB** establishes the document's identity and triggers the next processing stage.

## Step 3: Queueing the Extraction Job

After persistence, a **DocumentUploadJob** is enqueued on the **DocumentTextExtractorQueue**—an SQS queue defined in [`crates/macro_queues/src/lib.rs`](https://github.com/macro-inc/macro/blob/main/crates/macro_queues/src/lib.rs):

```rust
// crates/macro_queues/src/lib.rs
pub async fn enqueue_document_text_extractor(
    &self,
    s3_key: &str,
    document_id: i64,
) -> anyhow::Result<()> {
    let msg = serde_json::json!({
        "s3_key": s3_key,
        "document_id": document_id,
    });
    self.sqs
        .send_message(&self.document_text_extractor_queue, msg.to_string())
        .await
}

```

The queue definition lives in the shared **macro_queues** crate, ensuring consistent queue semantics across services.

## Step 4: Lambda Trigger and File Download

The **document‑text‑extractor** Lambda wakes when SQS delivers the message. Its infrastructure definition in [`infra/stacks/document-text-extractor/document-text-extractor.ts`](https://github.com/macro-inc/macro/blob/main/infra/stacks/document-text-extractor/document-text-extractor.ts) bundles the **pdfium** native library and configures environment variables for S3 access.

The Lambda handler downloads the raw file using the `s3_key` from the job message.

## Step 5: Format‑Specific Text Extraction

Extraction logic varies by file type:

- **PDF** → processed with **pdfium** via the bundled native library in [`services/document_text_extractor/src/main.rs`](https://github.com/macro-inc/macro/blob/main/services/document_text_extractor/src/main.rs)
- **DOCX** → unzipped by the **docx_unzip_handler** service, then parsed
- **Plain‑text** → read directly without transformation

The extractor emits UTF‑8 text chunks suitable for database storage and search indexing. The core extraction flow appears in:

```rust
// services/document_text_extractor/src/main.rs
#[tokio::main]
async fn main() -> anyhow::Result<()> {
    let event = sqs::receive_message().await?;
    let job: DocumentUploadJob = serde_json::from_str(&event.body)?;
    let raw = s3::download(&job.s3_key).await?;
    let text = pdfium::extract_text(&raw)?;
    db::store_extracted_text(job.document_id, &text).await?;
    search::queue_indexing(job.document_id).await?;
    Ok(())
}

```

DOCX handling splits across services: [`services/docx_unzip_handler/src/service/document/notify_docx_upload_job.rs`](https://github.com/macro-inc/macro/blob/main/services/docx_unzip_handler/src/service/document/notify_docx_upload_job.rs) manages the unzip notification flow.

## Step 6: Storing Extracted Text

The processed text is persisted in two places:

1. **MacroDB** as **DocumentChunk** records (via [`crates/documents/src/domain/upload_finalize.rs`](https://github.com/macro-inc/macro/blob/main/crates/documents/src/domain/upload_finalize.rs))
2. **S3** as a backup object for reconstruction or re‑indexing

The [`upload_finalize.rs`](https://github.com/macro-inc/macro/blob/main/upload_finalize.rs) module in the documents crate coordinates this dual write.

## Step 7: OpenSearch Indexing

With text extracted, a **SearchEvent** enters the **search‑upload** queue. The **search‑upload** Lambda—defined in [`infra/stacks/search-upload/search-upload-lambda.ts`](https://github.com/macro-inc/macro/blob/main/infra/stacks/search-upload/search-upload-lambda.ts)—consumes this event and pushes the document into OpenSearch:

```typescript
// infra/stacks/search-upload/search-upload-lambda.ts
import { OpenSearchClient } from '@opensearch-project/client';
import { getDocumentChunks } from './document-store';

export const handler = async (event: any) => {
  const docId = event.document_id;
  const chunks = await getDocumentChunks(docId);
  const client = new OpenSearchClient({ node: process.env.OPENSEARCH_ENDPOINT });
  await client.index({
    index: 'documents',
    id: `${docId}`,
    body: { content: chunks.join('\n') },
  });
};

```

Field‑level analyzers in the OpenSearch mapping enable full‑text search, highlighting, and faceting.

## Step 8: Querying the Search Index

Once indexed, the document becomes searchable. The **lexical‑service** exposes an HTTP endpoint that forwards queries to OpenSearch and returns highlighted results:

```typescript
// services/lexical-service/src/endpoints/search-text.ts
// The HTTP endpoint that forwards queries to OpenSearch

```

Frontend search requests hit this **Axum**‑powered API, which translates user queries into OpenSearch DSL and formats the response.

## Key Architecture Decisions

Understanding why Macro structured its document processing this way reveals important engineering trade‑offs:

| Decision | Implementation | Benefit |
|----------|---------------|---------|
| **Event‑driven async** | SQS between every stage | Natural back‑pressure, horizontal scaling, independent service deployment |
| **Native PDF processing** | pdfium bundled in Lambda | Fast, high‑fidelity text extraction without external API dependencies |
| **Dual storage** | MacroDB + S3 for extracted text | Durability, re‑indexing capability without re‑extraction |
| **OpenSearch as search store** | Dedicated indexing Lambda | Dedicated indexing path, optimized for query performance |

## Summary

- **Presigned URLs** in [`apps/web/src/lib/service-clients/service-storage/util/upload.ts`](https://github.com/macro-inc/macro/blob/main/apps/web/src/lib/service-clients/service-storage/util/upload.ts) enable direct browser‑to‑S3 uploads
- **Document metadata** is written by [`services/document_storage_service/src/api/documents/simple_save.rs`](https://github.com/macro-inc/macro/blob/main/services/document_storage_service/src/api/documents/simple_save.rs) before extraction begins
- **SQS queues** defined in [`crates/macro_queues/src/lib.rs`](https://github.com/macro-inc/macro/blob/main/crates/macro_queues/src/lib.rs) orchestrate the asynchronous pipeline
- **pdfium** powers PDF extraction in [`services/document_text_extractor/src/main.rs`](https://github.com/macro-inc/macro/blob/main/services/document_text_extractor/src/main.rs)
- **DOCX files** pass through the dedicated `docx_unzip_handler` service
- **Extracted text** is finalized via [`crates/documents/src/domain/upload_finalize.rs`](https://github.com/macro-inc/macro/blob/main/crates/documents/src/domain/upload_finalize.rs)
- **OpenSearch indexing** happens in [`infra/stacks/search-upload/search-upload-lambda.ts`](https://github.com/macro-inc/macro/blob/main/infra/stacks/search-upload/search-upload-lambda.ts)
- **Search queries** are served by [`services/lexical-service/src/endpoints/search-text.ts`](https://github.com/macro-inc/macro/blob/main/services/lexical-service/src/endpoints/search-text.ts)

## Frequently Asked Questions

### What queue system does Macro use for document processing?

Macro uses **Amazon SQS** for all inter‑service communication in the document processing pipeline. The `DocumentTextExtractorQueue` and search‑upload queue are defined in [`crates/macro_queues/src/lib.rs`](https://github.com/macro-inc/macro/blob/main/crates/macro_queues/src/lib.rs) and consumed by Lambda functions. This event‑driven approach decouples services and provides automatic retries for failed extractions.

### How does Macro handle different file formats?

Each format has a dedicated code path. **PDFs** are processed with the **pdfium** native library bundled into the document‑text‑extractor Lambda. **DOCX files** are routed through a separate **docx_unzip_handler** service that unzips the archive and extracts the document XML. Plain‑text files bypass transformation and are read directly. This modular approach allows adding new format handlers without modifying existing code.

### Where does Macro store extracted document text?

Extracted text is stored in **two locations**: as **DocumentChunk** records in **MacroDB** for operational queries, and as objects in **S3** for durability and potential re‑indexing. The [`upload_finalize.rs`](https://github.com/macro-inc/macro/blob/main/upload_finalize.rs) module in the documents crate coordinates these writes before triggering the search indexing step.

### Which search engine powers Macro's document search?

**OpenSearch** indexes all extracted document content. The `search‑upload` Lambda pushes text chunks into an OpenSearch index with configured analyzers, while the **lexical‑service** API in [`services/lexical-service/src/endpoints/search-text.ts`](https://github.com/macro-inc/macro/blob/main/services/lexical-service/src/endpoints/search-text.ts) forwards user queries and returns highlighted results.