How Macro's Document Processing Pipeline Works: From Upload to Search Indexing

Macro converts uploaded files into searchable text through an event‑driven, micro‑services pipeline that streams binary data to S3, extracts content via Lambda workers, and indexes results into OpenSearch.

The document processing pipeline in the macro‑inc/macro repository handles everything from the moment a user drops a file into the web UI to the point where that document appears in search results. This article breaks down each step using actual source paths and implementation details from the codebase.

Step 1: Client Upload via Presigned S3 URL

The pipeline begins in the browser. When a user selects a file, the web client requests a presigned URL from the backend and streams the binary directly to S3.

In apps/web/src/lib/service-clients/service-storage/util/upload.ts, the uploadFile function orchestrates this:

// apps/web/src/lib/service-clients/service-storage/util/upload.ts
import { uploadToPresignedUrl } from './uploadToPresignedUrl';

export async function uploadFile(file: File, folderId: string) {
  // 1️⃣ Ask the backend for a presigned URL
  const { url, key } = await fetch(
    `/api/storage/${folderId}/presigned-upload`,
    { method: 'POST', body: JSON.stringify({ name: file.name }) }
  ).then(r => r.json());

  // 2️⃣ Stream the file to S3
  await uploadToPresignedUrl(url, file);

  // 3️⃣ Return the S3 key so the UI can display the new document
  return key;
}

This approach keeps the Macro application servers out of the data path—handling large files without consuming application memory or bandwidth.

Step 2: Metadata Persistence in MacroDB

Once the file lands in the document‑storage‑service S3 bucket, the API records its metadata. The simple_save.rs module in the document storage service creates the database entry:

// services/document_storage_service/src/api/documents/simple_save.rs
// Inserts a row into the Document table with owner, type, and S3 key

This write to MacroDB establishes the document's identity and triggers the next processing stage.

Step 3: Queueing the Extraction Job

After persistence, a DocumentUploadJob is enqueued on the DocumentTextExtractorQueue—an SQS queue defined in crates/macro_queues/src/lib.rs:

// crates/macro_queues/src/lib.rs
pub async fn enqueue_document_text_extractor(
    &self,
    s3_key: &str,
    document_id: i64,
) -> anyhow::Result<()> {
    let msg = serde_json::json!({
        "s3_key": s3_key,
        "document_id": document_id,
    });
    self.sqs
        .send_message(&self.document_text_extractor_queue, msg.to_string())
        .await
}

The queue definition lives in the shared macro_queues crate, ensuring consistent queue semantics across services.

Step 4: Lambda Trigger and File Download

The document‑text‑extractor Lambda wakes when SQS delivers the message. Its infrastructure definition in infra/stacks/document-text-extractor/document-text-extractor.ts bundles the pdfium native library and configures environment variables for S3 access.

The Lambda handler downloads the raw file using the s3_key from the job message.

Step 5: Format‑Specific Text Extraction

Extraction logic varies by file type:

  • PDF → processed with pdfium via the bundled native library in services/document_text_extractor/src/main.rs
  • DOCX → unzipped by the docx_unzip_handler service, then parsed
  • Plain‑text → read directly without transformation

The extractor emits UTF‑8 text chunks suitable for database storage and search indexing. The core extraction flow appears in:

// services/document_text_extractor/src/main.rs
#[tokio::main]
async fn main() -> anyhow::Result<()> {
    let event = sqs::receive_message().await?;
    let job: DocumentUploadJob = serde_json::from_str(&event.body)?;
    let raw = s3::download(&job.s3_key).await?;
    let text = pdfium::extract_text(&raw)?;
    db::store_extracted_text(job.document_id, &text).await?;
    search::queue_indexing(job.document_id).await?;
    Ok(())
}

DOCX handling splits across services: services/docx_unzip_handler/src/service/document/notify_docx_upload_job.rs manages the unzip notification flow.

Step 6: Storing Extracted Text

The processed text is persisted in two places:

  1. MacroDB as DocumentChunk records (via crates/documents/src/domain/upload_finalize.rs)
  2. S3 as a backup object for reconstruction or re‑indexing

The upload_finalize.rs module in the documents crate coordinates this dual write.

Step 7: OpenSearch Indexing

With text extracted, a SearchEvent enters the search‑upload queue. The search‑upload Lambda—defined in infra/stacks/search-upload/search-upload-lambda.ts—consumes this event and pushes the document into OpenSearch:

// infra/stacks/search-upload/search-upload-lambda.ts
import { OpenSearchClient } from '@opensearch-project/client';
import { getDocumentChunks } from './document-store';

export const handler = async (event: any) => {
  const docId = event.document_id;
  const chunks = await getDocumentChunks(docId);
  const client = new OpenSearchClient({ node: process.env.OPENSEARCH_ENDPOINT });
  await client.index({
    index: 'documents',
    id: `${docId}`,
    body: { content: chunks.join('\n') },
  });
};

Field‑level analyzers in the OpenSearch mapping enable full‑text search, highlighting, and faceting.

Step 8: Querying the Search Index

Once indexed, the document becomes searchable. The lexical‑service exposes an HTTP endpoint that forwards queries to OpenSearch and returns highlighted results:

// services/lexical-service/src/endpoints/search-text.ts
// The HTTP endpoint that forwards queries to OpenSearch

Frontend search requests hit this Axum‑powered API, which translates user queries into OpenSearch DSL and formats the response.

Key Architecture Decisions

Understanding why Macro structured its document processing this way reveals important engineering trade‑offs:

Decision Implementation Benefit
Event‑driven async SQS between every stage Natural back‑pressure, horizontal scaling, independent service deployment
Native PDF processing pdfium bundled in Lambda Fast, high‑fidelity text extraction without external API dependencies
Dual storage MacroDB + S3 for extracted text Durability, re‑indexing capability without re‑extraction
OpenSearch as search store Dedicated indexing Lambda Dedicated indexing path, optimized for query performance

Summary

Frequently Asked Questions

What queue system does Macro use for document processing?

Macro uses Amazon SQS for all inter‑service communication in the document processing pipeline. The DocumentTextExtractorQueue and search‑upload queue are defined in crates/macro_queues/src/lib.rs and consumed by Lambda functions. This event‑driven approach decouples services and provides automatic retries for failed extractions.

How does Macro handle different file formats?

Each format has a dedicated code path. PDFs are processed with the pdfium native library bundled into the document‑text‑extractor Lambda. DOCX files are routed through a separate docx_unzip_handler service that unzips the archive and extracts the document XML. Plain‑text files bypass transformation and are read directly. This modular approach allows adding new format handlers without modifying existing code.

Where does Macro store extracted document text?

Extracted text is stored in two locations: as DocumentChunk records in MacroDB for operational queries, and as objects in S3 for durability and potential re‑indexing. The upload_finalize.rs module in the documents crate coordinates these writes before triggering the search indexing step.

OpenSearch indexes all extracted document content. The search‑upload Lambda pushes text chunks into an OpenSearch index with configured analyzers, while the lexical‑service API in services/lexical-service/src/endpoints/search-text.ts forwards user queries and returns highlighted results.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →