# How the Document Cognition Service Analyzes and Processes Uploaded Documents in Macro

> Discover how the Macro Document Cognition Service analyzes and processes uploaded documents through ingestion, text extraction, and AI-driven analysis for searchable, enhanced content.

- Repository: [Macro/macro](https://github.com/macro-inc/macro)
- Tags: deep-dive
- Published: 2026-08-16

---

**The Document Cognition Service (DCS) transforms raw uploads into searchable, AI-enhanced content through a three-stage pipeline: ingestion, text extraction, and AI-driven analysis.**

The **Document Cognition Service** is the core engine in the [macro-inc/macro](https://github.com/macro-inc/macro) repository that converts user-uploaded files into enriched, queryable documents. Built in Rust and deployed on AWS ECS Fargate, DCS coordinates storage, extraction, and AI processing to power Macro's intelligent document workflows.

## Document Ingestion and Storage

When a user initiates an upload, the **Document Storage Service (DSS)** generates a pre-signed S3 URL to securely receive the file. This happens through the `S3UploadUrlAdapter` struct, which targets either `config.document_storage_bucket` for standard files or `config.docx_document_upload_bucket` for Word documents.

In [`services/document_cognition_service/src/main.rs`](https://github.com/macro-inc/macro/blob/main/services/document_cognition_service/src/main.rs) at lines 66-71, the adapter is instantiated:

```rust
let s3_upload_adapter = S3UploadUrlAdapter::new(
    s3_client,
    config.document_storage_bucket.to_string(),
    config.docx_document_upload_bucket.to_string(),
);

```

The client uploads directly to S3, minimizing load on the service and enabling resumable transfers.

## Text Extraction Pipeline

Once the object lands in S3, DSS emits a message to the **Document Text Extractor** SQS queue. DCS establishes a handle to this queue in [`main.rs`](https://github.com/macro-inc/macro/blob/main/main.rs) at lines 102-110:

```rust
let document_text_extractor_queue = macro_queues::DocumentTextExtractorQueue::new();
let sqs_client = sqs_client::SQS::new(queue_aws_client)
    .document_text_extractor_queue(&document_text_extractor_queue);

```

A dedicated Lambda function (`document_text_extractor`) performs the actual extraction:

- **PDF files** — processed with PDFium for robust text extraction
- **DOCX files** — unzipped and parsed from the underlying XML

The Lambda writes the extracted plain text back to S3, triggering the next stage.

## AI-Driven Document Analysis

After text extraction completes, DCS receives a notification and invokes the **Document Tool Context** (`DocumentToolContext`). This context composes three critical dependencies defined in [`main.rs`](https://github.com/macro-inc/macro/blob/main/main.rs) at lines 108-124:

1. **`DocumentServiceImpl`** — orchestrates storage operations, synchronization, and property updates
2. **`AiTools`** — provides access to summarization, embeddings, and projections via `ai_tools::all_tools()`
3. **`task_properties_service`** — propagates AI-derived properties to associated tasks

The service executes multiple AI pipelines in parallel:

### Summarization and Classification

The `ai_projections` subsystem generates concise document previews and assigns intelligent tags for categorization.

### Embedding Generation

Vector embeddings are produced and stored in the `ai_projection` table. These embeddings feed the **Search Service**, enabling semantic retrieval beyond keyword matching.

### Property Propagation

Task-related attributes—status, priority, due date, and assignee—are automatically inferred and synchronized through `task_properties_service` at lines 90-99.

## Persistence and Real-Time Notification

All derived metadata flows through the repository layer. The `AiProjectionRepositoryImpl` (wrapped by `AiProjectionServiceImpl`) persists previews, embeddings, and AI-generated properties to PostgreSQL, as referenced in [`main.rs`](https://github.com/macro-inc/macro/blob/main/main.rs) at lines 64-71.

DCS then publishes a notification to the **AI Projection Queue** (`ai_projection_queue`), allowing downstream clients like the web UI to refresh document state in real time. The [`ai_projections/worker.rs`](https://github.com/macro-inc/macro/blob/main/ai_projections/worker.rs) file handles queue consumption and materialization.

## HTTP API Exposure

Processed documents are accessible through REST endpoints defined in `src/api/`:

| Endpoint | Source File | Purpose |
|----------|-------------|---------|
| `GET /documents/{id}/preview` | [`api/preview/get_batch_preview.rs`](https://github.com/macro-inc/macro/blob/main/api/preview/get_batch_preview.rs) | Returns AI-generated previews in batch |
| `GET /documents/{id}/embedding` | `api/attachments/*` | Retrieves vector embeddings for search |

These handlers are wired into the Axum server through `api::setup_and_serve` in the main entry point.

## Summary

- **Ingestion**: `S3UploadUrlAdapter` generates pre-signed URLs for direct S3 uploads to configurable buckets
- **Extraction**: The `document_text_extractor_queue` delegates PDF/DOCX processing to a dedicated Lambda
- **AI Analysis**: `DocumentToolContext` orchestrates `DocumentServiceImpl` and `AiTools` for summarization, embeddings, and property propagation
- **Persistence**: `AiProjectionRepositoryImpl` stores results in PostgreSQL with real-time queue notifications
- **API**: Axum-powered endpoints in `src/api/` expose processed content to clients

## Frequently Asked Questions

### How does Macro handle different file formats during text extraction?

The `document_text_extractor` Lambda uses **PDFium** for PDF parsing and a **DOCX unzip routine** for Word documents. Both paths output standardized plain text to S3, where DCS picks up the result for AI processing.

### Where are document embeddings stored and how are they used?

Embeddings are stored in the `ai_projection` table via `AiProjectionServiceImpl`. The **Search Service** queries these vectors for semantic document retrieval, enabling natural language search beyond exact keyword matches.

### What triggers the AI analysis stage after upload?

DCS listens to the `document_text_extractor_queue` SQS queue. When the Lambda completes extraction and writes text to S3, a message notifies DCS to invoke `DocumentToolContext` and begin the AI pipeline.

### Can downstream services react to completed document processing?

Yes. DCS pushes notifications to the `ai_projection_queue` after persisting AI results. The [`ai_projections/worker.rs`](https://github.com/macro-inc/macro/blob/main/ai_projections/worker.rs) component and other subscribers consume these messages to update UIs or trigger additional workflows.