How the Document Cognition Service Analyzes and Processes Uploaded Documents in Macro
The Document Cognition Service (DCS) transforms raw uploads into searchable, AI-enhanced content through a three-stage pipeline: ingestion, text extraction, and AI-driven analysis.
The Document Cognition Service is the core engine in the macro-inc/macro repository that converts user-uploaded files into enriched, queryable documents. Built in Rust and deployed on AWS ECS Fargate, DCS coordinates storage, extraction, and AI processing to power Macro's intelligent document workflows.
Document Ingestion and Storage
When a user initiates an upload, the Document Storage Service (DSS) generates a pre-signed S3 URL to securely receive the file. This happens through the S3UploadUrlAdapter struct, which targets either config.document_storage_bucket for standard files or config.docx_document_upload_bucket for Word documents.
In services/document_cognition_service/src/main.rs at lines 66-71, the adapter is instantiated:
let s3_upload_adapter = S3UploadUrlAdapter::new(
s3_client,
config.document_storage_bucket.to_string(),
config.docx_document_upload_bucket.to_string(),
);
The client uploads directly to S3, minimizing load on the service and enabling resumable transfers.
Text Extraction Pipeline
Once the object lands in S3, DSS emits a message to the Document Text Extractor SQS queue. DCS establishes a handle to this queue in main.rs at lines 102-110:
let document_text_extractor_queue = macro_queues::DocumentTextExtractorQueue::new();
let sqs_client = sqs_client::SQS::new(queue_aws_client)
.document_text_extractor_queue(&document_text_extractor_queue);
A dedicated Lambda function (document_text_extractor) performs the actual extraction:
- PDF files — processed with PDFium for robust text extraction
- DOCX files — unzipped and parsed from the underlying XML
The Lambda writes the extracted plain text back to S3, triggering the next stage.
AI-Driven Document Analysis
After text extraction completes, DCS receives a notification and invokes the Document Tool Context (DocumentToolContext). This context composes three critical dependencies defined in main.rs at lines 108-124:
DocumentServiceImpl— orchestrates storage operations, synchronization, and property updatesAiTools— provides access to summarization, embeddings, and projections viaai_tools::all_tools()task_properties_service— propagates AI-derived properties to associated tasks
The service executes multiple AI pipelines in parallel:
Summarization and Classification
The ai_projections subsystem generates concise document previews and assigns intelligent tags for categorization.
Embedding Generation
Vector embeddings are produced and stored in the ai_projection table. These embeddings feed the Search Service, enabling semantic retrieval beyond keyword matching.
Property Propagation
Task-related attributes—status, priority, due date, and assignee—are automatically inferred and synchronized through task_properties_service at lines 90-99.
Persistence and Real-Time Notification
All derived metadata flows through the repository layer. The AiProjectionRepositoryImpl (wrapped by AiProjectionServiceImpl) persists previews, embeddings, and AI-generated properties to PostgreSQL, as referenced in main.rs at lines 64-71.
DCS then publishes a notification to the AI Projection Queue (ai_projection_queue), allowing downstream clients like the web UI to refresh document state in real time. The ai_projections/worker.rs file handles queue consumption and materialization.
HTTP API Exposure
Processed documents are accessible through REST endpoints defined in src/api/:
| Endpoint | Source File | Purpose |
|---|---|---|
GET /documents/{id}/preview |
api/preview/get_batch_preview.rs |
Returns AI-generated previews in batch |
GET /documents/{id}/embedding |
api/attachments/* |
Retrieves vector embeddings for search |
These handlers are wired into the Axum server through api::setup_and_serve in the main entry point.
Summary
- Ingestion:
S3UploadUrlAdaptergenerates pre-signed URLs for direct S3 uploads to configurable buckets - Extraction: The
document_text_extractor_queuedelegates PDF/DOCX processing to a dedicated Lambda - AI Analysis:
DocumentToolContextorchestratesDocumentServiceImplandAiToolsfor summarization, embeddings, and property propagation - Persistence:
AiProjectionRepositoryImplstores results in PostgreSQL with real-time queue notifications - API: Axum-powered endpoints in
src/api/expose processed content to clients
Frequently Asked Questions
How does Macro handle different file formats during text extraction?
The document_text_extractor Lambda uses PDFium for PDF parsing and a DOCX unzip routine for Word documents. Both paths output standardized plain text to S3, where DCS picks up the result for AI processing.
Where are document embeddings stored and how are they used?
Embeddings are stored in the ai_projection table via AiProjectionServiceImpl. The Search Service queries these vectors for semantic document retrieval, enabling natural language search beyond exact keyword matches.
What triggers the AI analysis stage after upload?
DCS listens to the document_text_extractor_queue SQS queue. When the Lambda completes extraction and writes text to S3, a message notifies DCS to invoke DocumentToolContext and begin the AI pipeline.
Can downstream services react to completed document processing?
Yes. DCS pushes notifications to the ai_projection_queue after persisting AI results. The ai_projections/worker.rs component and other subscribers consume these messages to update UIs or trigger additional workflows.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →