# Macro Architecture Overview: Rust-Based Cloud Storage Microservices Explained

> Explore the Macro architecture, a Rust-based cloud storage microservices platform. Learn how it uses PostgreSQL, S3, and AWS Lambda for efficient data and document processing.

- Repository: [Macro/macro](https://github.com/macro-inc/macro)
- Tags: architecture
- Published: 2026-08-21

---

**The Macro architecture is an 80+ crate Rust Cargo workspace implementing a cloud storage platform through domain-driven microservices, using PostgreSQL for transactional data, S3 for object storage, and AWS Lambda for serverless document processing.**

The `macro-inc/macro` repository implements a comprehensive document lifecycle management system as a cloud-native Rust application. According to the project documentation in [`AGENTS.md`](https://github.com/macro-inc/macro/blob/main/AGENTS.md), the architecture handles everything from document upload and AI-driven text extraction to search indexing and user notifications. The system is designed as a loosely coupled collection of specialized services communicating through HTTP APIs and message queues.

## Workspace Structure and Service Groups

The codebase is organized as a Cargo workspace containing over 80 individual crates, each representing a distinct service domain or shared library. The services are grouped by functional responsibility rather than technical layer, following domain-driven design principles.

### Core Storage Services

The central document management capabilities reside in four primary services:

- **`document-storage-service`** – Handles upload, retrieval, and lifecycle management of documents via REST APIs
- **`document-cognition-service`** – Provides AI-driven analysis and content understanding
- **`search_service`** – Exposes global search APIs for document discovery
- **`static_file_service`** – Serves static assets and processed file outputs

As implemented in [`crates/document-storage-service/src/main.rs`](https://github.com/macro-inc/macro/blob/main/crates/document-storage-service/src/main.rs), these services act as the primary interface for document operations, coordinating with downstream processing pipelines.

### Processing and Extraction Services

Background document processing is handled by specialized workers:

- **`convert_service`** – Transforms documents between formats
- **`document-text-extractor`** – Extracts searchable text from PDFs and DOCX files, implemented in [`crates/document-text-extractor/src/lib.rs`](https://github.com/macro-inc/macro/blob/main/crates/document-text-extractor/src/lib.rs)
- **`search_processing_service`** – Manages indexing pipelines for OpenSearch integration

### Communication and Infrastructure

Supporting services handle cross-cutting concerns:

- **`email_service`** – Parses incoming emails and manages thread-based communication, defined in [`crates/email_service/src/lib.rs`](https://github.com/macro-inc/macro/blob/main/crates/email_service/src/lib.rs)
- **`notification_service`** – Dispatches user alerts and system events
- **`authentication_service`** – Validates tokens and manages identity
- **`connection_gateway`** – Maintains WebSocket connections for real-time features

## Data Storage Architecture

The platform employs a polyglot persistence strategy, matching specific storage technologies to access patterns and data models.

### Primary Transactional Stores

**MacroDB** is the central PostgreSQL instance responsible for documents, users, projects, communication data, email threads, and notification history. **ContactsDB** is a separate PostgreSQL instance dedicated exclusively to user connections and contact management. This separation allows independent scaling of social graph operations from core document workflows.

### External Storage Systems

- **S3** – Persistent object storage for raw file blobs and document assets
- **Redis** – Low-latency caching layer and session state management
- **OpenSearch** – Full-text search indices for extracted document content, managed by `search_service`
- **DynamoDB** – High-throughput connection tracking tables

## Inter-Service Communication Patterns

Services communicate through multiple channels depending on latency and reliability requirements:

1. **HTTP APIs** – Synchronous REST calls between internal service clients for immediate data retrieval
2. **SQS queues** – Asynchronous message passing for background jobs and decoupled processing pipelines
3. **Lambda triggers** – Event-driven serverless functions handling format-specific tasks like DOCX unzipping or PDF processing
4. **Redis** – Pub/sub and caching for real-time session management

This multi-modal approach allows the `document-storage-service` to immediately acknowledge uploads while delegating heavy processing to Lambda functions via SQS.

## Database Implementation with SQLx

Each service owns its database client crate (e.g., `macro_db_client`, `comms_db_client`) located in [`crates/macro_db_client/src/lib.rs`](https://github.com/macro-inc/macro/blob/main/crates/macro_db_client/src/lib.rs). The architecture uses **SQLx** with compile-time query verification to ensure type safety across service boundaries.

Database migrations are managed per-service, and the workspace provides a `just prepare_db` command to refresh the `.sqlx` cache after schema changes. This prevents runtime SQL errors by validating queries against the actual PostgreSQL schema during compilation.

## AWS Integration Strategy

The architecture is deeply integrated with AWS cloud primitives:

- **S3** serves as the durable document store, with services generating pre-signed URLs for client uploads
- **Lambda** executes serverless handlers for compute-intensive format processing, triggered by S3 events or SQS messages
- **SQS** provides reliable message queuing between the synchronous API layer and asynchronous processors
- **DynamoDB** tracks connection states with millisecond latency requirements
- **OpenSearch** powers the full-text search capabilities exposed by `search_service`

## Development Workflow and Tooling

The repository uses `just` (a command runner) for workspace-wide task orchestration. Key commands include:

```bash

# Compile all 80+ crates including services and Lambda functions

just build

# Execute tests for a specific service crate

cargo test -p document-storage-service

# Initialize PostgreSQL with all migrations

just setup_macrodb

# Refresh SQLx metadata after query changes

just prepare_db

```

Environment configuration is injected via **Doppler** and accessed through the shared `macro_env_var` crate, ensuring consistent secrets management across local development and production deployments.

### Example Service Endpoint

A typical HTTP handler in the document storage service follows this pattern from [`src/document_storage_service/src/api.rs`](https://github.com/macro-inc/macro/blob/main/src/document_storage_service/src/api.rs):

```rust
#[tracing::instrument(err)]
pub async fn get_document(
    State(state): State<AppState>,
    Path(doc_id): Path<Uuid>,
) -> Result<Json<Document>, AppError> {
    let doc = state.db_client.get_document_by_id(doc_id).await?;
    Ok(Json(doc))
}

```

This leverages Axum's `State` extractor for dependency injection and `tracing::instrument` for structured observability.

## Summary

- **Macro** is an 80-crate Rust workspace implementing cloud storage microservices with domain-driven service boundaries
- **Data persistence** uses PostgreSQL (MacroDB and ContactsDB), S3, Redis, OpenSearch, and DynamoDB for different access patterns
- **Service communication** combines synchronous HTTP APIs with asynchronous SQS queues and AWS Lambda triggers
- **Type-safe database access** is enforced through SQLx with per-service client crates and compile-time query checking
- **Development workflow** relies on `just` commands, Doppler for secrets, and workspace-level database management tools

## Frequently Asked Questions

### What programming language and framework does Macro use?

The Macro platform is built entirely in **Rust** using the **Cargo workspace** structure. Services expose HTTP APIs using the **Axum** web framework, while database access is handled through **SQLx** for compile-time checked SQL queries.

### How does Macro handle different file formats like PDF and DOCX?

The architecture delegates format-specific processing to **AWS Lambda functions** triggered via SQS queues. The `document-text-extractor` and `convert_service` handle extraction and transformation, allowing the core API to remain responsive while heavy processing occurs serverlessly.

### Why does Macro use multiple databases instead of a single PostgreSQL instance?

The architecture separates **MacroDB** (documents, users, projects) from **ContactsDB** (social connections) to allow independent scaling and operational isolation. Additionally, specialized stores like **DynamoDB** (connections), **OpenSearch** (search), and **S3** (blobs) optimize for specific query patterns that relational databases handle poorly.

### How do developers manage database schema changes in the Macro workspace?

Each service owns its migrations in dedicated database client crates. Developers run `just setup_macrodb` to apply migrations and `just prepare_db` to regenerate SQLx query metadata after changing `.sql` files. This ensures compile-time verification catches schema mismatches before runtime.