Macro Architecture Overview: Rust-Based Cloud Storage Microservices Explained

The Macro architecture is an 80+ crate Rust Cargo workspace implementing a cloud storage platform through domain-driven microservices, using PostgreSQL for transactional data, S3 for object storage, and AWS Lambda for serverless document processing.

The macro-inc/macro repository implements a comprehensive document lifecycle management system as a cloud-native Rust application. According to the project documentation in AGENTS.md, the architecture handles everything from document upload and AI-driven text extraction to search indexing and user notifications. The system is designed as a loosely coupled collection of specialized services communicating through HTTP APIs and message queues.

Workspace Structure and Service Groups

The codebase is organized as a Cargo workspace containing over 80 individual crates, each representing a distinct service domain or shared library. The services are grouped by functional responsibility rather than technical layer, following domain-driven design principles.

Core Storage Services

The central document management capabilities reside in four primary services:

  • document-storage-service – Handles upload, retrieval, and lifecycle management of documents via REST APIs
  • document-cognition-service – Provides AI-driven analysis and content understanding
  • search_service – Exposes global search APIs for document discovery
  • static_file_service – Serves static assets and processed file outputs

As implemented in crates/document-storage-service/src/main.rs, these services act as the primary interface for document operations, coordinating with downstream processing pipelines.

Processing and Extraction Services

Background document processing is handled by specialized workers:

  • convert_service – Transforms documents between formats
  • document-text-extractor – Extracts searchable text from PDFs and DOCX files, implemented in crates/document-text-extractor/src/lib.rs
  • search_processing_service – Manages indexing pipelines for OpenSearch integration

Communication and Infrastructure

Supporting services handle cross-cutting concerns:

  • email_service – Parses incoming emails and manages thread-based communication, defined in crates/email_service/src/lib.rs
  • notification_service – Dispatches user alerts and system events
  • authentication_service – Validates tokens and manages identity
  • connection_gateway – Maintains WebSocket connections for real-time features

Data Storage Architecture

The platform employs a polyglot persistence strategy, matching specific storage technologies to access patterns and data models.

Primary Transactional Stores

MacroDB is the central PostgreSQL instance responsible for documents, users, projects, communication data, email threads, and notification history. ContactsDB is a separate PostgreSQL instance dedicated exclusively to user connections and contact management. This separation allows independent scaling of social graph operations from core document workflows.

External Storage Systems

  • S3 – Persistent object storage for raw file blobs and document assets
  • Redis – Low-latency caching layer and session state management
  • OpenSearch – Full-text search indices for extracted document content, managed by search_service
  • DynamoDB – High-throughput connection tracking tables

Inter-Service Communication Patterns

Services communicate through multiple channels depending on latency and reliability requirements:

  1. HTTP APIs – Synchronous REST calls between internal service clients for immediate data retrieval
  2. SQS queues – Asynchronous message passing for background jobs and decoupled processing pipelines
  3. Lambda triggers – Event-driven serverless functions handling format-specific tasks like DOCX unzipping or PDF processing
  4. Redis – Pub/sub and caching for real-time session management

This multi-modal approach allows the document-storage-service to immediately acknowledge uploads while delegating heavy processing to Lambda functions via SQS.

Database Implementation with SQLx

Each service owns its database client crate (e.g., macro_db_client, comms_db_client) located in crates/macro_db_client/src/lib.rs. The architecture uses SQLx with compile-time query verification to ensure type safety across service boundaries.

Database migrations are managed per-service, and the workspace provides a just prepare_db command to refresh the .sqlx cache after schema changes. This prevents runtime SQL errors by validating queries against the actual PostgreSQL schema during compilation.

AWS Integration Strategy

The architecture is deeply integrated with AWS cloud primitives:

  • S3 serves as the durable document store, with services generating pre-signed URLs for client uploads
  • Lambda executes serverless handlers for compute-intensive format processing, triggered by S3 events or SQS messages
  • SQS provides reliable message queuing between the synchronous API layer and asynchronous processors
  • DynamoDB tracks connection states with millisecond latency requirements
  • OpenSearch powers the full-text search capabilities exposed by search_service

Development Workflow and Tooling

The repository uses just (a command runner) for workspace-wide task orchestration. Key commands include:


# Compile all 80+ crates including services and Lambda functions

just build

# Execute tests for a specific service crate

cargo test -p document-storage-service

# Initialize PostgreSQL with all migrations

just setup_macrodb

# Refresh SQLx metadata after query changes

just prepare_db

Environment configuration is injected via Doppler and accessed through the shared macro_env_var crate, ensuring consistent secrets management across local development and production deployments.

Example Service Endpoint

A typical HTTP handler in the document storage service follows this pattern from src/document_storage_service/src/api.rs:

#[tracing::instrument(err)]
pub async fn get_document(
    State(state): State<AppState>,
    Path(doc_id): Path<Uuid>,
) -> Result<Json<Document>, AppError> {
    let doc = state.db_client.get_document_by_id(doc_id).await?;
    Ok(Json(doc))
}

This leverages Axum's State extractor for dependency injection and tracing::instrument for structured observability.

Summary

  • Macro is an 80-crate Rust workspace implementing cloud storage microservices with domain-driven service boundaries
  • Data persistence uses PostgreSQL (MacroDB and ContactsDB), S3, Redis, OpenSearch, and DynamoDB for different access patterns
  • Service communication combines synchronous HTTP APIs with asynchronous SQS queues and AWS Lambda triggers
  • Type-safe database access is enforced through SQLx with per-service client crates and compile-time query checking
  • Development workflow relies on just commands, Doppler for secrets, and workspace-level database management tools

Frequently Asked Questions

What programming language and framework does Macro use?

The Macro platform is built entirely in Rust using the Cargo workspace structure. Services expose HTTP APIs using the Axum web framework, while database access is handled through SQLx for compile-time checked SQL queries.

How does Macro handle different file formats like PDF and DOCX?

The architecture delegates format-specific processing to AWS Lambda functions triggered via SQS queues. The document-text-extractor and convert_service handle extraction and transformation, allowing the core API to remain responsive while heavy processing occurs serverlessly.

Why does Macro use multiple databases instead of a single PostgreSQL instance?

The architecture separates MacroDB (documents, users, projects) from ContactsDB (social connections) to allow independent scaling and operational isolation. Additionally, specialized stores like DynamoDB (connections), OpenSearch (search), and S3 (blobs) optimize for specific query patterns that relational databases handle poorly.

How do developers manage database schema changes in the Macro workspace?

Each service owns its migrations in dedicated database client crates. Developers run just setup_macrodb to apply migrations and just prepare_db to regenerate SQLx query metadata after changing .sql files. This ensures compile-time verification catches schema mismatches before runtime.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →