How Macro Inc. Handles Data Storage and Persistence: A Multi-Layer Architecture

Macro Inc. implements a polyglot persistence architecture that combines PostgreSQL for transactional business data, Amazon S3 for immutable document blobs, Redis for session caching, OpenSearch for full-text search, and DynamoDB for real-time connection state management.

The macro-inc/macro repository demonstrates a production-grade approach to data storage and persistence, utilizing specialized storage engines optimized for distinct data access patterns and consistency requirements. This architecture separates critical business records, large binary objects, and search indexes across multiple AWS-managed services and open-source technologies, all accessed through type-safe Rust clients.

Relational Database Architecture with PostgreSQL

Macro Inc. maintains several PostgreSQL databases to handle transactional business data with ACID guarantees. The platform operates multiple logical databases including MacroDB for core application data, CommsDB for project communications, EmailDB for message threading, and ContactsDB for relationship management.

Schema evolution is managed through versioned migration files stored in the database client crate:

Type-Safe Database Access with SQLx

All PostgreSQL interactions utilize the SQLx crate with compile-time query verification. Services obtain connections through a shared PgPool, executing transactions via sqlx::Transaction<'_, sqlx::Postgres>. The codebase leverages sqlx::query! and sqlx::query_scalar! macros, which are validated offline using the just prepare_db command to ensure type safety before deployment.

The following pattern from crates/crm/src/outbound/companies_repo.rs demonstrates the standard persistence workflow:

pub async fn insert_company(
    tx: &mut sqlx::Transaction<'_, sqlx::Postgres>,
    name: &str,
) -> Result<Company, sqlx::Error> {
    let row = sqlx::query!(
        r#"INSERT INTO crm_companies (name) VALUES ($1) RETURNING id, name, created_at"#,
        name
    )
    .fetch_one(&mut *tx)
    .await?;

    Ok(Company {
        id: row.id,
        name: row.name,
        created_at: row.created_at,
    })
}

This implementation acquires a pooled connection, executes a compiled insert statement within a transaction boundary, and returns strongly-typed results validated at compile time.

Immutable Document Storage Using Amazon S3

For binary large objects (BLOBs) such as PDFs, DOCX files, and images, Macro utilizes Amazon S3 via the aws-sdk-s3 crate. The crates/document_storage_service/src/outbound/s3_client.rs file implements the upload and download logic for document blobs, separating immutable file content from searchable metadata stored in PostgreSQL.

The S3 integration handles multipart uploads and presigned URL generation:

pub async fn upload_document(
    client: &aws_sdk_s3::Client,
    bucket: &str,
    key: &str,
    body: bytes::Bytes,
) -> Result<(), aws_sdk_s3::Error> {
    client
        .put_object()
        .bucket(bucket)
        .key(key)
        .body(body.into())
        .send()
        .await?;
    Ok(())
}

This approach ensures durable, highly available storage for user-generated content while keeping the relational databases optimized for query performance.

High-Performance Caching with Redis

To reduce database load and improve latency for frequently accessed data, Macro employs Redis as a distributed cache layer. The crates/redis_cache/src/lib.rs module provides a wrapper around Redis commands, managing session tokens, recent search results, and temporary processing state with configurable TTL (time-to-live) values.

The cache implementation supports atomic operations for session management:

pub async fn set_session(
    conn: &mut redis::aio::Connection,
    session_id: &str,
    data: &str,
    ttl_secs: usize,
) -> redis::RedisResult<()> {
    redis::cmd("SETEX")
        .arg(session_id)
        .arg(ttl_secs)
        .arg(data)
        .query_async(conn)
        .await?;
    Ok(())
}

Redis connections are pooled and injected into service handlers, providing sub-millisecond access to hot data.

Full-Text Search Indexing via OpenSearch

Document content and metadata require complex query capabilities beyond PostgreSQL's relational strengths. Macro integrates OpenSearch (Elasticsearch-compatible) through crates/search_service/src/outbound/opensearch_client.rs to power full-text search across extracted document content.

The OpenSearch client handles index mapping creation, document ingestion pipelines, and complex query DSL construction. This separation allows the platform to execute fuzzy searches, aggregations, and relevance scoring without impacting transactional database performance.

Real-Time Connection State with DynamoDB

For tracking websocket connections and serverless worker states, Macro utilizes Amazon DynamoDB as a key-value store. The crates/connection_gateway/src/outbound/dynamo_repository.rs module manages connection identifiers, heartbeat timestamps, and routing information using the aws-sdk-dynamodb crate.

DynamoDB's consistent low-latency reads and TTL support make it ideal for ephemeral connection data that requires high throughput but does not need complex relational constraints.

Reliable Event Delivery with Amazon SQS

Asynchronous processing pipelines leverage Amazon SQS for guaranteed message delivery. The crates/webhook/src/outbound/sqs_queue.rs implementation enqueues webhook payloads and document processing jobs, ensuring reliable event handling even during traffic spikes or service interruptions.

Summary

Macro Inc.'s data storage and persistence architecture follows these key principles:

  • PostgreSQL serves as the source of truth for transactional business data, with strict schema management via SQLx migrations and compile-time query verification
  • Amazon S3 provides durable, cost-effective storage for immutable document binaries, accessed through the AWS SDK for Rust
  • Redis caches session data and hot objects to minimize database round-trips and improve response times
  • OpenSearch indexes document content separately from storage, enabling sophisticated full-text search without impacting transactional performance
  • DynamoDB tracks ephemeral connection state for real-time features, leveraging its consistent single-digit millisecond latency
  • Amazon SQS guarantees reliable delivery of asynchronous events and background processing tasks

All storage configurations are externalized via the macro_env_var crate, ensuring credentials and connection strings remain outside the codebase while maintaining environment-specific flexibility.

Frequently Asked Questions

What database system does Macro Inc. use for primary business data?

Macro Inc. uses PostgreSQL as the primary relational database for all transactional data, including user accounts, projects, emails, and contacts. The platform maintains separate logical databases (MacroDB, CommsDB, EmailDB, ContactsDB) with schema defined in migration files located in crates/macro_db_client/migrations/. All database access occurs through the SQLx crate with compile-time query verification.

How does Macro store uploaded documents and files?

User-uploaded documents are stored in Amazon S3 rather than the relational database. The crates/document_storage_service/src/outbound/s3_client.rs file implements the S3 client logic using aws-sdk-s3, handling upload operations via the put_object API. This separation ensures PostgreSQL remains optimized for metadata queries while S3 handles the heavy lifting of binary storage and retrieval.

What caching technology does Macro use for session management?

Macro utilizes Redis for caching session tokens, recent search results, and temporary processing state. The crates/redis_cache/src/lib.rs module provides a Rust wrapper around Redis commands, implementing operations like SETEX for time-bound session storage. This cache layer reduces database load and provides sub-millisecond access to frequently requested data.

How does Macro implement full-text search across documents?

Full-text search is implemented using OpenSearch, accessed through the client defined in crates/search_service/src/outbound/opensearch_client.rs. This search engine indexes extracted document content and metadata separately from the primary storage, enabling complex queries, fuzzy matching, and relevance scoring without impacting the performance of the PostgreSQL transactional databases.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →