# MacroDB, CommsDB, EmailDB, and ContactsDB Architecture Explained: 4-Database Design in the Macro Repository

> Explore the Macro Repository's 4-database design: MacroDB, CommsDB, EmailDB, and ContactsDB. Learn how this architecture isolates data in dedicated Rust crates for optimal performance.

- Repository: [Macro/macro](https://github.com/macro-inc/macro)
- Tags: architecture
- Published: 2026-08-16

---

**Macro's data layer splits across four PostgreSQL databases—MacroDB, CommsDB, EmailDB, and ContactsDB—each isolated in its own Rust crate with dedicated connection pools, migrations, and SQLx query definitions.**

This modular architecture lets services scale independently while maintaining type-safe database access through the `sqlx` ecosystem. The `macro-inc/macro` repository implements this pattern to separate core document data from communication, email, and contact management concerns.

## The Four Database Architecture

The Macro platform partitions data by domain into four distinct PostgreSQL instances. Each database ships as a standalone Rust crate under the `crates/` directory, exposing a public API for connection pooling and query execution.

### MacroDB: The Primary Data Store

**MacroDB** lives in `crates/macro_db_client` and serves as the central repository for documents, users, projects, and cross-cutting entities like messages and notification history.

- **Key responsibilities**: User accounts, project metadata, document content, channel records, notification preferences
- **Connection source**: `DATABASE_URL_MACRO_DB` environment variable loaded via `macro_env_var`
- **Migration path**: `crates/macro_db_client/migrations/`

Core services including `document-storage-service`, `search_service`, and `notification_service` import this crate and invoke its query helpers.

```rust
// Fetching a user from MacroDB
use macro_db_client::pool::get_pool;
use macro_db_client::queries::get_user_by_id;

#[tokio::main]
async fn main() -> Result<(), anyhow::Error> {
    let pool = get_pool().await?;                // pulls DATABASE_URL_MACRO_DB
    let user = get_user_by_id(&pool, 42).await?; // sqlx! query defined in crate
    println!("Found user: {}", user.email);
    Ok(())
}

```

The pool initialization and query macros are defined in [[`crates/macro_db_client/Cargo.toml`](https://github.com/macro-inc/macro/blob/main/crates/macro_db_client/Cargo.toml)](https://github.com/macro-inc/macro/blob/main/crates/macro_db_client/Cargo.toml), which declares `sqlx` and `tokio` dependencies alongside PostgreSQL-specific feature flags.

### CommsDB: Communication-Scoped Storage

**CommsDB** resides in `crates/comms_db_client` and isolates conversation-level data from the main MacroDB schema. This separation allows the communication stack to scale independently of core document operations.

- **Key responsibilities**: Conversations, message threads, participant mappings
- **Primary consumers**: `communication_service` and related background workers
- **Connection source**: `DATABASE_URL_COMMS_DB`

Services call into this crate for operations like thread creation, participant management, and message pagination without touching MacroDB tables.

### EmailDB: Email-Specific Entities

**EmailDB** is packaged in `crates/email_db_client` and manages email thread state, message bodies, and mailbox metadata.

- **Key responsibilities**: Email threads, individual messages, mailbox configuration
- **Primary consumers**: `email_service`, scheduled email handlers, inbound mail processors
- **Connection source**: `DATABASE_URL_EMAIL_DB`

```rust
// Inserting a new email thread into EmailDB
use email_db_client::pool::get_pool;
use email_db_client::queries::create_email_thread;

#[tokio::main]
async fn main() -> Result<(), anyhow::Error> {
    let pool = get_pool().await?; // uses DATABASE_URL_EMAIL_DB
    let thread_id = create_email_thread(&pool, "Welcome".into(), 1).await?;
    println!("Created email thread {}", thread_id);
    Ok(())
}

```

The crate structure mirrors MacroDB's pattern with its own migration directory and offline query validation via `.sqlx/` metadata. See [[`crates/email_db_client/Cargo.toml`](https://github.com/macro-inc/macro/blob/main/crates/email_db_client/Cargo.toml)](https://github.com/macro-inc/macro/blob/main/crates/email_db_client/Cargo.toml) for dependency specifications.

### ContactsDB: Relationship and Contact Management

**ContactsDB** is implemented in `crates/contacts` (with supporting `contacts_service`) and tracks user connections, contact lists, and relationship metadata.

- **Key responsibilities**: Contact records, contact groups, connection graphs
- **Primary consumers**: `contacts_service` for lookup resolution and group management
- **Schema scope**: Isolated from user authentication data in MacroDB

This crate follows the same architectural conventions as the other three DB clients, as declared in [[`crates/contacts/Cargo.toml`](https://github.com/macro-inc/macro/blob/main/crates/contacts/Cargo.toml)](https://github.com/macro-inc/macro/blob/main/crates/contacts/Cargo.toml).

## Architectural Patterns Across All Four Databases

### SQLx Offline Validation

Every DB client crate uses `sqlx::query!` and `sqlx::query_as!` macros with compile-time schema checking. The `.sqlx/` directory in each crate caches prepared metadata so builds succeed without a live database. Refresh the cache across all crates with:

```bash
just prepare_db

```

This command, defined in the repository's top-level `just` file, regenerates query metadata for MacroDB, CommsDB, EmailDB, and ContactsDB simultaneously.

### Environment-Driven Connection Configuration

The `macro_env_var` crate centralizes secret loading from Doppler. No service calls `std::env::var` directly; instead, each DB client requests its connection string through this abstraction. The environment variable naming convention follows `DATABASE_URL_{DB}_{SUFFIX}` patterns, as documented in [[`crates/macro_env_var/README.md`](https://github.com/macro-inc/macro/blob/main/crates/macro_env_var/README.md)](https://github.com/macro-inc/macro/blob/main/crates/macro_env_var/README.md).

### Migration Independence

Each database manages its own schema evolution:

| Database | Migration Command |
|----------|-----------------|
| MacroDB | `just setup_macrodb` |
| CommsDB | `just setup_commsdb` |
| EmailDB | `just setup_emaildb` |
| ContactsDB | `just setup_contactsdb` |

These commands invoke `sqlx migrate` against the appropriate target database, ensuring zero-downtime schema changes per service boundary.

### Cross-Database Coordination

Some workflows—like notifying a user after email receipt—span multiple databases. Services handle this by importing multiple client crates and coordinating operations at the application layer rather than through distributed transactions. Each operation respects its target database's isolation level, with idempotency keys and retry logic managed in the service code.

### Testing Infrastructure

The repository provides Docker Compose configurations that spin isolated PostgreSQL containers for each database. The `just test` command runs the full test suite against these live schemas. CI pipelines use `SQLX_OFFLINE=true` only for compile verification, never for test execution, ensuring query validity against actual table structures.

## Source File Reference

| Path | Purpose |
|------|---------|
| [[`crates/macro_db_client/Cargo.toml`](https://github.com/macro-inc/macro/blob/main/crates/macro_db_client/Cargo.toml)](https://github.com/macro-inc/macro/blob/main/crates/macro_db_client/Cargo.toml) | Primary database client dependencies and features |
| [[`crates/comms_db_client/Cargo.toml`](https://github.com/macro-inc/macro/blob/main/crates/comms_db_client/Cargo.toml)](https://github.com/macro-inc/macro/blob/main/crates/comms_db_client/Cargo.toml) | Communication database client configuration |
| [[`crates/email_db_client/Cargo.toml`](https://github.com/macro-inc/macro/blob/main/crates/email_db_client/Cargo.toml)](https://github.com/macro-inc/macro/blob/main/crates/email_db_client/Cargo.toml) | Email database client configuration |
| [[`crates/contacts/Cargo.toml`](https://github.com/macro-inc/macro/blob/main/crates/contacts/Cargo.toml)](https://github.com/macro-inc/macro/blob/main/crates/contacts/Cargo.toml) | Contacts database layer |
| [`just`](https://github.com/macro-inc/macro/blob/main/just) | Task runner with database setup and migration commands |

## Summary

- **Four PostgreSQL databases**—MacroDB, CommsDB, EmailDB, ContactsDB—partition data by domain in the `macro-inc/macro` repository
- **Crate-per-database isolation** keeps schemas, migrations, and connection pools independent under `crates/`
- **SQLx macros** provide compile-time query validation with offline mode support via `.sqlx/` metadata
- **`macro_env_var` abstraction** centralizes secret loading from Doppler for all connection strings
- **Independent migrations** let each database evolve on its own release schedule
- **Cross-database workflows** import multiple clients and coordinate at the service layer

## Frequently Asked Questions

### Why four separate databases instead of one large PostgreSQL instance?

Domain isolation lets teams scale, migrate, and deploy services independently. The communication stack (CommsDB) can handle high write throughput without impacting document search queries (MacroDB). Email processing (EmailDB) runs on different backup schedules than contact relationship data (ContactsDB). This separation also limits blast radius—schema corruption in one domain doesn't cascade to others.

### How does the Macro repository handle database connections securely?

All DB clients use the `macro_env_var` crate to load `DATABASE_URL_*` values from Doppler. This prevents stray `std::env::var` calls and enables secret rotation without code changes. Connection pools are instantiated once per service startup and shared across request handlers.

### What happens when a query needs data from multiple databases?

Services import the relevant client crates—e.g., `macro_db_client` and `email_db_client`—and execute sequential or parallel queries. The architecture avoids distributed transactions; instead, services use idempotency keys, outbox patterns, or saga orchestration for cross-database consistency. Each query runs against its native connection pool with standard PostgreSQL isolation semantics.