# Architecture of the Document Storage Service and Its S3 Integration Patterns: A Deep Dive into Macro's Implementation

> Explore Macro's Document Storage Service architecture. Learn about S3 integration patterns, CloudFront signed URLs, and direct client uploads in this deep dive.

- Repository: [Macro/macro](https://github.com/macro-inc/macro)
- Tags: architecture
- Published: 2026-08-17

---

**The Document Storage Service in the Macro repository is a stateless Rust microservice that abstracts Amazon S3 operations behind a typed `S3Client`, using CloudFront-signed URLs for secure downloads and presigned PUT URLs for direct client uploads, with fallback server-side streaming for internal migrations.**

The `document_storage_service` binary powers document persistence in the [macro-inc/macro](https://github.com/macro-inc/macro) codebase. Understanding the architecture of the document storage service and its S3 integration patterns reveals how the platform achieves horizontal scalability while offloading binary traffic directly to S3 to minimize latency.

## Core Architecture Components

The service is organized around distinct components that separate configuration, storage abstraction, and HTTP handling.

- **`Config`** – Defined in [`services/document_storage_service/src/config.rs`](https://github.com/macro-inc/macro/blob/main/services/document_storage_service/src/config.rs), this struct loads environment variables including the three S3 bucket names (`document_storage_bucket`, `docx_upload_bucket`, `upload_staging_bucket`), CloudFront distribution settings, and JWT secrets. It also resolves sensitive values from AWS Secrets Manager at startup.

- **`S3Client` (`service::s3::S3`)** – Located in [`services/document_storage_service/src/service/s3/mod.rs`](https://github.com/macro-inc/macro/blob/main/services/document_storage_service/src/service/s3/mod.rs), this thin wrapper around the AWS SDK S3 client exposes high-level helpers for `get`, `upload`, presigned-URL creation, `copy`, `delete`, existence checks, and folder listings. It holds references to the three operational buckets and manages all S3 I/O.

- **`S3UploadUrlAdapter`** – Implemented in [`services/document_storage_service/src/outbound/mod.rs`](https://github.com/macro-inc/macro/blob/main/services/document_storage_service/src/outbound/mod.rs), this component exposes an Axum-compatible endpoint that returns presigned PUT URLs. Clients use these URLs to upload binaries directly to S3 without streaming through the service binary.

- **`DocumentServiceImpl`** – The domain orchestrator defined in [`services/document_storage_service/src/main.rs`](https://github.com/macro-inc/macro/blob/main/services/document_storage_service/src/main.rs) (around line 99) receives the `S3Client` and `S3UploadUrlAdapter` via dependency injection. It coordinates metadata database writes, CloudFront URL building, S3 interactions, and event publishing.

- **API Layer** – Axum handlers in `services/document_storage_service/src/api/documents/` (e.g., [`get.rs`](https://github.com/macro-inc/macro/blob/main/get.rs)) expose REST endpoints under `/documents/*`, `/pins/*`, and `/recents/*`. Each handler extracts the `ApiContext`, validates JWTs, and delegates to `DocumentServiceImpl`.

## S3 Integration Patterns

The service implements five distinct patterns for S3 interaction, balancing security, performance, and operational flexibility.

### Direct GET via CloudFront-Signed URLs

For document retrieval, the service generates CloudFront-signed URLs rather than proxying bytes. The `CloudFrontConfig` struct—initialized in [`services/document_storage_service/src/main.rs`](https://github.com/macro-inc/macro/blob/main/services/document_storage_service/src/main.rs) at line 36—holds the distribution URL, signer key ID, private key, and expiry settings. The `S3Client` methods `put_snapshot_presigned_url` and `get_snapshot_presigned_url` embed CloudFront signatures, allowing clients to fetch files via HTTPS without further authentication while benefiting from CloudFront's edge caching.

### Presigned PUT for Client-Side Uploads

The primary upload path minimizes service load by directing clients straight to S3. When a client calls `/documents/upload-url`, the `S3UploadUrlAdapter` returns a presigned PUT URL scoped to the appropriate bucket (`document_storage_bucket`, `docx_upload_bucket`, or `upload_staging_bucket`). The client then executes a `PUT` request directly against S3, bypassing the service binary entirely and reducing upload latency.

### Server-Side Upload Fallback

When internal services or migrations require server-side handling, the code path uses `S3Client::upload_document`. This method accepts a `ByteStream` and streams the payload into S3. This pattern is invoked for thumbnail generation, internal API ingestion, or scenarios where presigned URLs are impractical. The implementation resides in [`services/document_storage_service/src/service/s3/mod.rs`](https://github.com/macro-inc/macro/blob/main/services/document_storage_service/src/service/s3/mod.rs).

### Folder Enumeration and Bulk Existence Checks

The `S3Client` provides utility methods for bulk operations. `get_folder_content_names` lists objects under a logical prefix, returning tuples of `(file_name, file_type)` for each entry. For deduplication workflows, `shas_exist` performs concurrent `HEAD` requests for a batch of SHA hashes. To prevent S3 throttling, this method uses a semaphore limited to **10 concurrent permits**, ensuring the service remains within AWS rate limits during large-scale existence checks.

### Copy and Delete Operations

Document versioning and cleanup rely on `S3Client::copy_document` and `S3Client::delete_document`. The copy helper moves objects between logical folders or buckets, while the delete helper removes all objects associated with a specific user/document combination. Background SQS workers—such as the `delete_document_worker`—reuse the same `S3Client` to perform these operations asynchronously.

## Dependency Injection and Concurrency Management

The service composes its components using Rust's ownership model and `Arc` for safe sharing across async contexts. In [`services/document_storage_service/src/main.rs`](https://github.com/macro-inc/macro/blob/main/services/document_storage_service/src/main.rs), the bootstrap sequence wraps the S3 client in an `Arc` before injecting it into `DocumentServiceImpl`:

```rust
let s3_client = macro_aws_config::s3_client().await;
let s3 = Arc::new(S3::new(
    s3_client,
    config.document_storage_bucket.as_ref(),
    config.docx_document_upload_bucket.as_ref(),
    config.upload_staging_bucket.as_ref(),
));
let document_service = Arc::new(DocumentServiceImpl::new(
    document_repo,
    cloudfront_config,
    sync_service_client.as_ref().clone(),
    s3_upload_adapter,
    /* other deps … */
));

```

This pattern ensures the `S3Client` can be shared safely across concurrent request handlers and background workers without cloning the underlying AWS SDK client.

## Scalability and Background Processing

The Document Storage Service is designed for horizontal scaling. Because all state persists in PostgreSQL, Redis, DynamoDB, or S3, the service itself is stateless. Multiple instances can run behind a load balancer without session affinity. Background SQS workers consume deletion and processing queues, using the same injected `S3Client` to perform batch operations without blocking the HTTP API.

## Practical Code Examples

### Generating a Presigned Upload URL

```rust
// Assume `s3` is the injected Arc<S3>
let key = format!("users/{}/documents/{}", user_id, filename);
let sha = "abc123..."; // SHA of the content, used for deduplication
let content_type = ContentType::Pdf; // or whatever matches the file
let url = s3.put_document_storage_presigned_url(&key, sha, content_type).await?;

```

### Streaming a Server-Side Upload

```rust
let key = format!("users/{}/documents/{}", user_id, filename);
let content: Vec<u8> = ...; // binary payload obtained from the request
s3.upload_document(&key, content).await?;

```

### Checking SHA Existence Before Upload

```rust
let shas = vec!["sha1".to_string(), "sha2".to_string()];
let all_exist = s3.shas_exist(&shas).await?;
if !all_exist {
    // Proceed with upload for missing SHAs
}

```

### Listing Directory Contents

```rust
let folder = format!("users/{}/documents/", user_id);
let entries = s3.get_folder_content_names(&folder).await?;
for (name, typ) in entries {
    println!("{} ({})", name, typ);
}

```

## Summary

- The **Document Storage Service** is a stateless Rust binary running in the Macro workspace, with source configuration in [`services/document_storage_service/src/config.rs`](https://github.com/macro-inc/macro/blob/main/services/document_storage_service/src/config.rs) and S3 abstraction in [`services/document_storage_service/src/service/s3/mod.rs`](https://github.com/macro-inc/macro/blob/main/services/document_storage_service/src/service/s3/mod.rs).
- It uses **three distinct S3 buckets** (`document_storage_bucket`, `docx_upload_bucket`, `upload_staging_bucket`) to segregate production data, Word document imports, and staging uploads.
- **CloudFront-signed URLs** provide secure, cacheable download paths without proxying bytes through the service.
- **Presigned PUT URLs** enable direct client-to-S3 uploads, offloading bandwidth from the service binary.
- **`S3Client::shas_exist`** implements concurrency limiting (10 permits) to prevent AWS throttling during bulk deduplication checks.
- **Dependency injection** via `Arc` allows safe sharing of the S3 client across async handlers and background SQS workers.

## Frequently Asked Questions

### How does Macro handle direct client uploads to S3 securely?

The service exposes an endpoint via `S3UploadUrlAdapter` in [`src/outbound/mod.rs`](https://github.com/macro-inc/macro/blob/main/src/outbound/mod.rs) that generates presigned PUT URLs scoped to specific buckets and object keys. These URLs include time-limited AWS signatures and optional CloudFront signatures, allowing the client to upload directly to S3 while the service retains control over the destination path and access policy.

### What is the purpose of the three separate S3 buckets in the Document Storage Service?

According to the `Config` struct in [`src/config.rs`](https://github.com/macro-inc/macro/blob/main/src/config.rs), the `document_storage_bucket` holds production user documents, `docx_upload_bucket` isolates Word document imports for processing, and `upload_staging_bucket` provides a temporary landing zone for asynchronous ingestion workflows. This separation enforces data isolation and allows different lifecycle policies per document type.

### How does the service prevent S3 throttling during bulk operations?

The `shas_exist` method in [`src/service/s3/mod.rs`](https://github.com/macro-inc/macro/blob/main/src/service/s3/mod.rs) uses a semaphore with **10 concurrent permits** to limit the number of simultaneous `HEAD` requests issued against S3. This rate limiting protects the service from AWS throttling errors when checking the existence of hundreds or thousands of SHA hashes during deduplication scans.

### What role does CloudFront play in the document retrieval architecture?

CloudFront serves as both a content delivery network and a security layer. The `CloudFrontConfig` initialized in [`src/main.rs`](https://github.com/macro-inc/macro/blob/main/src/main.rs) provides signing credentials for the `get_snapshot_presigned_url` method, which generates signed URLs that expire after a configurable duration. This allows offline clients and browsers to download documents directly from edge locations without exposing S3 bucket endpoints or requiring request-time authentication.