How to Design an S3-Like Object Storage System: 9 Critical Architecture Components

An S3-like object storage system achieves petabyte-scale durability and availability through a stateless API layer, strict separation of metadata and data paths, consistent hashing for placement, and configurable redundancy strategies ranging from 3-way replication to erasure coding.

Building an S3-compatible storage service requires architectural decisions that balance cost, latency, and durability guarantees. The liquidslr/system-design-notes repository outlines a complete design in 24. S3-like Object Storage/README.md that demonstrates how to implement 99.999999999% (11 nines) durability while exposing a simple RESTful API. This guide examines the nine fundamental design aspects that enable massive scalability without sacrificing consistency.

Core Architectural Components

Flat Namespace and Global Buckets

Objects in an S3-like system live within a flat namespace where buckets serve as globally unique containers rather than directory hierarchies. This design eliminates filesystem complexity and enables virtually unlimited scalability by avoiding recursive directory traversals. As defined in lines 37-40 of 24. S3-like Object Storage/README.md, buckets provide the top-level partitioning mechanism that allows the system to distribute objects across the entire storage cluster without maintaining complex tree structures.

Stateless API Layer for Horizontal Scaling

The API tier implements a stateless request handling pattern behind a load balancer, allowing horizontal scaling without session affinity. According to the high-level design diagram (lines 15-18), incoming REST requests route through stateless API services that authenticate users, validate permissions, and forward operations to either the metadata store or data store. This architecture ensures that adding API nodes increases throughput linearly without requiring sticky sessions or connection pooling.


# Upload a small object via the stateless API

curl -X PUT "https://api.example.com/bucket-name/hello.txt" \
     -H "Authorization: <sig>" \
     -H "Content-Type: text/plain" \
     -d "Hello, world!"

# Retrieve the object

curl -X GET "https://api.example.com/bucket-name/hello.txt" \
     -H "Authorization: <sig>"

Separation of Metadata and Data Stores

The architecture strictly separates metadata (object names, sizes, timestamps, versioning info) from raw data bytes, enabling independent scaling of read-heavy metadata queries and write-intensive storage operations. Lines 19-20 define the metadata store as a high-performance index, while lines 85-90 describe the data store as a distributed blob storage layer. This separation allows the system to optimize each tier specifically—metadata stores use fast SSDs for rapid lookups, while data stores utilize cost-effective HDDs or flash for bulk storage.

-- Query the metadata store for objects with a specific prefix
SELECT object_id, object_name, size, last_modified
FROM object
WHERE bucket_id = '12345' 
  AND object_name LIKE 'photos/2024/%';

Data Distribution and Durability Strategies

Consistent Hashing for Object Placement

Objects map to physical data nodes using consistent hashing of their UUIDs, ensuring even distribution across the cluster and minimal rebalancing during node additions or removals. Lines 32-34 explain how this deterministic placement algorithm routes objects to specific replication groups without requiring a central coordinator, reducing hotspots and enabling elastic capacity expansion.

Replication vs. Erasure Coding Trade-offs

The system supports two durability mechanisms: 3-way replication for low-latency reads and 8+4 erasure coding for cost-effective cold storage. As detailed in lines 25-31, 3-way replication provides immediate availability from any of three copies but incurs 200% storage overhead, while 8+4 erasure coding splits data into 8 fragments with 4 parity blocks, reducing overhead to 50% while maintaining >6 nines durability. The design allows per-bucket policies selecting the appropriate strategy based on access patterns and cost constraints.

Advanced Object Management

Immutable Versioning with Tombstone Deletion

Every write operation creates a new immutable version of the object, while deletes generate tombstone markers rather than immediate removal. Lines 87-93 describe this versioning schema, which supports audit trails, point-in-time recovery, and compliance requirements. The metadata store maintains the version chain, allowing clients to retrieve specific historical versions by specifying version IDs in their API requests.

Multipart Upload for Resilient Large File Handling

Large objects utilize multipart uploads, splitting files into chunks (typically 5-100 MB) that upload in parallel and assemble atomically upon completion. Lines 4-16 outline the three-phase process: initiate upload (receiving an UploadId), upload parts with associated ETag checksums, and complete the assembly via a manifest XML. This approach improves throughput through parallelization and provides resiliency against network failures—failed parts restart individually without retransmitting the entire object.

import requests, hashlib

# 1. Initiate multipart upload

init_resp = requests.post(
    "https://api.example.com/bucket/file.bin?uploads",
    headers={"Authorization": "<sig>"}
)
upload_id = init_resp.text.split("<UploadId>")[1].split("</UploadId>")[0]

# 2. Upload parts sequentially or in parallel

part_etags = []
with open("file.bin", "rb") as f:
    part_number = 1
    while chunk := f.read(5 * 1024 * 1024):  # 5 MiB parts

        resp = requests.put(
            f"https://api.example.com/bucket/file.bin?partNumber={part_number}&uploadId={upload_id}",
            data=chunk,
            headers={"Authorization": "<sig>", 
                     "Content-MD5": hashlib.md5(chunk).hexdigest()}
        )
        part_etags.append((part_number, resp.headers["ETag"]))
        part_number += 1

# 3. Complete the multipart upload

complete_xml = "<CompleteMultipartUpload>" + "".join(
    f"<Part><PartNumber>{pn}</PartNumber><ETag>{etag}</ETag></Part>"
    for pn, etag in part_etags
) + "</CompleteMultipartUpload>"

requests.post(
    f"https://api.example.com/bucket/file.bin?uploadId={upload_id}",
    data=complete_xml,
    headers={"Authorization": "<sig>", "Content-Type": "application/xml"}
)

Garbage Collection and Compaction

Background processes perform garbage collection to remove orphaned multipart parts, deleted object versions (tombstones), and consolidate small files into larger blocks. Lines 30-34 describe the compaction process that runs periodically on data nodes, reclaiming storage from incomplete uploads and expired versions while maintaining read performance by defragmenting physical storage layouts.

Consistency and Performance

Strong Consistency Guarantees on Writes

The implementation provides strong consistency by requiring the primary data node to replicate data to secondary nodes before acknowledging write success. Lines 31-34 in the consistency diagram illustrate this write flow: the API layer returns HTTP 200 only after the metadata store commits the transaction and the data store confirms replication to the target durability level (typically 2 additional copies or parity calculation). This ensures that subsequent read requests immediately observe the latest written state without session stickiness.

Summary

  • Flat namespaces with globally unique buckets eliminate hierarchy limitations and enable unlimited scaling.
  • Stateless API tiers allow horizontal scaling of request handling without session affinity or connection state.
  • Metadata/data separation optimizes storage costs by tiering hot metadata on fast storage and cold data on dense disks.
  • Consistent hashing provides uniform distribution and elastic cluster expansion without massive data rebalancing.
  • Configurable durability through 3-way replication (low latency) or 8+4 erasure coding (storage efficiency) adapts to workload requirements.
  • Immutable versioning with tombstone markers supports compliance, audit trails, and point-in-time recovery.
  • Multipart uploads parallelize large file transfers and provide fault tolerance against partial failures.
  • Background compaction prevents storage bloat and maintains node performance through garbage collection.
  • Strong write consistency guarantees immediate visibility of committed writes across all read paths.

Frequently Asked Questions

How does an S3-like object storage system handle large file uploads?

Large files utilize multipart uploads, where the client splits the object into 5-100 MB parts that upload independently and in parallel. According to 24. S3-like Object Storage/README.md (lines 4-16), the system assigns an UploadId for tracking, validates each part via MD5 checksums, and atomically assembles the final object only after receiving all parts and a completion manifest. This approach maximizes throughput and allows resumption of interrupted transfers without retransmitting completed parts.

What is the difference between replication and erasure coding in object storage?

Replication stores multiple complete copies (typically 3) of an object across failure domains, providing low-latency reads from the nearest copy but consuming 200% storage overhead. Erasure coding (such as 8+4 configuration) splits data into fragments and calculates parity, reducing overhead to 50% while maintaining equivalent durability through mathematical reconstruction. As implemented in the repository design (lines 25-31), replication serves hot data requiring millisecond access, while erasure coding optimizes cost for archival storage.

How does versioning work in an S3-compatible storage system?

Each PUT operation generates a new immutable version with a unique version ID, while DELETE operations create tombstone markers rather than removing data. Lines 87-93 of the README specify that the metadata store maintains the version chain, allowing retrieval of any historical state by specifying the version ID. This immutable approach supports regulatory compliance, accidental deletion recovery, and concurrent write safety without locking mechanisms.

Why separate metadata and data stores in object storage architecture?

The separation enables independent scaling and optimization of two distinct workload patterns: metadata operations require high IOPS for frequent small lookups (benefiting from SSDs), while data storage requires high throughput for large sequential transfers (suited for dense HDDs). As described in lines 19-20 and 85-90, this decoupling allows the metadata tier to scale based on query volume while the data tier scales based on storage capacity, preventing either workload from starving the other of resources.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →