# How Durability Is Achieved in S3-Like Object Storage: Replication, Erasure Coding, and Checksums

> Discover how S3-like object storage achieves high durability through data replication, erasure coding, and checksums. Learn about replication, erasure coding, and checksums.

- Repository: [Gaurav Kumar/system-design-notes](https://github.com/liquidslr/system-design-notes)
- Tags: deep-dive
- Published: 2026-09-11

---

**Durability in S3-like object storage is achieved by replicating data across three independent Availability Zones and optionally using erasure coding schemes like 8+4, combined with failure-domain isolation and cryptographic checksum verification.**

This architectural deep-dive into the `liquidslr/system-design-notes` repository explains how distributed object stores achieve eleven-nines durability. The design principles described in `24. S3-like Object Storage/README.md` demonstrate how modern cloud storage systems balance cost, performance, and data persistence guarantees.

## Multi-AZ Replication for Six-Nines Durability

The foundation of durability in S3-like systems starts with **multi-AZ replication**. According to the design documentation in `24. S3-like Object Storage/README.md`, the system stores three complete copies of every object, with each replica residing in a different Availability Zone (AZ).

This triplication strategy yields **99.999999% (six-nines) durability**, allowing the system to survive the complete failure of an entire data center while maintaining data accessibility. By distributing replicas across independent power, networking, and cooling domains, the architecture eliminates single points of failure at the physical infrastructure layer.

## The Replication Strategy and Placement Service

Durability is enforced through a coordinated write path managed by a **placement service**. As implemented in the source design, this service uses **consistent hashing** to determine the primary storage node for any given object key.

The persistence flow follows these steps:

1. The **primary node** receives the write request and persists the data locally.
2. Before acknowledging the write to the client, the system replicates the payload to **two secondary nodes** located in different AZs.
3. Only after successful replication to all three nodes does the system return a success response, ensuring **strong consistency** and high durability.

This approach ensures that even if two nodes fail simultaneously, the data remains available from the surviving replica.

### Replication Implementation Example

The following Go pseudo-code from the repository illustrates the replication logic:

```go
func StoreObject(bucket, key string, data []byte) (objID string, err error) {
    // 1️⃣ Choose primary node via consistent hashing
    primary := placementService.SelectPrimary(bucket, key)

    // 2️⃣ Persist locally
    objID = primary.Write(data)

    // 3️⃣ Replicate to secondary nodes
    replicas := placementService.SelectReplicas(primary, 2) // two secondaries
    for _, node := range replicas {
        if err = node.Replicate(objID, data); err != nil {
            // trigger async retry / re‑balance
            go retryReplication(node, objID, data)
        }
    }
    // 4️⃣ Record checksum in metadata store
    checksum := sha256.Sum256(data)
    metadataStore.Put(objID, checksum[:])
    return objID, nil
}

```

Notice that the function distributes replicas across distinct failure domains and records a cryptographic checksum for integrity verification.

## Erasure Coding for Cost-Efficient Eleven-Nines Durability

For cold storage and cost-sensitive workloads, the system optionally replaces simple replication with **erasure coding**. The design implements an **8+4 erasure-coding scheme**, splitting objects into eight data fragments and four parity fragments.

This configuration achieves **99.999999999% (eleven-nines) durability** while reducing storage overhead by approximately 50% compared to triple replication. The trade-off involves increased CPU usage during encoding/decoding and slightly higher read latency due to the reconstruction process.

The fragments are distributed across distinct failure domains, allowing the system to reconstruct the original data even if up to four fragments become unavailable.

### Erasure Coding Configuration

When enabled, the storage policy is configured via JSON as shown in the repository:

```json
{
  "erasure_coding": {
    "data_shards": 8,
    "parity_shards": 4,
    "placement_policy": "spread_across_azs"
  }
}

```

## Failure-Domain Isolation and Automated Recovery

Durability depends on geographic and topological diversity. The placement service ensures that replicas—or erasure-coded fragments—are distributed across separate **failure domains**, including different racks, data centers, and network segments.

The system continuously monitors node health through periodic heartbeats. When the placement service detects a node failure, it automatically triggers **re-replication** or re-encoding to restore the desired redundancy level. This self-healing mechanism maintains durability guarantees even when multiple nodes within the same domain fail simultaneously.

## Data Integrity Through Cryptographic Checksums

Beyond redundancy, durability requires protection against bit-rot and silent corruption. Each stored object or fragment is accompanied by a **cryptographic checksum** (typically SHA-256).

During read operations, the system validates data integrity against the stored checksum. If a mismatch is detected, the system immediately triggers re-replication from healthy nodes. This verification happens during write operations as well, ensuring that corrupted data never enters the storage system.

## Summary

- **Multi-AZ replication** stores three copies of data across independent Availability Zones to achieve six-nines durability.
- The **placement service** uses consistent hashing to distribute primary and secondary replicas, requiring successful replication to all nodes before write acknowledgment.
- **Erasure coding (8+4)** provides eleven-nines durability with 50% lower storage costs, though with higher computational overhead.
- **Failure-domain isolation** places replicas across racks, data centers, and networks, with automated re-replication restoring redundancy after node failures.
- **Cryptographic checksums** detect and prevent silent data corruption through continuous integrity verification.

## Frequently Asked Questions

### What is the standard durability target for S3-like object storage?

The standard durability target for S3-like object storage is **99.999999% (six-nines)**, meaning the probability of data loss is one in one million objects over a year. This is achieved through triple replication across Availability Zones. With erasure coding enabled, systems can reach **99.999999999% (eleven-nines)** durability, or one in one hundred billion objects lost per year.

### How does erasure coding improve durability compared to simple replication?

Erasure coding improves durability by mathematically distributing redundancy across more failure domains than simple replication allows. While triple replication can tolerate two simultaneous failures, an **8+4 erasure coding scheme** can tolerate the loss of any four fragments out of twelve. This mathematical redundancy, combined with wider distribution across nodes, increases durability from six-nines to eleven-nines while reducing storage overhead by approximately 50%.

### What happens when a data node fails in a replicated object storage system?

When a data node fails, the **placement service** detects the outage through missing heartbeats. The system then initiates **re-replication**, copying data from healthy replicas to new nodes in different failure domains to restore the required redundancy level. This process happens automatically without client intervention, ensuring that the durability SLA is maintained even during infrastructure failures.

### How does the placement service ensure replicas are geographically distributed?

The placement service uses a **failure-domain-aware consistent hashing algorithm** that maps object keys to primary nodes while explicitly selecting secondary nodes from different Availability Zones, racks, and network segments. The `SelectReplicas` function enforces anti-affinity rules, preventing the system from placing multiple replicas in the same physical location or on nodes sharing power or cooling infrastructure.