How Durability Is Achieved in S3-Like Object Storage: Replication, Erasure Coding, and Checksums
Durability in S3-like object storage is achieved by replicating data across three independent Availability Zones and optionally using erasure coding schemes like 8+4, combined with failure-domain isolation and cryptographic checksum verification.
This architectural deep-dive into the liquidslr/system-design-notes repository explains how distributed object stores achieve eleven-nines durability. The design principles described in 24. S3-like Object Storage/README.md demonstrate how modern cloud storage systems balance cost, performance, and data persistence guarantees.
Multi-AZ Replication for Six-Nines Durability
The foundation of durability in S3-like systems starts with multi-AZ replication. According to the design documentation in 24. S3-like Object Storage/README.md, the system stores three complete copies of every object, with each replica residing in a different Availability Zone (AZ).
This triplication strategy yields 99.999999% (six-nines) durability, allowing the system to survive the complete failure of an entire data center while maintaining data accessibility. By distributing replicas across independent power, networking, and cooling domains, the architecture eliminates single points of failure at the physical infrastructure layer.
The Replication Strategy and Placement Service
Durability is enforced through a coordinated write path managed by a placement service. As implemented in the source design, this service uses consistent hashing to determine the primary storage node for any given object key.
The persistence flow follows these steps:
- The primary node receives the write request and persists the data locally.
- Before acknowledging the write to the client, the system replicates the payload to two secondary nodes located in different AZs.
- Only after successful replication to all three nodes does the system return a success response, ensuring strong consistency and high durability.
This approach ensures that even if two nodes fail simultaneously, the data remains available from the surviving replica.
Replication Implementation Example
The following Go pseudo-code from the repository illustrates the replication logic:
func StoreObject(bucket, key string, data []byte) (objID string, err error) {
// 1️⃣ Choose primary node via consistent hashing
primary := placementService.SelectPrimary(bucket, key)
// 2️⃣ Persist locally
objID = primary.Write(data)
// 3️⃣ Replicate to secondary nodes
replicas := placementService.SelectReplicas(primary, 2) // two secondaries
for _, node := range replicas {
if err = node.Replicate(objID, data); err != nil {
// trigger async retry / re‑balance
go retryReplication(node, objID, data)
}
}
// 4️⃣ Record checksum in metadata store
checksum := sha256.Sum256(data)
metadataStore.Put(objID, checksum[:])
return objID, nil
}
Notice that the function distributes replicas across distinct failure domains and records a cryptographic checksum for integrity verification.
Erasure Coding for Cost-Efficient Eleven-Nines Durability
For cold storage and cost-sensitive workloads, the system optionally replaces simple replication with erasure coding. The design implements an 8+4 erasure-coding scheme, splitting objects into eight data fragments and four parity fragments.
This configuration achieves 99.999999999% (eleven-nines) durability while reducing storage overhead by approximately 50% compared to triple replication. The trade-off involves increased CPU usage during encoding/decoding and slightly higher read latency due to the reconstruction process.
The fragments are distributed across distinct failure domains, allowing the system to reconstruct the original data even if up to four fragments become unavailable.
Erasure Coding Configuration
When enabled, the storage policy is configured via JSON as shown in the repository:
{
"erasure_coding": {
"data_shards": 8,
"parity_shards": 4,
"placement_policy": "spread_across_azs"
}
}
Failure-Domain Isolation and Automated Recovery
Durability depends on geographic and topological diversity. The placement service ensures that replicas—or erasure-coded fragments—are distributed across separate failure domains, including different racks, data centers, and network segments.
The system continuously monitors node health through periodic heartbeats. When the placement service detects a node failure, it automatically triggers re-replication or re-encoding to restore the desired redundancy level. This self-healing mechanism maintains durability guarantees even when multiple nodes within the same domain fail simultaneously.
Data Integrity Through Cryptographic Checksums
Beyond redundancy, durability requires protection against bit-rot and silent corruption. Each stored object or fragment is accompanied by a cryptographic checksum (typically SHA-256).
During read operations, the system validates data integrity against the stored checksum. If a mismatch is detected, the system immediately triggers re-replication from healthy nodes. This verification happens during write operations as well, ensuring that corrupted data never enters the storage system.
Summary
- Multi-AZ replication stores three copies of data across independent Availability Zones to achieve six-nines durability.
- The placement service uses consistent hashing to distribute primary and secondary replicas, requiring successful replication to all nodes before write acknowledgment.
- Erasure coding (8+4) provides eleven-nines durability with 50% lower storage costs, though with higher computational overhead.
- Failure-domain isolation places replicas across racks, data centers, and networks, with automated re-replication restoring redundancy after node failures.
- Cryptographic checksums detect and prevent silent data corruption through continuous integrity verification.
Frequently Asked Questions
What is the standard durability target for S3-like object storage?
The standard durability target for S3-like object storage is 99.999999% (six-nines), meaning the probability of data loss is one in one million objects over a year. This is achieved through triple replication across Availability Zones. With erasure coding enabled, systems can reach 99.999999999% (eleven-nines) durability, or one in one hundred billion objects lost per year.
How does erasure coding improve durability compared to simple replication?
Erasure coding improves durability by mathematically distributing redundancy across more failure domains than simple replication allows. While triple replication can tolerate two simultaneous failures, an 8+4 erasure coding scheme can tolerate the loss of any four fragments out of twelve. This mathematical redundancy, combined with wider distribution across nodes, increases durability from six-nines to eleven-nines while reducing storage overhead by approximately 50%.
What happens when a data node fails in a replicated object storage system?
When a data node fails, the placement service detects the outage through missing heartbeats. The system then initiates re-replication, copying data from healthy replicas to new nodes in different failure domains to restore the required redundancy level. This process happens automatically without client intervention, ensuring that the durability SLA is maintained even during infrastructure failures.
How does the placement service ensure replicas are geographically distributed?
The placement service uses a failure-domain-aware consistent hashing algorithm that maps object keys to primary nodes while explicitly selecting secondary nodes from different Availability Zones, racks, and network segments. The SelectReplicas function enforces anti-affinity rules, preventing the system from placing multiple replicas in the same physical location or on nodes sharing power or cooling infrastructure.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →