# Designing for High Availability: A 7-Step Blueprint from Production Systems

> Learn to design for high availability with our 7-step blueprint. Eliminate single points of failure using replication, health checks, and consensus protocols for reliable production systems.

- Repository: [Gaurav Kumar/system-design-notes](https://github.com/liquidslr/system-design-notes)
- Tags: how-to-guide
- Published: 2026-09-11

---

**Designing for high availability requires eliminating single points of failure through strategic replication, automated failure detection via heartbeats and health checks, and deterministic recovery using consensus protocols like Raft and event sourcing.**

The `liquidslr/system-design-notes` repository documents architectural patterns from latency-critical systems such as stock exchanges and distributed key-value stores. This guide synthesizes those proven strategies into a practical framework for designing for high availability, targeting the industry standard of **99.99% uptime** (approximately 8 minutes of downtime per year).

## Define Availability Targets and Eliminate Single Points of Failure

Start by quantifying your **Service Level Agreements (SLAs)** using availability percentages or "nines." According to `02. Back-of-the-Envelope Estimation/Readme.md`, each additional nine drastically reduces allowed downtime—moving from 99.9% (52 minutes/year) to 99.99% (8 minutes/year) requires an order of magnitude improvement in reliability.

Next, audit your architecture for **Single Points of Failure (SPOFs)**. In `28. Stock Exchange/README.md`, the authors identify the **matching engine**, **sequencer**, and **event store** as critical components that would halt trading if they failed. Map these dependencies in your own system to determine what requires immediate replication.

## Replication Strategies for Stateless and Stateful Services

Different components demand different replication strategies. The repository outlines three distinct patterns in `28. Stock Exchange/README.md` and `06. Key-Value Store/README.md`:

**Stateless services** (e.g., client gateways, API servers) scale horizontally behind load balancers. These can be duplicated infinitely without consistency concerns, as shown in the stock exchange's stateless gateway tier.

**Stateful services** (e.g., order managers, databases) require **active-active replication** with leader election. The stock exchange implementation uses **Raft consensus** to maintain a consistent leader across replicas while allowing automatic failover if the leader stops sending heartbeats.

**Event stores and logs** replicate across data centers using reliable UDP or consensus protocols. This ensures that even during regional outages, the event stream remains durable and can reconstruct system state.

## Implement Automated Failure Detection and Failover

High availability depends on detecting failures faster than human response times. The `28. Stock Exchange/README.md` describes a multi-layered detection strategy:

- **Heartbeat mechanisms**: Services broadcast periodic "I’m alive" messages via UDP multicast. When a replica misses three consecutive heartbeats (typically 6 seconds), surviving nodes trigger a leader election.
- **Health-check endpoints**: Expose `/healthz` endpoints that probe essential dependencies. A degraded service reports itself unhealthy, causing the load balancer to route traffic to healthy replicas.
- **Circuit breakers**: Wrap outbound calls to external services so that cascading failures trigger an open circuit, protecting the core system from dependency crashes.

```go
// Simplified Raft-style leader election from system-design-notes
type Node struct {
    ID        string
    IsLeader  bool
    LastSeen  map[string]time.Time
    Heartbeat chan string
}

func (n *Node) monitorPeers() {
    for {
        time.Sleep(5 * time.Second)
        for id, ts := range n.LastSeen {
            if time.Since(ts) > 6*time.Second {
                n.startElection() // Peer considered down
            }
        }
    }
}

```

## Ensure Deterministic Recovery with Event Sourcing

When failures occur, systems must recover to a consistent state without data loss. The repository emphasizes **event sourcing** as implemented in `28. Stock Exchange/README.md`—where all state changes are appended to an immutable log. By replaying events from the event store, a failed node reconstructs its exact previous state deterministically.

For distributed data, **quorum consensus** ensures durability before acknowledging writes. As detailed in `06. Key-Value Store/README.md`, writes must persist to a majority (N/2 + 1) of replicas before returning success to the client. This guarantees that even if minority partitions fail, committed data survives.

```java
// Event-sourced write path from Key-Value Store design
class LogEntry {
    UUID id;
    String key;
    String value;
    long timestamp;
}

public void put(String key, String value) {
    LogEntry e = new LogEntry(UUID.randomUUID(), key, value, System.currentTimeMillis());
    quorumWrite(e); // Write to majority of replicas
    map.put(key, value); // Apply to in-memory state
}

```

## Validate Resilience with Chaos Engineering and Monitoring

Designing for high availability is incomplete without verification. The `28. Stock Exchange/README.md` references **Chaos Engineering**—intentionally terminating random replicas to verify that automated failover functions correctly under real conditions.

Complement chaos testing with comprehensive observability from `20. Metrics Monitoring and Alerting System/README.md`. Track **99th and 99.9th percentile latencies** alongside availability metrics; a system is not truly available if latency spikes render it unusable.

```python

# Health-check implementation pattern

from flask import Flask, jsonify
import psutil

app = Flask(__name__)

@app.route("/healthz")
def health():
    db_ok = check_database()
    cpu_ok = psutil.cpu_percent() < 80
    return jsonify({
        "status": "UP" if db_ok and cpu_ok else "DOWN",
        "db": db_ok,
        "cpu": cpu_ok
    })

```

## Summary

Designing for high availability follows a systematic methodology distilled from the `liquidslr/system-design-notes` repository:

- **Quantify targets** using availability "nines" to drive architectural decisions
- **Eliminate SPOFs** by replicating stateless services horizontally and stateful services using Raft leader election
- **Automate detection** through heartbeats, health checks, and circuit breakers
- **Ensure recovery** via event sourcing and quorum consensus for deterministic state reconstruction
- **Validate continuously** using chaos engineering and latency percentile monitoring

## Frequently Asked Questions

### What is the difference between high availability and fault tolerance?

High availability focuses on minimizing downtime and ensuring rapid recovery when components fail, typically through redundancy and automated failover. Fault tolerance aims to prevent any service interruption entirely, often through more expensive hardware or complex software masking of errors. According to the repository's stock exchange design, achieving high availability involves accepting brieffailover windows (seconds) while maintaining data consistency through Raft consensus, whereas true fault tolerance would require hot-spare systems with zero switchover time.

### How does Raft consensus specifically enable high availability?

Raft consensus, as implemented in `28. Stock Exchange/README.md`, maintains a distributed log of commands across multiple replicas. By requiring a majority quorum for write operations and using heartbeats to detect leader failures, Raft ensures that committed writes survive even if minority nodes crash. When a leader fails, remaining nodes elect a new leader within milliseconds using the replicated log, allowing the system to continue processing transactions without manual intervention or data inconsistency.

### What are the most common single points of failure in distributed systems?

The repository identifies three critical SPOFs that appear across multiple designs: the **primary database** (single write node), the **message sequencer** (single ordering authority), and the **event store** (single source of truth for state reconstruction). In `28. Stock Exchange/README.md`, the matching engine represents a compute SPOF, while in `19. Distributed Message Queue/README.md`, the centralized broker topology creates a messaging SPOF that must be addressed through clustered broker architectures.

### How do you calculate the allowed downtime for a specific availability target?

Use the reference table in `02. Back-of-the-Envelope Estimation/Readme.md` to calculate downtime windows. For a target of **99.9% availability** (three nines), you have approximately 8.76 hours of allowed downtime per year. For **99.99%** (four nines), this drops to 52.56 minutes per year. When designing for high availability, divide this budget across planned maintenance, deployment risks, and unexpected failures to determine if your architecture can realistically meet the SLA.