# How to Set Up High Availability for Self-Hosted Services: A Layered Architecture Guide

> Learn how to set up high availability for self-hosted services. Eliminate single points of failure with redundancy, distributed data, and load balancing for continuous operation.

- Repository: [Michael Royal/Self-Hosting-Guide](https://github.com/mikeroyal/Self-Hosting-Guide)
- Tags: architecture
- Published: 2026-06-17

---

**High availability for self-hosted services requires eliminating single points of failure through redundant control planes, distributed data stores, and automated load balancing to ensure continuous operation during node or network failures.**

Setting up high availability for self-hosted services ensures your infrastructure remains operational during hardware failures or network outages. The mikeroyal/Self-Hosting-Guide repository provides a comprehensive roadmap for building resilient systems using battle-tested open-source tools documented in the README.md High Availability section (lines ~410-420). This guide walks through the specific architecture patterns and configuration files needed to deploy a fault-tolerant self-hosted environment.

## Understanding High Availability Layers

The Self-Hosting-Guide recommends a layered approach that combines **redundant control planes**, **distributed data stores**, and **reliable front-end load balancers**. This architecture ensures that no single component can disable the entire system.

### Control-Plane HA

Control-plane high availability involves running multiple master or manager nodes so the orchestration layer survives individual node failures. The repository highlights **k3s-ansible** (line 414) as a fully automated HA installer that deploys k3s with etcd and adds a virtual IP via **kube-vip** alongside **MetalLB** for load balancing. For non-Kubernetes clusters, the **Corosync Cluster Engine** (documented at line 420) provides group communication systems to coordinate state across HA clusters.

### Data-Layer HA

Data replication prevents storage node loss from corrupting datasets. **Apache Cassandra** (listed at line 1151) offers distributed NoSQL storage with automatic replication across nodes. For Kubernetes deployments, **etcd** stores cluster state in a quorum of nodes, ensuring the control plane remains consistent even if individual nodes fail.

### Load-Balancing and Front-End HA

Distributing traffic across service instances requires robust load balancers. **HAProxy** (line 556) functions as a software load balancer that terminates TLS, performs health checks, and routes to multiple backends. For bare-metal Kubernetes, **MetalLB** works with kube-vip to provide external IPs, while **Traefik** serves as a dynamic reverse-proxy that discovers services via Kubernetes Ingress.

### Application-Level HA

Running multiple replicas of each container ensures service continuity. **Docker Compose** supports this through the `replicas` directive in compose files, while **Kubernetes Deployments** use `spec.replicas` to maintain desired pod counts across the cluster.

### Automated Fail-Over

Detection and recovery mechanisms include **Autoheal** for restarting unhealthy Docker containers based on health-check status, and **WatchTower** for monitoring new image versions and redeploying containers automatically.

## Configuring k3s-ansible for Control-Plane Redundancy

The k3s-ansible setup creates a three-node control plane with etcd quorum. Create an inventory file at [`inventory/hosts.ini`](https://github.com/mikeroyal/Self-Hosting-Guide/blob/main/inventory/hosts.ini):

```ini
[master]
master1 ansible_host=10.0.0.1
master2 ansible_host=10.0.0.2
master3 ansible_host=10.0.0.3

[node]
worker1 ansible_host=10.0.0.4
worker2 ansible_host=10.0.0.5

[k3s_cluster:children]
master
node

```

Execute the playbook with variables enabling HA features:

```bash
ansible-playbook -i inventory/hosts.ini site.yml \
  -e "k3s_control_master=true k3s_etcd_datadir=/var/lib/etcd" \
  -e "k3s_kube_vip_address=10.0.0.100" \
  -e "k3s_metallb_range=10.0.0.200-10.0.0.210"

```

This configuration deploys kube-vip for virtual IP management and MetalLB for service load balancing across the cluster.

## Setting Up HAProxy for Front-End Load Balancing

For the ingress layer, deploy HAProxy instances in active-passive mode behind a virtual IP. This configuration performs health checks on backend services as documented in the repository:

```haproxy
global
    log /dev/log local0
    maxconn 2000
    daemon

defaults
    log     global
    mode    http
    option  httplog
    timeout connect 5s
    timeout client  30s
    timeout server  30s

frontend http_in
    bind 0.0.0.0:80
    default_backend app_servers

backend app_servers
    balance roundrobin
    option httpchk GET /healthz
    server app1 10.0.0.11:80 check
    server app2 10.0.0.12:80 check

```

Deploy two instances behind a virtual IP using keepalived or similar to eliminate the load balancer as a single point of failure.

## Implementing Application Replicas in Docker Compose

For container-level high availability without Kubernetes, use Docker Compose with replica scaling:

```yaml
version: "3.8"
services:
  web:
    image: nginx:stable
    deploy:
      replicas: 3
      restart_policy:
        condition: on-failure
    ports:
      - "8080:80"

```

Run with `docker compose up -d`. When paired with a reverse-proxy like HAProxy or Traefik, this setup provides redundancy at the container level.

## Deploying High Availability in Kubernetes

Kubernetes Deployments natively support high availability through the replicas field. This manifest ensures four pods remain running across the cluster:

```yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: myapp
spec:
  replicas: 4
  selector:
    matchLabels:
      app: myapp
  template:
    metadata:
      labels:
        app: myapp
    spec:
      containers:
      - name: myapp
        image: myorg/myapp:latest
        ports:
        - containerPort: 80

```

Apply with `kubectl apply -f deployment.yaml`. The control-plane (k3s) automatically spreads pods across nodes and reschedules them if nodes become unavailable.

## Summary

- **Redundant control planes**: Use k3s-ansible with three master nodes and etcd quorum to prevent orchestration failures.
- **Distributed data stores**: Implement Apache Cassandra or etcd replication to protect against storage node loss.
- **Load balancing**: Deploy HAProxy with health checks or MetalLB with kube-vip for traffic distribution and virtual IP management.
- **Application replicas**: Configure Docker Compose with `replicas: 3` or Kubernetes Deployments with `spec.replicas: 4` to maintain service availability.
- **Auto-recovery**: Implement Autoheal for container health checks and WatchTower for automated updates.

## Frequently Asked Questions

### What constitutes a single point of failure in self-hosted services?

A single point of failure occurs when one component—such as a single server, database instance, or load balancer—can disable the entire service if it stops working. The Self-Hosting-Guide eliminates these through redundant control planes, replicated data stores, and multiple load balancer instances behind virtual IPs.

### How does k3s-ansible provide high availability?

The k3s-ansible installer (documented at line 414 of README.md) automatically configures three control-plane nodes running etcd in a quorum, deploys kube-vip for virtual IP failover, and installs MetalLB to distribute external traffic across the cluster.

### Should I use Docker Compose or Kubernetes for high availability?

Docker Compose with replicas suits smaller deployments requiring simple container redundancy, while Kubernetes provides superior orchestration capabilities including automatic pod rescheduling, node affinity rules, and integrated service discovery. The guide recommends Kubernetes via k3s for production high availability.

### How do I monitor high availability health?

Implement Prometheus and Alertmanager to watch node health, configure HAProxy health checks to remove failed backends automatically, and use Autoheal to restart unhealthy containers. These tools provide visibility into cluster state and automated recovery from failures.