How to Set Up High Availability for Self-Hosted Services: A Layered Architecture Guide

High availability for self-hosted services requires eliminating single points of failure through redundant control planes, distributed data stores, and automated load balancing to ensure continuous operation during node or network failures.

Setting up high availability for self-hosted services ensures your infrastructure remains operational during hardware failures or network outages. The mikeroyal/Self-Hosting-Guide repository provides a comprehensive roadmap for building resilient systems using battle-tested open-source tools documented in the README.md High Availability section (lines ~410-420). This guide walks through the specific architecture patterns and configuration files needed to deploy a fault-tolerant self-hosted environment.

Understanding High Availability Layers

The Self-Hosting-Guide recommends a layered approach that combines redundant control planes, distributed data stores, and reliable front-end load balancers. This architecture ensures that no single component can disable the entire system.

Control-Plane HA

Control-plane high availability involves running multiple master or manager nodes so the orchestration layer survives individual node failures. The repository highlights k3s-ansible (line 414) as a fully automated HA installer that deploys k3s with etcd and adds a virtual IP via kube-vip alongside MetalLB for load balancing. For non-Kubernetes clusters, the Corosync Cluster Engine (documented at line 420) provides group communication systems to coordinate state across HA clusters.

Data-Layer HA

Data replication prevents storage node loss from corrupting datasets. Apache Cassandra (listed at line 1151) offers distributed NoSQL storage with automatic replication across nodes. For Kubernetes deployments, etcd stores cluster state in a quorum of nodes, ensuring the control plane remains consistent even if individual nodes fail.

Load-Balancing and Front-End HA

Distributing traffic across service instances requires robust load balancers. HAProxy (line 556) functions as a software load balancer that terminates TLS, performs health checks, and routes to multiple backends. For bare-metal Kubernetes, MetalLB works with kube-vip to provide external IPs, while Traefik serves as a dynamic reverse-proxy that discovers services via Kubernetes Ingress.

Application-Level HA

Running multiple replicas of each container ensures service continuity. Docker Compose supports this through the replicas directive in compose files, while Kubernetes Deployments use spec.replicas to maintain desired pod counts across the cluster.

Automated Fail-Over

Detection and recovery mechanisms include Autoheal for restarting unhealthy Docker containers based on health-check status, and WatchTower for monitoring new image versions and redeploying containers automatically.

Configuring k3s-ansible for Control-Plane Redundancy

The k3s-ansible setup creates a three-node control plane with etcd quorum. Create an inventory file at inventory/hosts.ini:

[master]
master1 ansible_host=10.0.0.1
master2 ansible_host=10.0.0.2
master3 ansible_host=10.0.0.3

[node]
worker1 ansible_host=10.0.0.4
worker2 ansible_host=10.0.0.5

[k3s_cluster:children]
master
node

Execute the playbook with variables enabling HA features:

ansible-playbook -i inventory/hosts.ini site.yml \
  -e "k3s_control_master=true k3s_etcd_datadir=/var/lib/etcd" \
  -e "k3s_kube_vip_address=10.0.0.100" \
  -e "k3s_metallb_range=10.0.0.200-10.0.0.210"

This configuration deploys kube-vip for virtual IP management and MetalLB for service load balancing across the cluster.

Setting Up HAProxy for Front-End Load Balancing

For the ingress layer, deploy HAProxy instances in active-passive mode behind a virtual IP. This configuration performs health checks on backend services as documented in the repository:

global
    log /dev/log local0
    maxconn 2000
    daemon

defaults
    log     global
    mode    http
    option  httplog
    timeout connect 5s
    timeout client  30s
    timeout server  30s

frontend http_in
    bind 0.0.0.0:80
    default_backend app_servers

backend app_servers
    balance roundrobin
    option httpchk GET /healthz
    server app1 10.0.0.11:80 check
    server app2 10.0.0.12:80 check

Deploy two instances behind a virtual IP using keepalived or similar to eliminate the load balancer as a single point of failure.

Implementing Application Replicas in Docker Compose

For container-level high availability without Kubernetes, use Docker Compose with replica scaling:

version: "3.8"
services:
  web:
    image: nginx:stable
    deploy:
      replicas: 3
      restart_policy:
        condition: on-failure
    ports:
      - "8080:80"

Run with docker compose up -d. When paired with a reverse-proxy like HAProxy or Traefik, this setup provides redundancy at the container level.

Deploying High Availability in Kubernetes

Kubernetes Deployments natively support high availability through the replicas field. This manifest ensures four pods remain running across the cluster:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: myapp
spec:
  replicas: 4
  selector:
    matchLabels:
      app: myapp
  template:
    metadata:
      labels:
        app: myapp
    spec:
      containers:
      - name: myapp
        image: myorg/myapp:latest
        ports:
        - containerPort: 80

Apply with kubectl apply -f deployment.yaml. The control-plane (k3s) automatically spreads pods across nodes and reschedules them if nodes become unavailable.

Summary

  • Redundant control planes: Use k3s-ansible with three master nodes and etcd quorum to prevent orchestration failures.
  • Distributed data stores: Implement Apache Cassandra or etcd replication to protect against storage node loss.
  • Load balancing: Deploy HAProxy with health checks or MetalLB with kube-vip for traffic distribution and virtual IP management.
  • Application replicas: Configure Docker Compose with replicas: 3 or Kubernetes Deployments with spec.replicas: 4 to maintain service availability.
  • Auto-recovery: Implement Autoheal for container health checks and WatchTower for automated updates.

Frequently Asked Questions

What constitutes a single point of failure in self-hosted services?

A single point of failure occurs when one component—such as a single server, database instance, or load balancer—can disable the entire service if it stops working. The Self-Hosting-Guide eliminates these through redundant control planes, replicated data stores, and multiple load balancer instances behind virtual IPs.

How does k3s-ansible provide high availability?

The k3s-ansible installer (documented at line 414 of README.md) automatically configures three control-plane nodes running etcd in a quorum, deploys kube-vip for virtual IP failover, and installs MetalLB to distribute external traffic across the cluster.

Should I use Docker Compose or Kubernetes for high availability?

Docker Compose with replicas suits smaller deployments requiring simple container redundancy, while Kubernetes provides superior orchestration capabilities including automatic pod rescheduling, node affinity rules, and integrated service discovery. The guide recommends Kubernetes via k3s for production high availability.

How do I monitor high availability health?

Implement Prometheus and Alertmanager to watch node health, configure HAProxy health checks to remove failed backends automatically, and use Autoheal to restart unhealthy containers. These tools provide visibility into cluster state and automated recovery from failures.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →