# How Easegress Implements Raft Consensus for High Availability

> Discover how Easegress ensures high availability by leveraging etcd for Raft consensus, simplifying leader election distributed locking and state replication with easy APIs.

- Repository: [MegaEase/easegress](https://github.com/megaease/easegress)
- Tags: internals
- Published: 2026-03-07

---

**Easegress achieves high availability by embedding an etcd server in each node, delegating all Raft consensus operations to the battle-tested etcd implementation while exposing simple APIs for leader election, distributed locking, and state replication.**

Easegress is a Cloud Native traffic gateway designed for high availability through a distributed, fault-tolerant cluster architecture. Rather than reinventing distributed consensus, the platform leverages the `cluster` package to embed a full etcd server into every node, utilizing etcd’s production-grade Raft engine for metadata replication and leader election. This approach allows Easegress to focus on lifecycle management and developer-friendly abstractions while inheriting etcd’s proven reliability for distributed coordination.

## Embedded etcd Architecture

Easegress does not re-implement the Raft algorithm. Instead, each node runs an embedded etcd instance using `embed.StartEtcd` from the `go.etcd.io/etcd/server/v3/embed` package. The `cluster` package in [`pkg/cluster/cluster.go`](https://github.com/megaease/easegress/blob/main/pkg/cluster/cluster.go) manages the complete lifecycle of this embedded server, from initial configuration to shutdown, while providing a unified interface for the rest of the application to interact with the consensus layer.

### Cluster Configuration and Bootstrap

The bootstrap process begins in [`pkg/cluster/config.go`](https://github.com/megaease/easegress/blob/main/pkg/cluster/config.go), where `CreateStaticClusterEtcdConfig` generates the etcd configuration including client URLs, peer URLs, and snapshot tuning parameters. When `cluster.New(opt)` is invoked, it initializes the cluster layout and launches the `run()` goroutine, which orchestrates node startup through `getReady()` and `startServer()`.

Nodes operate in two distinct roles determined by the `--cluster-role` flag:

- **Primary nodes** start a full embedded etcd server, becoming voting members of the Raft group. They register the cluster name under the `/eg/cluster/name` key.
- **Secondary nodes** create only an etcd client connection, joining the existing cluster without participating in Raft log replication. They validate the cluster name and initialize a lease via `initLease()` before becoming operational.

This distinction allows flexible deployment topologies where only a subset of nodes bear the storage and consensus overhead, while others act as stateless workers.

## Leader Election and Node Coordination

Raft leader election occurs automatically within the embedded etcd layer. Easegress provides thin wrappers to expose this state to application logic, ensuring that exclusive operations execute only on the current leader.

### Detecting the Raft Leader

The `IsLeader()` method in [`pkg/cluster/cluster.go`](https://github.com/megaease/easegress/blob/main/pkg/cluster/cluster.go) (lines 61-68) determines whether the local node holds leadership by comparing `server.Server.Leader()` with `server.Server.ID()`. This boolean check enables components to gate singleton tasks such as log defragmentation or scheduled maintenance on the elected leader.

```go
if cls.IsLeader() {
    fmt.Println("This node is the Raft leader")
} else {
    fmt.Println("This node is a follower")
}

```

### Member Status and Heartbeats

Cluster membership health is maintained through a heartbeat mechanism implemented in [`pkg/cluster/cluster.go`](https://github.com/megaease/easegress/blob/main/pkg/cluster/cluster.go). Every node spawns a goroutine that periodically invokes `heartbeat()`, which in turn calls `syncStatus()` to write the node's `MemberStatus` to a lease-backed key at `/status/members/<memberName>`.

The lease mechanism is critical for fault detection: each node maintains a time-to-live lease that must be renewed periodically. If a node crashes, its lease expires, the status key disappears, and remaining members detect the failure through the absence of the key. This provides a consistent, consensus-backed view of cluster topology without custom failure detectors.

## Distributed Primitives

Building on the embedded etcd foundation, Easegress implements higher-level distributed coordination primitives that ensure safety across the cluster.

### Cluster-wide Mutex Implementation

Distributed mutual exclusion is provided through [`pkg/cluster/mutex.go`](https://github.com/megaease/easegress/blob/main/pkg/cluster/mutex.go), which wraps the `go.etcd.io/etcd/client/v3/concurrency` library. The `Mutex` struct creates a session backed by an etcd lease, ensuring that if a holding node crashes or becomes partitioned, the lock automatically releases when the lease expires.

```go
m, err := cls.Mutex("critical-section-name")
if err != nil {
    log.Fatalf("cannot create mutex: %v", err)
}
if err = m.Lock(); err != nil {
    log.Fatalf("lock acquisition failed: %v", err)
}
// Execute exclusive work here...
if err = m.Unlock(); err != nil {
    log.Fatalf("lock release failed: %v", err)
}

```

### Configuration Syncing and Watching

The `Syncer` interface in [`pkg/cluster/syncer.go`](https://github.com/megaease/easegress/blob/main/pkg/cluster/syncer.go) enables components to watch for configuration changes across the cluster. It supports both polling intervals and real-time watch streams, delivering updates through Go channels when keys under `/config/` change.

```go
syncer, _ := cls.Syncer(5 * time.Second) // Poll every 5s
ch, _ := syncer.Sync("/config/objects/my-pipeline")

go func() {
    for cfg := range ch {
        fmt.Printf("Pipeline configuration updated: %s\n", *cfg)
    }
}()

```

## Key Namespace and Data Layout

Consistency across the Raft-backed store relies on a deterministic key namespace defined in [`pkg/cluster/layout.go`](https://github.com/megaease/easegress/blob/main/pkg/cluster/layout.go). All nodes agree on hierarchical paths such as `/status/*` for member health, `/config/*` for object definitions, and `/leases/*` for session management. This convention ensures that every read and write operation targets the same logical location, regardless of which node serves the request.

## Practical Implementation Examples

### Bootstrapping a Primary Node

To initialize a new Easegress cluster, start a primary node with the appropriate flags:

```go
opt, _ := option.Parse([]string{
    "--cluster-role=primary",
    "--listen-peer-urls=http://localhost:2380",
    "--listen-client-urls=http://localhost:2379",
})

cls, err := cluster.New(opt)
if err != nil {
    log.Fatalf("cluster startup failed: %v", err)
}
defer cls.Close()

```

This sequence triggers `run()` → `getReady()` → `startServer()`, launching the embedded etcd instance and registering the node as a Raft voting member.

## Summary

- Easegress embeds an etcd server per node rather than implementing Raft consensus directly, inheriting etcd’s proven fault tolerance and consistency guarantees.
- The `cluster` package manages the embedded etcd lifecycle through `cluster.New()` and `startServer()`, handling configuration generation via `CreateStaticClusterEtcdConfig`.
- Leader detection is exposed through the `IsLeader()` wrapper, which queries the embedded server's Raft state to determine election status.
- Node health is maintained via lease-backed heartbeats written to `/status/members/`, allowing automatic failure detection without custom protocols.
- Distributed coordination relies on etcd's concurrency library for mutexes and the `Syncer` interface for configuration watching, both utilizing session leases for crash safety.
- A deterministic key namespace defined in [`layout.go`](https://github.com/megaease/easegress/blob/main/layout.go) ensures consistent data organization across the replicated state machine.

## Frequently Asked Questions

### Does Easegress implement its own Raft algorithm?

No. Easegress delegates all consensus responsibilities to the embedded etcd server (`go.etcd.io/etcd/server/v3/embed`), which provides a production-grade Raft implementation. The Easegress codebase focuses on lifecycle management and providing convenient wrappers around etcd's client APIs.

### How does Easegress handle node failover and leader election?

Failover is handled automatically by the embedded etcd's Raft engine. When the current leader fails, etcd elects a new leader from the remaining voting members. Easegress exposes this state through the `IsLeader()` method in [`pkg/cluster/cluster.go`](https://github.com/megaease/easegress/blob/main/pkg/cluster/cluster.go), allowing nodes to determine which instance should execute exclusive tasks like maintenance operations or garbage collection.

### What is the difference between primary and secondary nodes in Easegress?

Primary nodes run a full embedded etcd server and participate in the Raft consensus group as voting members with persistent storage. Secondary nodes create only an etcd client connection to join the existing cluster without storing Raft logs locally or participating in leader election. The distinction is determined at startup by the `--cluster-role` flag and handled in the `getReady()` function within [`pkg/cluster/cluster.go`](https://github.com/megaease/easegress/blob/main/pkg/cluster/cluster.go).

### How are distributed locks implemented in Easegress?

Distributed locks are implemented using etcd's concurrency library (`go.etcd.io/etcd/client/v3/concurrency`) as shown in [`pkg/cluster/mutex.go`](https://github.com/megaease/easegress/blob/main/pkg/cluster/mutex.go). Each lock is tied to an etcd session backed by a time-to-live lease, ensuring that if a holding node crashes, the lock is automatically released when the lease expires, preventing deadlocks in the distributed system.