How Easegress Implements Raft Consensus for High Availability
Easegress achieves high availability by embedding an etcd server in each node, delegating all Raft consensus operations to the battle-tested etcd implementation while exposing simple APIs for leader election, distributed locking, and state replication.
Easegress is a Cloud Native traffic gateway designed for high availability through a distributed, fault-tolerant cluster architecture. Rather than reinventing distributed consensus, the platform leverages the cluster package to embed a full etcd server into every node, utilizing etcd’s production-grade Raft engine for metadata replication and leader election. This approach allows Easegress to focus on lifecycle management and developer-friendly abstractions while inheriting etcd’s proven reliability for distributed coordination.
Embedded etcd Architecture
Easegress does not re-implement the Raft algorithm. Instead, each node runs an embedded etcd instance using embed.StartEtcd from the go.etcd.io/etcd/server/v3/embed package. The cluster package in pkg/cluster/cluster.go manages the complete lifecycle of this embedded server, from initial configuration to shutdown, while providing a unified interface for the rest of the application to interact with the consensus layer.
Cluster Configuration and Bootstrap
The bootstrap process begins in pkg/cluster/config.go, where CreateStaticClusterEtcdConfig generates the etcd configuration including client URLs, peer URLs, and snapshot tuning parameters. When cluster.New(opt) is invoked, it initializes the cluster layout and launches the run() goroutine, which orchestrates node startup through getReady() and startServer().
Nodes operate in two distinct roles determined by the --cluster-role flag:
- Primary nodes start a full embedded etcd server, becoming voting members of the Raft group. They register the cluster name under the
/eg/cluster/namekey. - Secondary nodes create only an etcd client connection, joining the existing cluster without participating in Raft log replication. They validate the cluster name and initialize a lease via
initLease()before becoming operational.
This distinction allows flexible deployment topologies where only a subset of nodes bear the storage and consensus overhead, while others act as stateless workers.
Leader Election and Node Coordination
Raft leader election occurs automatically within the embedded etcd layer. Easegress provides thin wrappers to expose this state to application logic, ensuring that exclusive operations execute only on the current leader.
Detecting the Raft Leader
The IsLeader() method in pkg/cluster/cluster.go (lines 61-68) determines whether the local node holds leadership by comparing server.Server.Leader() with server.Server.ID(). This boolean check enables components to gate singleton tasks such as log defragmentation or scheduled maintenance on the elected leader.
if cls.IsLeader() {
fmt.Println("This node is the Raft leader")
} else {
fmt.Println("This node is a follower")
}
Member Status and Heartbeats
Cluster membership health is maintained through a heartbeat mechanism implemented in pkg/cluster/cluster.go. Every node spawns a goroutine that periodically invokes heartbeat(), which in turn calls syncStatus() to write the node's MemberStatus to a lease-backed key at /status/members/<memberName>.
The lease mechanism is critical for fault detection: each node maintains a time-to-live lease that must be renewed periodically. If a node crashes, its lease expires, the status key disappears, and remaining members detect the failure through the absence of the key. This provides a consistent, consensus-backed view of cluster topology without custom failure detectors.
Distributed Primitives
Building on the embedded etcd foundation, Easegress implements higher-level distributed coordination primitives that ensure safety across the cluster.
Cluster-wide Mutex Implementation
Distributed mutual exclusion is provided through pkg/cluster/mutex.go, which wraps the go.etcd.io/etcd/client/v3/concurrency library. The Mutex struct creates a session backed by an etcd lease, ensuring that if a holding node crashes or becomes partitioned, the lock automatically releases when the lease expires.
m, err := cls.Mutex("critical-section-name")
if err != nil {
log.Fatalf("cannot create mutex: %v", err)
}
if err = m.Lock(); err != nil {
log.Fatalf("lock acquisition failed: %v", err)
}
// Execute exclusive work here...
if err = m.Unlock(); err != nil {
log.Fatalf("lock release failed: %v", err)
}
Configuration Syncing and Watching
The Syncer interface in pkg/cluster/syncer.go enables components to watch for configuration changes across the cluster. It supports both polling intervals and real-time watch streams, delivering updates through Go channels when keys under /config/ change.
syncer, _ := cls.Syncer(5 * time.Second) // Poll every 5s
ch, _ := syncer.Sync("/config/objects/my-pipeline")
go func() {
for cfg := range ch {
fmt.Printf("Pipeline configuration updated: %s\n", *cfg)
}
}()
Key Namespace and Data Layout
Consistency across the Raft-backed store relies on a deterministic key namespace defined in pkg/cluster/layout.go. All nodes agree on hierarchical paths such as /status/* for member health, /config/* for object definitions, and /leases/* for session management. This convention ensures that every read and write operation targets the same logical location, regardless of which node serves the request.
Practical Implementation Examples
Bootstrapping a Primary Node
To initialize a new Easegress cluster, start a primary node with the appropriate flags:
opt, _ := option.Parse([]string{
"--cluster-role=primary",
"--listen-peer-urls=http://localhost:2380",
"--listen-client-urls=http://localhost:2379",
})
cls, err := cluster.New(opt)
if err != nil {
log.Fatalf("cluster startup failed: %v", err)
}
defer cls.Close()
This sequence triggers run() → getReady() → startServer(), launching the embedded etcd instance and registering the node as a Raft voting member.
Summary
- Easegress embeds an etcd server per node rather than implementing Raft consensus directly, inheriting etcd’s proven fault tolerance and consistency guarantees.
- The
clusterpackage manages the embedded etcd lifecycle throughcluster.New()andstartServer(), handling configuration generation viaCreateStaticClusterEtcdConfig. - Leader detection is exposed through the
IsLeader()wrapper, which queries the embedded server's Raft state to determine election status. - Node health is maintained via lease-backed heartbeats written to
/status/members/, allowing automatic failure detection without custom protocols. - Distributed coordination relies on etcd's concurrency library for mutexes and the
Syncerinterface for configuration watching, both utilizing session leases for crash safety. - A deterministic key namespace defined in
layout.goensures consistent data organization across the replicated state machine.
Frequently Asked Questions
Does Easegress implement its own Raft algorithm?
No. Easegress delegates all consensus responsibilities to the embedded etcd server (go.etcd.io/etcd/server/v3/embed), which provides a production-grade Raft implementation. The Easegress codebase focuses on lifecycle management and providing convenient wrappers around etcd's client APIs.
How does Easegress handle node failover and leader election?
Failover is handled automatically by the embedded etcd's Raft engine. When the current leader fails, etcd elects a new leader from the remaining voting members. Easegress exposes this state through the IsLeader() method in pkg/cluster/cluster.go, allowing nodes to determine which instance should execute exclusive tasks like maintenance operations or garbage collection.
What is the difference between primary and secondary nodes in Easegress?
Primary nodes run a full embedded etcd server and participate in the Raft consensus group as voting members with persistent storage. Secondary nodes create only an etcd client connection to join the existing cluster without storing Raft logs locally or participating in leader election. The distinction is determined at startup by the --cluster-role flag and handled in the getReady() function within pkg/cluster/cluster.go.
How are distributed locks implemented in Easegress?
Distributed locks are implemented using etcd's concurrency library (go.etcd.io/etcd/client/v3/concurrency) as shown in pkg/cluster/mutex.go. Each lock is tied to an etcd session backed by a time-to-live lease, ensuring that if a holding node crashes, the lock is automatically released when the lease expires, preventing deadlocks in the distributed system.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →