How Dragonboat Guarantees Linearizability for Concurrent Reads and Writes

Dragonboat achieves linearizability by coupling every write with client sessions that enforce at-most-once semantics, while using the Raft Read-Index protocol to ensure reads only execute after a quorum confirms the latest committed write index.

Dragonboat is a high-performance Go implementation of the Raft consensus algorithm designed for building strongly consistent distributed systems. Understanding how dragonboat linearizability works is essential for developers who need strict consistency guarantees when applications issue concurrent read and write operations across distributed nodes. This article examines the architectural mechanisms, source code implementation, and practical usage patterns that enable these guarantees.

Architectural Foundations of Dragonboat Linearizability

Dragonboat implements linearizability through two complementary mechanisms: at-most-once write semantics via client sessions, and quorum-confirmed reads via the Read-Index protocol.

At-Most-Once Writes via Client Sessions

Every write operation in Dragonboat is associated with a client session that provides unique request identification and duplicate detection. Located in client/session.go, the session structure tracks the highest request ID that has been successfully proposed for each client.

When a client calls SyncPropose, the proposal carries the session ID and a monotonically increasing request ID. The leader in internal/raft/raft.go checks this ID against the session's last applied request. If the ID is less than or equal to the last applied value, the request is a duplicate and is ignored, preventing double execution. This ensures that retries caused by network timeouts do not violate linearizability.

The Read-Index Protocol for Linearizable Reads

Linearizable reads in Dragonboat use the Read-Index protocol defined in the Raft thesis (section 6.4). When a read request arrives, the leader must confirm that it is still the leader and that the read will observe all writes committed before the read started.

The implementation in internal/raft/readindex.go manages read index tracking through the readIndex and readStatus structures. When NodeHost.linearizableRead is invoked (defined in nodehost.go), it issues a ReadIndex request to the Raft layer. The leader stores this request and broadcasts a heartbeat containing a hint (the SystemCtx) to all followers via broadcastHeartbeatMessageWithHint in raft.go.

Followers respond with ReadIndexResp messages. Once the leader receives a quorum of responses confirming they have applied entries up to the hinted index, the read is confirmed via handleLeaderReadIndex (lines 1841-1863 in raft.go). Only then does the user-provided read function execute against the state machine.

Deep Dive into the Read-Index Implementation

The Read-Index implementation spans multiple files and involves careful coordination between the NodeHost API and the core Raft logic.

In nodehost.go, the linearizableRead method serves as the entry point for all linearizable read operations. It first invokes nh.readIndex to initiate the protocol, then blocks until the index is confirmed. Once confirmed, it executes the user-provided function f(node) which receives a *dragonboat.Node providing access to the underlying state machine.

The internal/raft/readindex.go file contains the readIndex structure which maintains a queue of pending read requests. Each request is associated with a SystemCtx value that serves as the hint broadcast to followers. The addRequest method records the current commit index when the read is received, while confirm processes the quorum responses and marks requests as ready for execution.

In internal/raft/raft.go, the leader's heartbeat mechanism is modified to include Read-Index hints. The broadcastHeartbeatMessageWithHint function (lines 47-61) attaches the SystemCtx to heartbeat messages. When followers respond, the handleLeaderReadIndex function (lines 1841-1863) validates that a quorum has confirmed the read index before allowing the read to proceed.

Practical Implementation: Code Examples

The following examples demonstrate how to use Dragonboat's APIs to achieve linearizability in real applications.

Setting Up NodeHost and Client Sessions

Before issuing writes, establish a NodeHost and obtain a client session:

import (
    "context"
    "time"
    "github.com/lni/dragonboat/v4"
)

// Initialize NodeHost with configuration
nh, err := dragonboat.NewNodeHost(cfg)
if err != nil {
    panic(err)
}

// Obtain a client session for shard 1, replica 1
session, err := nh.GetClientSession(1, 1)
if err != nil {
    panic(err)
}
defer session.Close()

Executing Linearizable Writes with SyncPropose

Use SyncPropose with a session to ensure at-most-once execution:

ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()

cmd := []byte("set key=foo value=bar")
result, err := nh.SyncPropose(ctx, 1, session, cmd)
if err != nil {
    // Handle ErrInvalidSession, ErrTimeout, etc.
    panic(err)
}

// Wait for proposal to be committed and applied
<-result.Completed()

The session's RequestID prevents duplicate execution if network timeouts cause retries.

Performing Linearizable Reads with SyncRead

Use SyncRead to ensure the read observes all prior writes:

result, err := nh.SyncRead(ctx, 1, func(node *dragonboat.Node) (interface{}, error) {
    // Access the state machine directly
    // Example: return sm.Get("foo")
    return node.UserState().(YourStateMachine).Get("foo")
})
if err != nil {
    panic(err)
}
fmt.Printf("value=%v\n", result)

Under the hood, this invokes NodeHost.linearizableRead, which executes the Read-Index protocol before running your function.

Handling Concurrent Read and Write Operations

Dragonboat safely handles concurrent operations:

// Goroutine 1: Write operation
go func() {
    ctx, cancel := context.WithTimeout(context.Background(), 3*time.Second)
    defer cancel()
    nh.SyncPropose(ctx, 1, session, []byte("inc counter"))
}()

// Goroutine 2: Read operation that must see the increment
go func() {
    ctx, cancel := context.WithTimeout(context.Background(), 3*time.Second)
    defer cancel()
    v, _ := nh.SyncRead(ctx, 1, func(node *dragonboat.Node) (interface{}, error) {
        return sm.Get("counter")
    })
    fmt.Println("counter =", v)
}()

Because SyncRead waits for the Read-Index quorum confirmation, it observes either the pre-increment or post-increment state, but never an inconsistent intermediate state.

Key Source Files for Dragonboat Linearizability

Understanding the implementation requires examining these specific files:

  • client/session.go – Implements client sessions with request ID tracking for at-most-once write semantics.
  • internal/raft/readindex.go – Contains the readIndex and readStatus structures that manage pending read requests and quorum confirmation.
  • nodehost.go – Houses NodeHost.linearizableRead and the public SyncRead/SyncPropose APIs.
  • internal/raft/raft.go – Implements leader-side Read-Index handling via broadcastHeartbeatMessageWithHint (lines 47-61) and handleLeaderReadIndex (lines 1841-1863).

Summary

  • Dragonboat linearizability relies on two core mechanisms: client sessions for writes and the Read-Index protocol for reads.
  • Client sessions in client/session.go provide at-most-once semantics by tracking request IDs, ensuring retried proposals do not execute twice.
  • The Read-Index protocol requires the leader to confirm a quorum of followers have applied entries up to a specific index before executing read operations.
  • Public APIs SyncPropose and SyncRead hide the complexity of these protocols while delivering strict linearizable guarantees.
  • Concurrent operations are safe because reads wait for quorum confirmation, ensuring they observe all writes committed before the read started.

Frequently Asked Questions

What is the difference between a client session and a regular proposal?

A client session is a persistent context created via GetClientSession that tracks the highest request ID successfully proposed by that client. When you use SyncPropose with a session, the proposal includes this request ID, allowing the leader to detect and reject duplicates. Regular proposals without sessions lack this at-most-once guarantee and may execute multiple times if the client retries due to network timeouts.

How does Dragonboat handle read requests during leader elections?

During leader elections, the cluster may temporarily lack a valid leader. If a SyncRead request arrives at a node that is not the leader, or if the leader steps down during the read process, the request returns an error (typically ErrNotLeader or ErrTimeout). The Read-Index protocol requires a stable leader to obtain the quorum confirmation, so reads issued during leadership transitions will fail fast rather than return stale data.

Can I use SyncRead on follower nodes?

No, SyncRead must execute on the leader node. The Read-Index protocol requires the leader to broadcast heartbeat hints and collect quorum responses from followers. While followers can serve stale reads using StaleRead or similar non-linearizable methods, only the leader can guarantee that a read observes all committed writes. Attempting to invoke linearizable read logic on a follower will result in an error redirecting the client to the leader.

What happens if a Read-Index request times out?

If a Read-Index request times out before the leader collects a quorum of confirmations, the SyncRead call returns an error (typically ErrTimeout). This can occur during network partitions, high latency between nodes, or when the leader is partitioned from a majority of followers. The timeout prevents the system from blocking indefinitely and allows the application to implement retry logic with exponential backoff or fail over to alternative strategies.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →