IStateMachine vs IOnDiskStateMachine in Dragonboat: Functional Differences Explained

IStateMachine maintains state purely in memory and rebuilds from snapshots on restart, while IOnDiskStateMachine persists state to disk with an explicit Open hook, batch updates, and finer concurrency control.

Dragonboat is a high-performance Raft consensus library for Go that provides two distinct state machine interfaces for managing replicated state. Understanding the functional differences between IStateMachine and IOnDiskStateMachine is critical for choosing the right persistence model for your application, whether you need fast in-memory counters or durable disk-backed storage.

Core Architectural Differences

State Persistence Model

The fundamental distinction lies in where the application state lives:

  • IStateMachine: State is kept primarily in-memory. All user data resides in RAM and is reconstructed from snapshots and Raft logs after a restart. This model is defined in statemachine/rsm.go.
  • IOnDiskStateMachine: State is persisted on-disk. The state is always written to files or a database, and the machine starts by opening existing data on disk rather than rebuilding from scratch. This interface is defined in statemachine/disk.go.

Lifecycle and Initialization

IStateMachine implementations start empty when the node boots. Dragonboat automatically rebuilds the state by replaying the latest snapshot and subsequent log entries. There is no explicit initialization hook.

IOnDiskStateMachine requires an explicit Open(stopc <-chan struct{}) (uint64, error) method. As implemented in statemachine/disk.go, this hook is called immediately after the Raft node starts. It must load persisted data (or create a new empty state) and return the last applied Raft index. This allows the node to resume from exactly where it left off without replaying the entire log.

Update Method Signatures

The update methods differ in how they process Raft log entries:

  • IStateMachine: Update(entry Entry) (Result, error) processes a single Raft log entry at a time. This is the classic approach described in the original Raft thesis.
  • IOnDiskStateMachine: Update([]Entry) ([]Entry, error) receives a batch of consecutive entries. As noted in the source analysis, this allows implementations to batch-write to disk, significantly improving throughput for durable storage engines.

Concurrency and Snapshot Behavior

Concurrency Guarantees

Dragonboat employs different locking strategies for each interface:

  • IStateMachine: A global sync.RWMutex guards the machine. Update, RecoverFromSnapshot, and Close acquire the write lock, while Lookup and SaveSnapshot acquire the read lock. This ensures deterministic replay but limits concurrency.
  • IOnDiskStateMachine: Concurrency is more fine-grained. According to the implementation in engine.go and statemachine/disk.go, Update may run concurrently with Lookup and SaveSnapshot. The system guarantees mutual exclusion only for methods that require it (Sync, PrepareSnapshot, RecoverFromSnapshot, Close).

Snapshot Mechanisms

The snapshot workflows differ significantly based on where state lives:

IStateMachine implementations must serialize all in-memory data during SaveSnapshot(io.Writer, …). The entire state is streamed to the writer, which can be expensive for large datasets.

IOnDiskStateMachine uses a two-phase approach defined in statemachine/disk.go:

  1. PrepareSnapshot() creates a small identifier (e.g., a version number or timestamp) representing the current disk state.
  2. SaveSnapshot only writes this identifier (often just a few kilobytes) because the bulk of the state already exists on disk.
  3. External files can be added to the snapshot through ISnapshotFileCollection.

This design enables high-throughput workloads that would otherwise be limited by RAM size or snapshot overhead.

Sync Operations

IStateMachine has no explicit sync operation—the in-memory state is always up-to-date after each Update returns.

IOnDiskStateMachine provides a Sync() method to explicitly flush any in-core buffers to durable storage. As implemented in typical disk-backed stores, this gives users control over when data is persisted, allowing for fsync policies that balance durability against performance.

Code Implementation Examples

In-Memory Counter (IStateMachine)

The following example from statemachine/rsm.go demonstrates a simple counter that implements IStateMachine:

type CounterSM struct {
    count uint64
}

// Ensure CounterSM implements IStateMachine
var _ statemachine.IStateMachine = (*CounterSM)(nil)

func (c *CounterSM) Update(entry statemachine.Entry) (statemachine.Result, error) {
    // entry.Cmd is expected to be a uint64 delta in little-endian format
    delta := binary.LittleEndian.Uint64(entry.Cmd)
    c.count += delta
    return statemachine.Result{Value: c.count}, nil
}

func (c *CounterSM) Lookup(arg interface{}) (interface{}, error) {
    // arg is ignored; just return the current count
    return c.count, nil
}

// Snapshot the count as an uint64
func (c *CounterSM) SaveSnapshot(w io.Writer,
    _ statemachine.ISnapshotFileCollection, _ <-chan struct{}) error {
    var buf [8]byte
    binary.LittleEndian.PutUint64(buf[:], c.count)
    _, err := w.Write(buf[:])
    return err
}

func (c *CounterSM) RecoverFromSnapshot(r io.Reader,
    _ []statemachine.SnapshotFile, _ <-chan struct{}) error {
    var buf [8]byte
    if _, err := io.ReadFull(r, buf[:]); err != nil {
        return err
    }
    c.count = binary.LittleEndian.Uint64(buf[:])
    return nil
}

func (c *CounterSM) Close() error { return nil }

Disk-Backed KV Store (IOnDiskStateMachine)

This example from statemachine/disk.go illustrates a BoltDB-backed key-value store implementing IOnDiskStateMachine:

type DiskKV struct {
    db *bolt.DB
}

// Ensure DiskKV implements IOnDiskStateMachine
var _ statemachine.IOnDiskStateMachine = (*DiskKV)(nil)

func (kv *DiskKV) Open(stopc <-chan struct{}) (uint64, error) {
    var err error
    kv.db, err = bolt.Open("kv.db", 0600, nil)
    if err != nil {
        return 0, err
    }
    // Retrieve last applied index from a meta bucket (if any)
    var lastIdx uint64
    kv.db.View(func(tx *bolt.Tx) error {
        b := tx.Bucket([]byte("meta"))
        if b != nil {
            lastIdx = binary.BigEndian.Uint64(b.Get([]byte("lastIdx")))
        }
        return nil
    })
    return lastIdx, nil
}

func (kv *DiskKV) Update(ents []statemachine.Entry) ([]statemachine.Entry, error) {
    for i, e := range ents {
        // e.Cmd expected to be: keyLen(4) | key | valueLen(4) | value
        // (omitting parsing for brevity)
        kv.db.Update(func(tx *bolt.Tx) error {
            b, _ := tx.CreateBucketIfNotExists([]byte("data"))
            // store value under key
            // b.Put(key, value)
            return nil
        })
        // Record the index in result
        ents[i].Result = statemachine.Result{Value: e.Index}
    }
    return ents, nil
}

func (kv *DiskKV) Lookup(arg interface{}) (interface{}, error) {
    key := arg.([]byte)
    var val []byte
    kv.db.View(func(tx *bolt.Tx) error {
        b := tx.Bucket([]byte("data"))
        if b != nil {
            v := b.Get(key)
            if v != nil {
                val = append([]byte{}, v...)
            }
        }
        return nil
    })
    if val == nil {
        return nil, errors.New("not found")
    }
    return val, nil
}

func (kv *DiskKV) Sync() error {
    return kv.db.Sync()
}

func (kv *DiskKV) PrepareSnapshot() (interface{}, error) {
    // Return snapshot identifier (e.g., a timestamp)
    return time.Now().UnixNano(), nil
}

func (kv *DiskKV) SaveSnapshot(stateID interface{}, w io.Writer,
    stopc <-chan struct{}) error {
    // Write a simple marker; real implementation would copy DB files.
    _, err := fmt.Fprintf(w, "snapshot-%d", stateID.(int64))
    return err
}

func (kv *DiskKV) RecoverFromSnapshot(r io.Reader, stopc <-chan struct{}) error {
    // In practice you would replace the DB file with the snapshot data.
    // Here we just read the marker.
    _, err := io.ReadAll(r)
    return err
}

func (kv *DiskKV) Close() error {
    return kv.db.Close()
}

When to Use Each Interface

Choose IStateMachine when your application requires fast, purely in-memory data structures such as counters, maps, or caches. This interface is ideal when datasets fit comfortably in RAM and the cost of rebuilding from a snapshot is acceptable. According to the implementation in statemachine/rsm.go, this is the classic Raft state machine model described in the original Raft thesis.

Choose IOnDiskStateMachine when dealing with large datasets that exceed available memory, or when integrating with external databases like RocksDB or LevelDB. This interface, defined in statemachine/disk.go, follows the "on-disk state machine" model (section 5.2 of the Raft thesis) and enables high-throughput workloads through batch processing and fine-grained concurrency control.

Summary

  • IStateMachine keeps all data in RAM and rebuilds from snapshots after restarts, using single-entry updates and coarse-grained locking via sync.RWMutex.
  • IOnDiskStateMachine persists state to disk, implements an Open hook to resume from existing data, processes batched entries via Update([]Entry), and provides fine-grained concurrency allowing reads during updates.
  • Snapshots for in-memory machines require serializing all data, while on-disk machines only snapshot lightweight identifiers since the bulk data already exists on disk.
  • Error handling differs through constants like ErrOpenStopped and ErrSnapshotStopped in statemachine/disk.go, allowing disk-based implementations to signal specific failure modes.

Frequently Asked Questions

What happens if an IOnDiskStateMachine Open method returns an error?

If the Open method returns an error, the Raft node will fail to start and report the error through the NodeHost. According to the implementation in nodehost.go, Dragonboat calls Open immediately after the node starts, and any failure to load persisted data or create the initial state prevents the node from joining the cluster. You can also return ErrOpenStopped if the operation is cancelled via the stop channel.

Can I use IStateMachine for large datasets that don't fit in memory?

No, IStateMachine is designed specifically for in-memory workloads. As implemented in statemachine/rsm.go, the interface assumes all state can be held in RAM and reconstructed from snapshots. If your dataset exceeds available memory or requires durability across restarts without full replay, you must implement IOnDiskStateMachine instead, which is optimized for large on-disk datasets through batching and efficient snapshotting.

How does snapshot recovery differ between the two interfaces?

For IStateMachine, RecoverFromSnapshot must deserialize the entire state from the snapshot reader and repopulate the in-memory structures, as shown in statemachine/rsm.go. For IOnDiskStateMachine, RecoverFromSnapshot typically replaces on-disk files with the snapshot data or applies incremental changes, since the bulk state is already persisted. The on-disk version also uses PrepareSnapshot to create lightweight identifiers rather than serializing all data, significantly reducing snapshot overhead for large datasets.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →