Best Practices for Snapshotting and Log Compaction in Dragonboat: A Complete Guide
Dragonboat uses automatic snapshotting and log compaction to prevent unbounded storage growth while maintaining fast replication, configurable through SnapshotEntries, CompactionOverhead, and manual RequestSnapshot or RequestCompaction calls.
Dragonboat, the high-performance Raft consensus library written in Go (lni/dragonboat), relies on two core mechanisms for storage management: snapshotting and log compaction. Properly configuring these features ensures that your Raft log does not grow indefinitely while keeping recovery and replication costs low. This guide explains the internal architecture of these systems and provides production-ready best practices based on the actual source code implementation.
How Snapshotting Works in Dragonboat
Dragonboat creates snapshots to capture the current state of your state machine, allowing nodes to recover without replaying the entire Raft log.
Automatic and Manual Snapshot Triggers
Snapshots are triggered in two ways. First, automatically every N applied entries, controlled by the Config.SnapshotEntries field defined in config/config.go. Second, manually via the NodeHost.RequestSnapshot method in nodehost.go. When requesting a manual snapshot, you can pass a SnapshotOption struct, which is validated by SnapshotOption.Validate to ensure parameters like CompactionOverhead and CompactionIndex are legal (nodehost.go lines 21-35).
The Snapshot Generation Pipeline
The actual snapshot creation is handled by the snapshotter component in snapshotter.go. The process follows this pipeline:
- Write: State machine data is written to a temporary directory.
- Compress: The data is compressed based on
SnapshotCompressionTypesettings. - Commit: Metadata is committed to the LogDB via
snapshotter.Save.
The full implementation of the save pipeline is found in snapshotter.Save (snapshotter.go lines 104-146).
Metadata Registration
After a successful write, the snapshot is registered in the LogDB through logdb.SaveSnapshots. This metadata allows the system to track which snapshots are available for recovery and compaction.
How Log Compaction Works in Dragonboat
Log compaction removes obsolete Raft log entries that have been superseded by a snapshot, reclaiming disk space.
Automatic Compaction Triggers
When a snapshot is successfully persisted and Config.DisableAutoCompactions is false (the default), the node automatically calls node.requestCompaction() in node.go. This method obtains the compact-to index from the snapshotState and invokes LogDB.CompactEntriesTo (node.go lines 72-86).
The Compaction Flow
The compaction process involves two main operations:
- Log Entry Removal:
LogDB.RemoveEntriesTodeletes Raft log entries up to the compact-to index. - Snapshot Directory Cleanup:
snapshotter.Compactremoves stale snapshot directories older than the target index. This is implemented insnapshotter.Compact(snapshotter.go lines 221-236).
Manual Compaction Requests
If you disable automatic compactions to control I/O bursts, you must manually trigger compaction using NodeHost.RequestCompaction. This method forwards the request to the node's internal requestCompaction method (nodehost.go lines 79-94).
Configuration Knobs for Storage Management
The config.Config struct in config/config.go (lines 99-136) defines the primary fields for controlling snapshotting and log compaction:
| Field | Purpose | Typical Values |
|---|---|---|
SnapshotEntries |
Frequency of automatic snapshots (entries per snapshot) | 0 disables auto-snapshot; otherwise 50000 to 1000000 depending on workload |
CompactionOverhead |
Number of recent log entries kept after compaction | 0 compacts up to snapshot index; 500 keeps recent entries for fast follower catch-up |
DisableAutoCompactions |
Disables automatic compaction after snapshots | false for normal operation; true only for explicit I/O control |
SnapshotOption.OverrideCompactionOverhead |
Per-snapshot override of compaction behavior | Use for critical snapshots requiring custom history retention |
Snapshot State Tracking
The snapshotState struct in snapshotstate.go manages the indices required for compaction and ensures only one snapshot operation runs concurrently. Key fields include compactLogTo and compactedTo, which track the target and completed compaction indices respectively (snapshotstate.go lines 65-70).
Access to these fields is synchronized via atomic helpers such as hasCompactLogTo, getCompactLogTo, and setCompactedTo (snapshotstate.go lines 133-151).
Best Practices for Production Workloads
Follow these guidelines to optimize storage management in Dragonboat:
-
Set a sensible
SnapshotEntriesvalue – Choose a value that balances snapshot size versus frequency. Very large numbers cause the Raft log to grow unbounded; very small numbers increase snapshot I/O overhead. -
Tune
CompactionOverhead– Keep a modest overhead (e.g., 200–1000 entries) so followers can catch up by log replication rather than downloading an entire snapshot, reducing network traffic. -
Leave auto-compaction enabled unless necessary – The node automatically compacts the log right after a snapshot, reclaiming disk space promptly. Disabling it adds management burden.
-
When disabling auto-compaction, always call
RequestCompaction– Otherwise the log will never be trimmed and storage will explode. -
Use
SnapshotOption.OverrideCompactionOverheadfor outlier snapshots – For critical snapshots you may want to keep more recent entries (or none) – supply a customSnapshotOptiontoRequestSnapshot. -
Monitor snapshot and compaction events – Dragonboat emits
SnapshotCompacted,LogCompacted, andLogDBCompactedevents (seeraftio.Listener). Hook into these for observability or alerting. -
Periodically clean up old snapshots – Even after log compaction, old snapshot directories remain. Use
snapshotter.Remove(called internally) or invokesnapshotter.Compactwith the newest index to purge older snapshots. -
Avoid overlapping snapshot requests – The
snapshotStatelogic ensures only one snapshot is in flight; if you receiveErrSystemBusyfromRequestSnapshot, back-off and retry. -
Test with realistic payload sizes – Compression (
SnapshotCompressionType) can affect snapshot size and compaction speed; experiment withSnappyversusNoCompressionin a staging environment.
Code Examples
Configuring NodeHost for Balanced Storage
cfg := config.NodeHostConfig{
// … other required fields …
RaftAddress: "127.0.0.1:63001",
LogDBFactory: logdb.New,
SnapshotEntries: 50000, // create a snapshot every 50 k entries
CompactionOverhead: 500, // keep the last 500 entries after compaction
// DisableAutoCompactions: false // default – keep auto‑compaction
}
nh, err := dragonboat.NewNodeHost(cfg)
if err != nil { /* handle error */ }
Links: SnapshotEntries and CompactionOverhead fields are defined in config/config.go (lines 99-136).
Requesting Manual Snapshots with Custom Compaction Overhead
opt := dragonboat.SnapshotOption{
OverrideCompactionOverhead: true,
CompactionOverhead: 0, // compact completely up to snapshot index
}
rs, err := nh.RequestSnapshot(shardID, opt, 5*time.Second)
if err != nil { /* handle error */ }
// wait for completion
result := <-rs.ResultC()
fmt.Printf("snapshot created at index %d\n", result.Value)
Link: SnapshotOption.Validate ensures the override is legal (nodehost.go lines 21-35).
Triggering Manual Log Compaction
// Assuming you previously disabled auto compaction in the config:
_, err = nh.RequestCompaction(shardID, replicaID)
if err != nil && err != dragonboat.ErrRejected {
// ErrRejected means there was nothing to compact
log.Fatalf("compaction failed: %v", err)
}
Link: NodeHost.RequestCompaction forwards the request to the node's internal requestCompaction method (nodehost.go lines 79-94).
Observing Compaction Events
type myListener struct{ dragonboat.SysEventListener }
func (l *myListener) LogCompacted(info raftio.EntryInfo) {
fmt.Printf("log compacted up to %d\n", info.Index)
}
nh.AddEventListener(&myListener{})
Link: The listener interface is defined in raftio/listener.go (listener definitions).
Key Source Files Reference
| File | Role | Direct Link |
|---|---|---|
config/config.go |
Holds global configuration (SnapshotEntries, CompactionOverhead, DisableAutoCompactions) |
https://github.com/lni/dragonboat/blob/master/config/config.go |
snapshotstate.go |
Tracks snapshot indices, compaction requests, and flags (compactLogTo, compactedTo) |
https://github.com/lni/dragonboat/blob/master/snapshotstate.go |
snapshotter.go |
Implements snapshot creation, saving, loading, and compacting of old snapshot files | https://github.com/lni/dragonboat/blob/master/snapshotter.go |
node.go |
Core Raft node logic – applies updates, triggers snapshotting, and performs log compaction (requestCompaction) |
https://github.com/lni/dragonboat/blob/master/node.go |
nodehost.go |
Public API for snapshot & compaction requests (RequestSnapshot, RequestCompaction) |
https://github.com/lni/dragonboat/blob/master/nodehost.go |
raftio/logdb.go |
Interface to the underlying log storage; defines CompactEntriesTo and RemoveEntriesTo |
https://github.com/lni/dragonboat/blob/master/raftio/logdb.go |
raftio/listener.go |
System‑event listener hooks for snapshot/compaction notifications | https://github.com/lni/dragonboat/blob/master/raftio/listener.go |
Summary
- Snapshotting and log compaction in Dragonboat work together to prevent unbounded storage growth while maintaining replication efficiency.
- Configure
SnapshotEntriesto balance snapshot frequency against log growth, typically between 50,000 and 1,000,000 entries depending on workload. - Set
CompactionOverheadto retain recent entries (e.g., 500-1000) so followers can catch up via log replication rather than full snapshot transfers. - Keep
DisableAutoCompactionsset tofalseunless you need explicit I/O control, in which case you must manually invokeRequestCompaction. - Use
SnapshotOption.OverrideCompactionOverheadfor specific snapshots that require custom retention policies. - Monitor system health by implementing the
raftio.Listenerinterface to receiveLogCompactedandSnapshotCompactedevents.
Frequently Asked Questions
What is the difference between snapshotting and log compaction in Dragonboat?
Snapshotting captures the current state of your state machine and saves it to disk via the snapshotter.Save method in snapshotter.go, creating a point-in-time image that can be used for recovery. Log compaction removes obsolete Raft log entries that are already covered by a snapshot through LogDB.RemoveEntriesTo and snapshotter.Compact, reclaiming disk space. While snapshotting creates recovery points, compaction cleans up the log entries that preceded those points.
How often should I configure SnapshotEntries for optimal performance?
The ideal value for SnapshotEntries in config/config.go depends on your write throughput and state machine size. For most workloads, values between 50,000 and 1,000,000 entries work well. Setting it too high (e.g., millions) causes the Raft log to grow large, increasing recovery time and disk usage. Setting it too low (e.g., thousands) generates excessive I/O from frequent snapshot creation. Test with realistic payload sizes using Snappy or NoCompression to find the balance for your specific hardware.
Can I disable automatic log compaction in Dragonboat?
Yes, you can disable automatic compaction by setting DisableAutoCompactions to true in your config.Config. This is useful when you need explicit control over I/O bursts or want to coordinate compaction with maintenance windows. However, you must manually call NodeHost.RequestCompaction periodically if you disable auto-compaction; otherwise, the log will never be trimmed and storage will grow unbounded. Always check for dragonboat.ErrRejected when calling this method, as it indicates there was nothing to compact.
How do I monitor snapshot and compaction activity in production?
Dragonboat emits system events through the raftio.Listener interface defined in raftio/listener.go. Implement this interface to receive notifications when compaction occurs:
LogCompactedfires when entries are removed from the log.SnapshotCompactedfires when old snapshot files are cleaned up.LogDBCompactedprovides low-level storage compaction events.
Attach your listener using NodeHost.AddEventListener(). Additionally, monitor for ErrSystemBusy when calling RequestSnapshot, as the snapshotState logic in snapshotstate.go ensures only one snapshot runs at a time, and busy errors indicate you need to back off and retry.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →