Dragonboat SnapshotEntries vs CompactionOverhead: Configuration and Behavior Differences

SnapshotEntries determines the frequency of automatic snapshots by counting applied Raft entries, while CompactionOverhead specifies how many recent log entries to preserve after compaction to support lagging followers.

In the lni/dragonboat Raft library, managing the trade-off between snapshot frequency and log retention is essential for cluster performance. The SnapshotEntries and CompactionOverhead configuration fields in config/config.go control distinct phases of log management—triggering snapshots versus preserving log history. Understanding how these settings interact helps prevent excessive snapshotting while ensuring nodes can recover without full state transfers.

What SnapshotEntries Configures

SnapshotEntries controls when automatic snapshots are triggered. Defined in config/config.go (lines 99-101), this field specifies the number of Raft log entries that must be applied to the state machine before Dragonboat automatically initiates a snapshot.

When Config.SnapshotEntries is set to a value greater than 0, the Raft node periodically checks the applied entry count. Once the threshold is reached, the system triggers a snapshot automatically. Setting this value to 0 (the default) disables automatic snapshotting, requiring manual intervention via NodeHost.RequestSnapshot or NodeHost.SyncRequestSnapshot.

Typical use cases include reducing snapshot frequency for workloads with high volumes of small proposals, or ensuring regular snapshots occur in write-heavy environments to prevent unbounded log growth.

What CompactionOverhead Configures

CompactionOverhead determines how much of the Raft log is preserved after a snapshot triggers log compaction. Defined in config/config.go (lines 120-124), this setting specifies the number of most recent log entries to retain after compaction, ensuring lagging followers can catch up using log entries rather than requiring a full snapshot transfer.

When compaction runs—either automatically following a snapshot or manually via RequestSnapshot—Dragonboat calculates the cutoff point as lastSnapshotIndex - CompactionOverhead. Entries older than this index are physically removed, while the specified overhead count is preserved.

The default value of 0 removes all entries up to the snapshot index, which minimizes disk usage but forces lagging nodes to receive full snapshots. Setting a non-zero overhead creates a "log tail" that facilitates faster recovery for followers that are slightly behind, and supports debugging or replay of recent history.

Key Differences and Interaction

While both settings relate to snapshot management, they control distinct phases with different override capabilities.

SnapshotEntries governs the trigger condition for automatic snapshots. It is a static configuration that cannot be overridden per request. Once set, the node monitors applied entry counts and initiates snapshots automatically when the threshold is reached.

CompactionOverhead governs the retention policy applied after snapshot completion. While it has a default node-wide value in Config, it can be overridden for individual snapshot requests using SnapshotOption.OverrideCompactionOverhead and SnapshotOption.CompactionOverhead (defined in nodehost.go, lines 98-104).

The interaction between these settings follows a sequential process: first, SnapshotEntries triggers the snapshot creation; second, after the snapshot persists, the compaction process applies the effective overhead (either from config or the override) to determine which log entries to preserve.

Configuration Examples

Automatic Snapshots with Log Retention

Configure automatic snapshots every 10,000 entries while preserving the last 500 entries for follower catch-up:

cfg := config.Config{
    // Trigger snapshot after 10,000 applied entries
    SnapshotEntries: 10000,
    // Preserve 500 recent entries after compaction
    CompactionOverhead: 500,
}

With this configuration, Dragonboat automatically snapshots after approximately 10,000 proposals and then compacts the log, leaving the last 500 entries on disk.

Overriding Compaction Overhead for a Single Snapshot

Temporarily increase log retention for a specific manual snapshot without changing the node configuration:

opt := nodehost.SnapshotOption{
    OverrideCompactionOverhead: true,
    CompactionOverhead:         1000,
}
err := nh.RequestSnapshot(shardID, replicaID, opt)

This request preserves the latest 1,000 entries regardless of the Config.CompactionOverhead setting, which remains unchanged for subsequent automatic snapshots.

Disabling Automatic Snapshots

For complete manual control over when snapshots occur:

cfg := config.Config{
    SnapshotEntries: 0, // Disable automatic snapshotting
}

With this configuration, the log grows until you explicitly call RequestSnapshot or SyncRequestSnapshot.

Summary

  • SnapshotEntries controls the frequency of automatic snapshots by counting applied Raft log entries. Set it to 0 to disable automatic snapshots and trigger them manually via RequestSnapshot.

  • CompactionOverhead determines how many recent log entries are preserved after snapshot compaction. It ensures lagging followers can catch up without full state transfers and can be overridden per-request using SnapshotOption.

  • These settings operate sequentially: SnapshotEntries triggers snapshot creation, then compaction applies the effective overhead (from Config or override) to retain a tail of recent entries.

  • Source locations: Config.SnapshotEntries and Config.CompactionOverhead are defined in config/config.go, while per-request override capabilities reside in SnapshotOption within nodehost.go.

Frequently Asked Questions

What happens if I set SnapshotEntries to 0?

When SnapshotEntries is set to 0 (the default), automatic snapshotting is completely disabled. The Raft log will grow indefinitely until you manually trigger a snapshot using NodeHost.RequestSnapshot or NodeHost.SyncRequestSnapshot. This configuration is useful when you need precise control over when snapshots occur, such as during maintenance windows or after specific business logic milestones.

Can I change CompactionOverhead without restarting the node?

No, the node-wide Config.CompactionOverhead cannot be changed dynamically without restarting the node. However, you can temporarily override the compaction behavior for a single snapshot request by setting SnapshotOption.OverrideCompactionOverhead to true and specifying a custom SnapshotOption.CompactionOverhead value. This override affects only that specific snapshot operation and does not modify the permanent node configuration.

How do these settings affect cluster recovery time?

SnapshotEntries indirectly affects recovery time by controlling snapshot freshness. Lower values create snapshots more frequently, reducing the number of log entries a crashed node must replay, but increasing I/O overhead. CompactionOverhead directly impacts recovery by determining whether lagging followers can catch up via incremental log entries or must receive a full snapshot. A higher overhead preserves more recent entries, enabling faster catch-up for slightly behind nodes, but increases disk usage for the log.

While the settings operate independently, practical deployments often consider their interaction. If SnapshotEntries is set to 10,000 and CompactionOverhead is set to 500, the system snapshots every 10,000 entries but keeps the last 500 of those in the log. Ensure CompactionOverhead is significantly smaller than SnapshotEntries to avoid negating the disk space benefits of snapshotting. There is no enforced relationship in the code, but maintaining a ratio where overhead is 5-10% of the snapshot frequency is a common operational guideline.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →