Twitter's Snowflake Unique ID Generator: 64-Bit Structure and Implementation

Twitter's Snowflake unique ID generator creates 64-bit integers that are globally unique and roughly time-ordered by partitioning bits into timestamp, machine identifiers, and a sequence counter, eliminating the need for distributed coordination.

Twitter developed the Snowflake algorithm to replace auto-incrementing database IDs with a scalable, distributed solution. According to the liquidslr/system-design-notes repository, the implementation documented in 07. Unique-Id Generator/Readme.md creates identifiers that sort chronologically while supporting thousands of nodes across multiple datacenters without single points of failure.

64-Bit ID Structure

Twitter's Snowflake allocates 64 bits into five distinct fields to ensure uniqueness and ordering:

Bits Field Purpose
1 Sign bit Always 0 to keep the ID positive as a signed 64-bit integer
41 Timestamp Milliseconds since custom epoch (1288834974657, November 4 2010)
5 Datacenter ID Identifies up to 32 distinct datacenters
5 Machine ID Identifies up to 32 machines per datacenter
12 Sequence Counts up to 4096 IDs per millisecond per machine

This bit allocation, detailed in liquidslr/system-design-notes at lines 74-80 of the Unique-Id Generator documentation, guarantees that the most significant bits contain temporal information, making IDs naturally sortable by creation time.

Timestamp and Epoch Configuration

The 41-bit timestamp field stores the elapsed milliseconds since Snowflake's custom epoch (1288834974657). By subtracting this fixed epoch from the current Unix time and shifting the result left by 22 bits, the algorithm positions temporal data as the most significant component of the ID.

This configuration supports:

  • 69 years of unique timestamps from the custom epoch
  • Chronological sorting without additional indexes
  • Roughly time-ordered generation across distributed systems

Datacenter and Machine Identification

Snowflake embeds 10 bits of location metadata directly into each ID:

  • Datacenter ID (5 bits): Supports 32 physical or logical datacenters
  • Machine ID (5 bits): Supports 32 workers per datacenter

These static identifiers are configured at process startup and shifted into their respective bit positions (datacenter shifted left by 17 bits, machine shifted left by 12 bits). This design eliminates the need for central coordination because each worker operates within a unique namespace defined by its datacenterId and machineId combination.

Sequence Counter and Overflow Handling

The 12-bit sequence number resolves collisions when multiple IDs generate within the same millisecond on the same machine. The sequence increments for each call within the same timestamp and resets to 0 when the millisecond changes.

Critical implementation details from the source code:

  • Maximum 4096 unique IDs per millisecond per machine
  • When the sequence exceeds 4095, the generator blocks until the next millisecond
  • The sequence occupies the least significant 12 bits (positions 0-11)

JavaScript Implementation Example

The following implementation demonstrates the complete bit-shifting logic described in the liquidslr/system-design-notes repository:

class Snowflake {
  constructor(datacenterId, machineId) {
    this.EPOCH = 1288834974657n;
    this.datacenterId = BigInt(datacenterId) & 0x1Fn;  // 5 bits
    this.machineId = BigInt(machineId) & 0x1Fn;      // 5 bits
    this.sequence = 0n;
    this.lastTimestamp = -1n;
  }

  _timestamp() {
    return BigInt(Date.now());
  }

  nextId() {
    let timestamp = this._timestamp();

    if (timestamp === this.lastTimestamp) {
      this.sequence = (this.sequence + 1n) & 0xFFFn; // 12 bits
      if (this.sequence === 0n) {
        // Wait until next millisecond
        while (timestamp <= this.lastTimestamp) {
          timestamp = this._timestamp();
        }
      }
    } else {
      this.sequence = 0n;
    }

    this.lastTimestamp = timestamp;

    // Compose the 64-bit ID
    return (
      ((timestamp - this.EPOCH) << 22n) |
      (this.datacenterId << 17n) |
      (this.machineId << 12n) |
      this.sequence
    );
  }
}

This implementation follows the exact bit layout: 22 bits of left shift for the timestamp (accommodating the 10-bit machine identifier and 12-bit sequence), 17 bits for the datacenter, and 12 bits for the machine ID.

Scalability and Performance Characteristics

The Snowflake architecture enables massive horizontal scaling:

  • 409,600 IDs per second per machine (4096/ms × 1000ms/1s)
  • 1,024 machines total capacity (32 datacenters × 32 machines)
  • No single point of failure because each worker generates IDs independently

Clock Synchronization Requirements

All participating machines must maintain synchronized clocks, typically via NTP. Clock drift can cause:

  • ID collisions if a machine's clock moves backward (mitigated by blocking or throwing exceptions)
  • Out-of-order IDs if clocks drift across datacenters

Summary

  • Twitter's Snowflake unique ID generator creates 64-bit integers combining timestamp, datacenter, machine, and sequence components.
  • The 41-bit timestamp field ensures roughly time-ordered generation for 69 years from the November 2010 epoch.
  • 5-bit datacenter and 5-bit machine identifiers support 1,024 concurrent workers without coordination.
  • The 12-bit sequence allows 4,096 IDs per millisecond per machine with automatic rollover handling.
  • Implementation requires BigInt operations in JavaScript to handle 64-bit arithmetic correctly.

Frequently Asked Questions

What is the maximum ID generation rate per machine in Twitter's Snowflake?

Each machine can generate 4,096 unique IDs per millisecond, translating to over 4 million IDs per second under continuous load. When this limit is reached, the generator blocks execution until the next millisecond tick, ensuring no collisions occur.

How does Snowflake handle clock skew or backward time jumps?

The algorithm tracks the lastTimestamp and blocks execution if the current time is less than the previous timestamp. This prevents duplicate IDs at the cost of availability during clock corrections. Production deployments require strict NTP synchronization to minimize such events.

Can I modify the bit allocation in Snowflake for my use case?

Yes, the bit structure is configurable depending on your infrastructure needs. You might allocate more bits to the sequence (for higher per-machine throughput) or more bits to the machine ID (for more nodes), though this reduces the timestamp range or other fields accordingly. The 64-bit total must remain constant for compatibility with standard integer types.

Why does Snowflake use a custom epoch instead of Unix epoch?

Twitter selected November 4, 2010 (1288834974657) as the epoch to maximize the 69-year lifespan of the 41-bit timestamp field relative to their production deployment date. Using the standard Unix epoch (1970) would waste approximately 40 years of timestamp range before the system even launched.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →