# How Chat Systems Handle Message Synchronization Across Multiple Devices: Architecture and Implementation

> Learn how chat systems synchronize messages across devices using push, storage, and conflict-free replication with append-only logs and sequence numbers. Understand the architecture and implementation.

- Repository: [Gaurav Kumar/system-design-notes](https://github.com/liquidslr/system-design-notes)
- Tags: deep-dive
- Published: 2026-09-09

---

**A modern chat system guarantees conversation consistency across smartphones, tablets, and desktops by combining real-time push, persistent storage, and conflict-free replication through an append-only log with monotonically increasing sequence numbers.**

Synchronizing messages across multiple devices requires more than simple database replication; it demands a distributed architecture that maintains total order despite network partitions and intermittent connectivity. According to the `liquidslr/system-design-notes` repository—which details these patterns in `12. Chat System/Readme.md` and the broader `03. System Design Framework/Readme.md`—production-grade chat services achieve **eventual consistency** by treating each message as an immutable event in a totally ordered stream.

## The Core Synchronization Pipeline

### Message Ingestion via Append-Only Logs

When a client sends a message, a stateless API gateway writes the payload to an **append-only log** (such as Kafka or Pulsar) and immediately returns an acknowledgment. As outlined in the system design framework, this approach ensures durability before downstream processing begins. The log assigns a **monotonically increasing sequence number** or ULID to each record, establishing a total order that prevents race conditions during replication.

### The Persistent Store as Source of Truth

A durable datastore—such as Cassandra, DynamoDB, or a relational database—persists the message alongside its sequence number, sender metadata, and conversation ID. This storage layer serves as the authoritative source for historical data when devices reconnect or new devices authenticate. The architecture in `12. Chat System/Readme.md` emphasizes that this store must support efficient range queries by sequence number to enable fast catch-up synchronization.

## Real-Time Delivery and Offline Synchronization

### WebSocket Push for Active Sessions

Online clients maintain persistent connections via WebSocket or gRPC streams. A **message broker** pushes new messages immediately to these connections, allowing devices to append events to their local view instantly. The client tracks the highest received sequence number locally, creating a checkpoint for incremental sync.

### Catching Up After Disconnection

When a device comes online or a new device logs in, it queries the persistent store for all messages **after the last sequence number** stored locally. The server returns the missing slice ordered by sequence, enabling the client to fast-forward its UI without reloading the entire conversation history. This mechanism, detailed in Chapter 12, minimizes bandwidth while ensuring that temporary offline periods do not create data gaps.

## Conflict-Free Replication and Consistency

### Total Ordering via Sequence Numbers

By using immutable message objects and a total order defined by sequence numbers, the system eliminates merge conflicts. Every device applies updates in the same sequence, ensuring that whether a user reads messages on a phone or laptop, the conversation state converges identically. The `03. System Design Framework/Readme.md` identifies this pattern as **conflict-free replicated data type (CRDT)** behavior achieved through linearizable logging.

### Handling Edits and Deletions with Tombstones

When a user edits or deletes a message, the system writes a new **tombstone event** with a higher sequence number rather than modifying the original record. All devices process this tombstone in order, updating their local views without ambiguity. This append-only approach to mutations ensures that offline devices eventually receive the edit when they synchronize, preventing "ghost" messages from reappearing.

## Scalability and Partitioning Strategy

### Sharding by Conversation ID

To handle millions of concurrent chats, the log shards by **conversation-ID**, distributing load across multiple partitions. This design enables horizontal scaling of both ingestion and delivery pipelines while maintaining ordering guarantees within each conversation. The repository illustrates this partitioning strategy in `12. Chat System/images/zookeeper.png`, showing how metadata coordination supports distributed message streaming.

### Coordination with ZooKeeper

A **coordinator service** (such as ZooKeeper) maintains metadata about partition leaders and handles leader election. As implemented in the reference architecture, this ensures high availability and fault tolerance without compromising the sequencing guarantees required for message synchronization across devices.

## Implementation Examples

The following pseudocode from `liquidslr/system-design-notes` demonstrates the core synchronization primitives: message ingestion, persistence, and device catch-up.

```python

# Message ingestion via API gateway

def send_message(user_id, chat_id, content):
    msg_id = uuid4()
    seq = kafka_producer.append(
        topic=chat_id,
        key=msg_id,
        value={
            "sender": user_id,
            "content": content,
            "timestamp": time.time(),
        },
    )
    return {"msg_id": msg_id, "seq": seq}

```

```python

# Persistence and real-time push

def consume_and_store(record):
    db.insert(
        table="messages",
        values={
            "msg_id": record.key,
            "chat_id": record.topic,
            "seq": record.offset,
            "sender": record.value["sender"],
            "content": record.value["content"],
            "ts": record.value["timestamp"],
        },
    )
    push_to_online_clients(chat_id=record.topic, payload=record.value)

```

```python

# Offline device synchronization

def sync_missing_messages(chat_id, last_seq):
    rows = db.query(
        f"SELECT * FROM messages WHERE chat_id=%s AND seq > %s ORDER BY seq ASC",
        (chat_id, last_seq),
    )
    return rows  # Client appends these to UI in order

```

```javascript
// Client-side real-time handling
socket.on('message', (msg) => {
  if (msg.seq > localLastSeq) {
    renderMessage(msg);
    localLastSeq = msg.seq;
  }
});

```

## Summary

- **Append-only logs** provide durability and total ordering through monotonically increasing sequence numbers.
- **Persistent storage** acts as the source of truth for offline devices and new device initialization.
- **WebSocket push** delivers real-time updates to online clients while sequence tracking enables efficient catch-up.
- **Tombstone events** handle edits and deletions without creating merge conflicts across distributed clients.
- **Conversation-ID sharding** and **ZooKeeper coordination** enable horizontal scaling while maintaining strict consistency guarantees.

## Frequently Asked Questions

### What happens if two devices send messages simultaneously?

The append-only log serializes all messages, assigning each a unique sequence number. Even if two devices send messages at the same millisecond, the log establishes a total order. Both devices receive both messages in the same sequence when they synchronize, ensuring identical conversation states regardless of network latency.

### How does the system handle message edits across devices?

Instead of updating existing records in place, the system appends a **tombstone event** with a higher sequence number that references the original message. All devices process this new event in order and update their local views accordingly, treating the edit as a new immutable fact rather than a destructive change that could cause synchronization conflicts.

### What storage technologies work best for chat message synchronization?

Cassandra, DynamoDB, and relational databases all support the required patterns, but the choice depends on scale and consistency requirements. The `liquidslr/system-design-notes` architecture emphasizes **partition tolerance** and **availability**, making distributed stores like Cassandra or DynamoDB suitable for high-throughput scenarios where horizontal scaling is critical.

### How does a new device synchronize years of chat history without overwhelming the network?

Rather than downloading the entire history, the new device queries the persistent store for messages **after sequence number zero** (or the latest available checkpoint), receiving them in paginated batches ordered by sequence. The client stores the highest received sequence number locally, enabling incremental synchronization on subsequent reconnections without re-fetching previously seen messages.