Communication Patterns Suitable for Chat Systems: A Complete Technical Guide
TLDR: Production chat systems combine HTTP POST for message ingestion, WebSockets for real-time delivery, and publish-subscribe patterns for presence management to achieve low latency at scale.
The liquidslr/system-design-notes repository outlines the communication patterns suitable for chat systems that support millions of daily active users. According to 12. Chat System/Readme.md, modern messaging architectures use a hybrid approach that balances connection state, latency requirements, and firewall constraints across multiple protocols.
HTTP POST for Stateless Message Ingestion
The entry point for new messages typically uses standard HTTP POST requests. In this pattern, the client issues a regular HTTP request that the API layer validates, assigns a unique ID, and writes to persistent storage before returning a quick acknowledgment.
This approach is stateless, works behind corporate firewalls, and scales horizontally behind load balancers. However, it is not suitable for real-time delivery since the client must open a new connection for each message.
# HTTP POST – message ingestion pattern
import requests, json
def send_message(api_url, auth_token, chat_id, sender_id, text):
payload = {
"chat_id": chat_id,
"sender_id": sender_id,
"text": text,
}
headers = {"Authorization": f"Bearer {auth_token}"}
r = requests.post(f"{api_url}/messages", json=payload, headers=headers)
r.raise_for_status()
return r.json() # { "message_id": 12345 }
Polling and Long Polling as Fallback Mechanisms
When persistent connections are not possible, chat implementations fall back to HTTP-based polling strategies.
Simple Polling
In the polling pattern, the client periodically issues GET /messages?lastSeen=… requests to check for new messages. While easy to implement and compatible with any HTTP client, this generates redundant traffic and introduces high latency for new messages.
Long Polling
Long polling improves efficiency by holding the HTTP request open until a new message arrives or a timeout expires. Once the server responds, the client immediately re-issues the request. This reduces unnecessary requests compared to plain polling and works over standard HTTP ports, though it still incurs request-response overhead.
# Long Polling – fallback when WebSockets are blocked
import requests, time
def long_poll(api_url, auth_token, last_seen):
while True:
r = requests.get(
f"{api_url}/messages?since={last_seen}",
headers={"Authorization": f"Bearer {auth_token}"},
timeout=35, # server holds request ~30s
)
msgs = r.json()
if msgs:
for m in msgs:
print(f"[{m['sender']}] {m['text']}")
last_seen = msgs[-1]["message_id"]
else:
# timeout → repeat request immediately
continue
WebSockets for Real-Time Bi-Directional Communication
WebSockets represent the default pattern for modern real-time chat services. After an initial HTTP handshake, the client and server exchange frames over a single TCP socket, enabling instant push of messages in both directions.
As shown in 12. Chat System/images/websocket.png, this pattern provides near-zero latency, low overhead, and full-duplex communication. The implementation requires handling connection upgrades, reconnection logic, and heartbeat mechanisms to detect silent disconnections.
# WebSocket – real-time delivery (client side)
import asyncio, websockets, json
async def chat_ws(uri, token):
async with websockets.connect(uri, extra_headers={"Authorization": f"Bearer {token}"}) as ws:
# send a heartbeat every 30s
async def heartbeat():
while True:
await ws.send(json.dumps({"type": "heartbeat"}))
await asyncio.sleep(30)
asyncio.create_task(heartbeat())
# receive incoming messages
async for raw in ws:
msg = json.loads(raw)
if msg["type"] == "chat":
print(f"[{msg['sender']}] {msg['text']}")
asyncio.run(chat_ws("wss://chat.example.com/ws", "user-jwt-token"))
Server-Sent Events for One-Way Server Push
Server-Sent Events (SSE) stream text/event data over an HTTP response, allowing the server to push updates to the client as they arrive. This pattern is simpler than WebSockets for uni-directional streams and works with standard HTTP infrastructure.
However, SSE does not support client-to-server messaging on the same channel, making it unsuitable for chat scenarios where both sides actively send messages.
Publish-Subscribe for Presence and Fan-Out
For distributing online/offline status, typing indicators, or group-chat updates, chat systems implement a publish-subscribe pattern. Each user relationship creates a logical channel; when a status changes, the server publishes to the channel and all subscribed friends receive the update via the fan-out model illustrated in 12. Chat System/images/fanout-presence.png.
This decouples producers from consumers and scales well for small groups, though it becomes costly for very large groups without careful sharding.
# Publish-Subscribe – presence fan-out (using Redis pub/sub)
import redis, json
r = redis.Redis(host="redis", port=6379)
def publish_presence(user_id, status):
channel = f"presence:{user_id}"
r.publish(channel, json.dumps({"user_id": user_id, "status": status}))
def subscribe_friends(friends):
pubsub = r.pubsub()
for fid in friends:
pubsub.subscribe(f"presence:{fid}")
for msg in pubsub.listen():
if msg['type'] == 'message':
data = json.loads(msg['data'])
print(f"Friend {data['user_id']} is now {data['status']}")
Heartbeat Mechanisms for Connection Liveness
To detect when a client has silently disconnected, chat systems implement heartbeat mechanisms. Clients periodically emit small "heartbeat" messages, and the presence service marks the user offline if a heartbeat is missed within a configurable threshold. This integrates with any transport layer, whether WebSocket, HTTP, or long polling, though it requires tuning the interval and timeout to balance accuracy against network traffic.
Summary
- HTTP POST provides stateless, firewall-friendly message ingestion that scales horizontally behind load balancers.
- Long polling offers a viable fallback when WebSockets are blocked by corporate proxies, reducing traffic compared to simple polling.
- WebSockets deliver near-zero latency bi-directional communication, making them the default for modern real-time chat.
- Publish-subscribe architectures handle presence updates and fan-out efficiently for small-to-medium user groups.
- Heartbeats maintain accurate online status detection across all connection types, preventing ghost user scenarios.
Frequently Asked Questions
What is the best communication pattern for real-time chat?
WebSockets are the optimal choice for real-time chat because they provide persistent, full-duplex TCP connections that enable instant message push with minimal overhead. According to the source code analysis in 12. Chat System/Readme.md, this pattern supports the sub-second latency required for 50 million daily active users while maintaining efficient resource usage through connection reuse.
When should I use long polling instead of WebSockets?
Use long polling when your deployment environment blocks WebSocket upgrade requests, such as certain corporate proxies or legacy network infrastructure. Long polling holds HTTP requests open until data arrives, providing better latency than simple polling while maintaining compatibility with standard HTTP ports, though it cannot match the efficiency of true persistent connections.
How do chat systems detect when a user goes offline?
Systems implement heartbeat mechanisms where clients send periodic ping messages at fixed intervals. The server maintains a threshold timer for each connection; if no heartbeat is received within the timeout window, the presence service marks the user as offline. This pattern works across WebSocket, HTTP, and long polling transports to prevent ghost sessions from consuming resources.
Why separate HTTP POST from the real-time delivery channel?
Separating HTTP POST (for sending messages) from the delivery channel follows the principle of separation of concerns: POST provides reliable, stateless ingestion that works behind firewalls and scales horizontally, while persistent connections handle the high-frequency, low-latency delivery requirements. This hybrid architecture allows the ingestion layer to focus on durability while the delivery layer optimizes for speed.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →