System design · build it to learn it

Design a Chat System

From "send" to ✓✓ — and everything the ticks are hiding.

"Design a chat system like WhatsApp. A sends B a message; A sees sent, delivered, read ticks. How does the message get there, what happens when B is offline, and how does this scale to hundreds of millions of users?"

On this page: takeaway · requirements · capacity math · architecture · deep dives · API + data model · trade-offs · failure modes · what I'd actually build · interview tips · live walkthrough

🎯 The takeaway, first

A chat system is a pipeline with receipts. Your message travels: your phone → a gateway holding your websocket → a message service that stamps it with an ID and writes it to storage before acknowledging → a fan-out step that finds the recipient's live connections → their phone. The ticks are just receipts flowing back along the same path: ✓ = "the server has it," ✓✓ = "their device has it," ✓✓ blue = "they opened the chat."

The interview sentence: "Persist before ack, idempotency keys for retries, fan-out to live connections with an offline queue and push fallback — the socket is the last mile, the guarantees are the system."

🚫 Misconception, busted

"It's just websockets — open a socket per user and push messages through." The socket is the easy 10%. The system is everything around it: who assigns the message ID (the server, so retries don't create duplicates), when you acknowledge (after the write hits disk, not before), where undelivered messages wait (an offline queue, not the void), and how you find which server holds the recipient's socket (a connection registry). Teams that "just use websockets" rediscover all of this during their first outage — usually at 2am, usually involving lost messages.

01Requirements

Functional

  • 1:1 and group messaging, text + media.
  • ✓ sent / ✓✓ delivered / ✓✓ read receipts, per message.
  • Typing indicators, presence (online/last seen).
  • Offline delivery: messages arrive when the recipient reconnects.
  • Full history sync on new devices.

Non-functional

  • 100M daily users, global, 99.99% availability.
  • Send-to-delivered latency p99 under 300ms.
  • Messages never lost, never duplicated on screen (idempotent display).
  • Per-conversation ordering guaranteed.

02Back-of-envelope math

Analogy first: the system is a postal service where every letter needs a signed receipt at each handoff. The math tells you how many post offices (connection servers) and how much warehouse (storage) you need.

QuantityAssumptionResult
Message rate100M DAU × 50 msgs/day5B msgs/day ≈ 58k/sec avg, ~3× peak = 175k/sec
Storage5B msgs/day × ~200 bytes (text + metadata)~1 TB/day — media lives in object storage, not the message DB
Connections100M concurrent websockets, 1M per gateway box~100 gateway servers — connections are the fleet-sizer
Fan-out, group of 1kOne message → 1,000 deliveriesFan-out service does 1k lookups + pushes per group message — groups dominate cost
Tick traffic3 receipts per messageReceipts are ~3× the message volume — batch and coalesce them

Takeaway after the math: connections and fan-out dominate, not message storage. Size the gateway fleet for sockets, and treat group fan-out as the expensive operation it is.

03Architecture

Think of it like registered mail with tracking: the post office (message service) logs the letter and gives you a tracking number before the truck leaves, then scans it at every handoff until it's in the recipient's hands.

flowchart LR
    A["Sender app"] --> B["Gateway
websocket holder"] B --> C["Message service
assign ID + validate"] C --> D[("Message store
persist first")] C --> E["Fan-out
find recipient conns"] E --> F["Recipient gateway"] F --> G["Recipient app"] G -. "ticks: delivered, read" .-> B H["Push service
APNs/FCM"] -. "offline fallback" .-> G E -. "recipient offline" .-> H
Go deeper: why the gateway and the message service are separate

Gateways are stateful and dumb: they hold a million sockets each and just shovel bytes. The message service is stateless and smart: it validates, assigns IDs, writes to storage. Separating them means you can scale sockets (add gateways) independently from message processing (add service replicas), and a gateway deploy doesn't risk message logic. The connection registry — "user X's socket lives on gateway 37" in Redis — is the glue between them.

04Component deep-dives

The send path: persist before ack

The client sends with a client_msg_id it generated. The server assigns the authoritative server_msg_id, writes to storage, then acks. If the ack is lost and the client retries, the server recognizes the client_msg_id and returns the existing ID instead of storing a duplicate. This is the whole reliability story in one exchange.

sequenceDiagram
    participant A as Sender app
    participant G as Gateway
    participant M as Message service
    participant S as Message store
    A->>G: send {client_msg_id: c-99, text}
    G->>M: forward + auth
    M->>M: assign server_msg_id s-1042
    M->>S: write message (durable)
    S->>M: write confirmed
    M->>G: ack {c-99 → s-1042}
    G->>A: ack → show ✓ sent
    Note over A,S: ack only after the write.
retry with same c-99 returns s-1042, no duplicate.

Delivery, ticks, and the offline path

Fan-out looks up the recipient's live connections; each connected device gets the message and acks back, which becomes your ✓✓. Opening the chat sends a read receipt — the blue ✓✓. No live connection? The message waits in the offline queue and a push notification goes out instead.

sequenceDiagram
    participant M as Message service
    participant F as Fan-out
    participant R as Registry
    participant G as Recipient gateway
    participant B as Recipient app
    participant P as Push service
    M->>F: new message s-1042 for user B
    F->>R: where is B connected?
    alt B online
        R->>F: gateway-37, device d-7
        F->>G: deliver s-1042
        G->>B: push over websocket
        B->>G: delivered ack
        G->>F: ✓✓ delivered
        B->>G: chat opened → read receipt
        G->>F: ✓✓ read (blue)
    else B offline
        F->>F: enqueue in offline queue
        F->>P: send push notification
        P->>B: "New message" banner
    end

05API + data model

API (websocket frames)

--> { "type": "send",
     "client_msg_id": "c-99",
     "conv_id": "dm:a:b",
     "text": "hey!" }
<-- { "type": "ack",
     "client_msg_id": "c-99",
     "server_msg_id": "s-1042" }   # ✓ sent

<-- { "type": "receipt",
     "server_msg_id": "s-1042",
     "status": "delivered" }        # ✓✓
<-- { "type": "receipt",
     "server_msg_id": "s-1042",
     "status": "read" }            # ✓✓ blue

Data model

server_msg_id (pk)Snowflake — ordered, unique
client_msg_id (unique)idempotency key per sender
conv_id + seqper-conversation ordering
sender / payloadtext or media pointer
statussent · delivered · read

Media goes to object storage; the message row holds only a pointer. History queries are WHERE conv_id = ? ORDER BY seq DESC LIMIT 50 — the one query pattern the schema serves.

06Trade-offs

DecisionOption AOption BVerdict
TransportWebSockets — bidirectional, matureSSE / long-poll — simpler, half-duplexWebSockets; you need client→server receipts anyway
Fan-outWrite-time: push to all devices nowRead-time: devices pull on openWrite-time for 1:1; hybrid for giant groups (push a "new message" ping, pull the batch)
EncryptionEnd-to-end — server sees ciphertextTLS + server-side plaintext — search, bots, moderation workE2E for 1:1 (it's the expectation now); server-side for group/business features
ID assignmentServer assigns — single ordering authorityClient assigns UUID — simpler, no orderingServer (Snowflake): ordering + dedup in one move

07Failure modes

08What I'd actually build

09Interview tips

10🔬 Live widget: send a message, watch it travel

A working miniature of the pipeline: type a message, send it, and watch it hop you → gateway → message service → storage → fan-out → Maya while the ticks flip from ✓ to ✓✓ to blue. Toggle Maya offline to see the push-fallback path.

M
Maya
● online

🚚 Pipeline

📜 Trace

send a message to trace its journey…
Go deeper: why the ticks sometimes lie

✓✓ "delivered" means a device received it — not that a human saw it. Multi-device makes this subtle: your message can be "delivered" to a laptop sitting closed in a bag while the phone is offline. That's why read receipts are per-conversation-open, not per-device, and why "delivered" is really "the server handed it to something you own." The blue ticks are the only ones that involve eyeballs.

Built as a single self-contained file · diagrams render locally, no CDN · back to the takeaway ↑

v2026.10.03-01