From "send" to ✓✓ — and everything the ticks are hiding.
"Design a chat system like WhatsApp. A sends B a message; A sees sent, delivered, read ticks. How does the message get there, what happens when B is offline, and how does this scale to hundreds of millions of users?"
A chat system is a pipeline with receipts. Your message travels: your phone → a gateway holding your websocket → a message service that stamps it with an ID and writes it to storage before acknowledging → a fan-out step that finds the recipient's live connections → their phone. The ticks are just receipts flowing back along the same path: ✓ = "the server has it," ✓✓ = "their device has it," ✓✓ blue = "they opened the chat."
The interview sentence: "Persist before ack, idempotency keys for retries, fan-out to live connections with an offline queue and push fallback — the socket is the last mile, the guarantees are the system."
"It's just websockets — open a socket per user and push messages through." The socket is the easy 10%. The system is everything around it: who assigns the message ID (the server, so retries don't create duplicates), when you acknowledge (after the write hits disk, not before), where undelivered messages wait (an offline queue, not the void), and how you find which server holds the recipient's socket (a connection registry). Teams that "just use websockets" rediscover all of this during their first outage — usually at 2am, usually involving lost messages.
Analogy first: the system is a postal service where every letter needs a signed receipt at each handoff. The math tells you how many post offices (connection servers) and how much warehouse (storage) you need.
| Quantity | Assumption | Result |
|---|---|---|
| Message rate | 100M DAU × 50 msgs/day | 5B msgs/day ≈ 58k/sec avg, ~3× peak = 175k/sec |
| Storage | 5B msgs/day × ~200 bytes (text + metadata) | ~1 TB/day — media lives in object storage, not the message DB |
| Connections | 100M concurrent websockets, 1M per gateway box | ~100 gateway servers — connections are the fleet-sizer |
| Fan-out, group of 1k | One message → 1,000 deliveries | Fan-out service does 1k lookups + pushes per group message — groups dominate cost |
| Tick traffic | 3 receipts per message | Receipts are ~3× the message volume — batch and coalesce them |
Takeaway after the math: connections and fan-out dominate, not message storage. Size the gateway fleet for sockets, and treat group fan-out as the expensive operation it is.
Think of it like registered mail with tracking: the post office (message service) logs the letter and gives you a tracking number before the truck leaves, then scans it at every handoff until it's in the recipient's hands.
flowchart LR
A["Sender app"] --> B["Gateway
websocket holder"]
B --> C["Message service
assign ID + validate"]
C --> D[("Message store
persist first")]
C --> E["Fan-out
find recipient conns"]
E --> F["Recipient gateway"]
F --> G["Recipient app"]
G -. "ticks: delivered, read" .-> B
H["Push service
APNs/FCM"] -. "offline fallback" .-> G
E -. "recipient offline" .-> H
Gateways are stateful and dumb: they hold a million sockets each and just shovel bytes. The message service is stateless and smart: it validates, assigns IDs, writes to storage. Separating them means you can scale sockets (add gateways) independently from message processing (add service replicas), and a gateway deploy doesn't risk message logic. The connection registry — "user X's socket lives on gateway 37" in Redis — is the glue between them.
The client sends with a client_msg_id it generated. The server assigns the authoritative server_msg_id, writes to storage, then acks. If the ack is lost and the client retries, the server recognizes the client_msg_id and returns the existing ID instead of storing a duplicate. This is the whole reliability story in one exchange.
sequenceDiagram
participant A as Sender app
participant G as Gateway
participant M as Message service
participant S as Message store
A->>G: send {client_msg_id: c-99, text}
G->>M: forward + auth
M->>M: assign server_msg_id s-1042
M->>S: write message (durable)
S->>M: write confirmed
M->>G: ack {c-99 → s-1042}
G->>A: ack → show ✓ sent
Note over A,S: ack only after the write.
retry with same c-99 returns s-1042, no duplicate.
Fan-out looks up the recipient's live connections; each connected device gets the message and acks back, which becomes your ✓✓. Opening the chat sends a read receipt — the blue ✓✓. No live connection? The message waits in the offline queue and a push notification goes out instead.
sequenceDiagram
participant M as Message service
participant F as Fan-out
participant R as Registry
participant G as Recipient gateway
participant B as Recipient app
participant P as Push service
M->>F: new message s-1042 for user B
F->>R: where is B connected?
alt B online
R->>F: gateway-37, device d-7
F->>G: deliver s-1042
G->>B: push over websocket
B->>G: delivered ack
G->>F: ✓✓ delivered
B->>G: chat opened → read receipt
G->>F: ✓✓ read (blue)
else B offline
F->>F: enqueue in offline queue
F->>P: send push notification
P->>B: "New message" banner
end
--> { "type": "send",
"client_msg_id": "c-99",
"conv_id": "dm:a:b",
"text": "hey!" }
<-- { "type": "ack",
"client_msg_id": "c-99",
"server_msg_id": "s-1042" } # ✓ sent
<-- { "type": "receipt",
"server_msg_id": "s-1042",
"status": "delivered" } # ✓✓
<-- { "type": "receipt",
"server_msg_id": "s-1042",
"status": "read" } # ✓✓ blue
server_msg_id (pk)Snowflake — ordered, uniqueclient_msg_id (unique)idempotency key per senderconv_id + seqper-conversation orderingsender / payloadtext or media pointerstatussent · delivered · readMedia goes to object storage; the message row holds only a pointer. History queries are WHERE conv_id = ? ORDER BY seq DESC LIMIT 50 — the one query pattern the schema serves.
| Decision | Option A | Option B | Verdict |
|---|---|---|---|
| Transport | WebSockets — bidirectional, mature | SSE / long-poll — simpler, half-duplex | WebSockets; you need client→server receipts anyway |
| Fan-out | Write-time: push to all devices now | Read-time: devices pull on open | Write-time for 1:1; hybrid for giant groups (push a "new message" ping, pull the batch) |
| Encryption | End-to-end — server sees ciphertext | TLS + server-side plaintext — search, bots, moderation work | E2E for 1:1 (it's the expectation now); server-side for group/business features |
| ID assignment | Server assigns — single ordering authority | Client assigns UUID — simpler, no ordering | Server (Snowflake): ordering + dedup in one move |
client_msg_id dedup — the retry returns the original ID.seq order.client_msg_id unique constraint = free dedup.client_msg_id idempotency before they ask about duplicates — juniors discover duplicates in production.A working miniature of the pipeline: type a message, send it, and watch it hop you → gateway → message service → storage → fan-out → Maya while the ticks flip from ✓ to ✓✓ to blue. Toggle Maya offline to see the push-fallback path.
✓✓ "delivered" means a device received it — not that a human saw it. Multi-device makes this subtle: your message can be "delivered" to a laptop sitting closed in a bag while the phone is offline. That's why read receipts are per-conversation-open, not per-device, and why "delivered" is really "the server handed it to something you own." The blue ticks are the only ones that involve eyeballs.
Built as a single self-contained file · diagrams render locally, no CDN · back to the takeaway ↑