Think Beyond the Happy Path
Designing a Slack-like chat app isn't just about sending messages — it's about scaling reliably, handling failure, and keeping user experience smooth under load. Before diving in, ask yourself:
- How would you ensure strict message ordering in group chats when messages are sent concurrently from distributed devices?
- How would you design real-time notifications for 100M+ users, supporting online (WebSocket) and offline (push) delivery without overwhelming your backend?
- How would you allow users to delete messages — and propagate that change to all clients — while supporting audit logs and retention policies?
- How would you scale your infrastructure to handle users frequently going online and offline across data centers — while maintaining low-latency fanout?
You don't need to answer all of these right away. A Senior/Staff+ Engineer doesn't stop at basic functionality — they anticipate edge cases, design for resilience, and push for production-grade reliability.
But great systems start with great questions. What would you ask next?
Problem Statement
Design a Slack-like chat system which supports:
- Send messages to a group of chats or to an individual user.
- Send notifications to users who are offline.
- Support for rich media (images or videos etc) in a chat message.
- Delete a message by sender. And the scale is billion user level.
Functional Requirements
FR1 – Users can send and receive messages in 1:1 or group chats
This covers both private and group conversations, with messages delivered in order and stored persistently. Group chats can support dynamic membership and maintain message history.
FR2 – Users can send rich media (e.g., images, videos, files) as part of a message
Messages can include text, attachments, or both. The system must support media upload, storage, and preview rendering across devices.
FR3 – Users can receive real-time notifications for new messages
When a user receives a new message—either in a direct chat or group—they should be notified through in-app indicators, push notifications, or email (based on settings and availability).
FR4 – Users can delete their own messages from a chat
A user should be able to remove a message they previously sent. The deletion may be visible (e.g., “This message was deleted”) or fully removed from view, depending on the UX and policy.
Non-Functional Requirements
NFR1 – High Scalability
The system should support 1 B users and 100k chats concurrently, especially in high-traffic group chats.
Illustration: Imagine a global company’s #all-hands channel with 100k active users — your system must fan out messages quickly without bottlenecks.
Pro Tip: Show Scalable Thinking with Numbers
When discussing non-functional requirements in interviews, don’t just say "high scalability" — you need to quantify it. For example, instead of saying: "The system should handle lots of users", you can say: "The system should support 10M+ users, with 5K concurrent messages per second in large chat rooms." Adding numeric expectations (like target QPS, latency percentiles, or fan-out scale) shows that you understand real-world constraints. This is a Staff+ signal — it shows not just that you can design systems, but that you've seen them operate at scale.
NFR2 – Low Latency in Chat Delivery
Aim for P95 end-to-end message latency under 200ms for online users.
Illustration: A user sends a message and expects to see it reflected across all devices in near real-time — delays over 500ms degrade the user experience.
Why 200ms Isn’t Just a Number?
Low latency in messaging systems isn’t about benchmarks — it’s about perceived real-time-ness. Once message delivery feels delayed, trust and engagement drop fast.
In practice, aiming for P95 latency under 200ms means thinking end-to-end:
- Client → API Gateway: 30ms
- Message Broker or Queue Delay: 10–50ms
- Fan-out to Recipients (DB or WebSocket push): 50–80ms
- Client render time: 20–40ms
What Staff+ candidates do well:
They don’t just say "low latency" — they outline the latency budget, identify hot paths, and explain what tradeoffs they’d make when things spike.
Bonus: If you can say “we’d fall back to degraded fan-out or batch sends at 500ms+ load,” that’s a signal you’ve seen the real fire.
NFR3 – CAP - Aware Messaging Guarantees
The system should prefer high availability, but must enforce strong ordering consistency within each chat room.
Illustration: Even if one server fails, the user should still be able to chat — and messages in a group must never arrive out of order.
CAP Theorem — A Critical Concept for Chat System Design
CAP Theorem states that in any distributed system, you can only guarantee two out of the following three properties at the same time:
- Consistency: All clients see the same data, even when there are concurrent updates.
- Availability: Every request receives a response, even when some servers fail.
- Partition Tolerance: The system continues to operate even if network delays or failures split it into disconnected parts.
Since network partitions are inevitable in real-world systems, the practical choice is between consistency and availability.
Many materials simplify this by choosing availability over consistency. For example, they might accept out-of-order messages to keep the chat responsive. But in messaging systems, users often expect messages to appear in order — especially in group chats or support conversations.
Our recommendation for Staff+ engineers: instead of picking a default, use this as an opportunity for deeper discussion. Ask:
- What level of consistency is truly required for this feature?
- Can we isolate consistency boundaries (e.g., enforce strong ordering within each chat room)?
- Can we design for availability while still guaranteeing a good user experience under partition?
Treat CAP not as a binary choice, but as a design framework. Good engineers make a decision. Great engineers justify it in context.
NFR4 – Message Durability
Messages must be safely persisted before confirming to the sender. No acknowledged message should be lost, even if the server crashes right after.
Illustration: If a user hits send and then immediately loses internet, the message must still exist when they reconnect.
When Is It Safe to Delete a Message?
Just because a message was sent — or even delivered — doesn’t mean it can be deleted. In real-world chat systems, a message must remain durably stored until one of the following conditions is met:
- The sender explicitly deletes the message (user-initiated deletion).
- The system enforces a retention policy, such as auto-deletion after 30 days (e.g., for compliance or storage limits).
Even if a recipient has "seen" the message, it must stay in backend storage for multi-device sync, scroll-back history, and recovery after app reload.
Staff+ engineers should treat durability and deletion as two separate concerns — never assume delivery means it’s safe to erase.
NFR5 – Consistent in Multi-Device
Users often stay logged in on multiple devices — desktop, phone, tablet — at the same time. The system must ensure that messages, read receipts, and typing indicators remain synchronized across all active sessions. Events such as message delivery, deletion, or read status must be reflected in real-time across all clients.
Illustration: A user reads a message on their laptop — seconds later, their phone reflects it as read, clears the notification badge, and doesn’t re-alert them. We put this in the end, is to make sure all above functions can be fulfilled first and cross-device session can hold them all.
Multi-Device Sync: A Silent Interview Killer
It’s easy to overlook this, but messaging systems must account for users with multiple simultaneous sessions. Without proper sync, users may:
- See stale read states
- Get duplicate notifications
- Lose trust in the product's reliability
Staff+ candidates stand out by calling out these edge cases and proposing:
- Session tracking per device
- Idempotent event delivery
- Consistent state propagation across all logged-in clients
If you bring this up in an interview — you’re showing real-world maturity.
Requirement Summary
| Functional Requirements (FRs) | |
|---|---|
| Name | Description |
1. Messaging in 1:1 and Group Chats | Users can send and receive messages in private or group conversations, with ordered delivery and persistent history. |
2. Rich Media Support | Messages can include images, videos, and files; media is uploaded, stored, and rendered across devices. |
3. Real-Time Notifications | Users receive in-app and push notifications when new messages are sent in chats they belong to. |
4. Message Deletion | Users can delete their own messages; deletions must reflect across all devices and preserve conversation flow. |
| Non-Functional Requirements (NFRs) | |
|---|---|
| Name | Description |
1. High Scalability | Support 1B+ MAU and 100K+ concurrent messages/sec across large chat rooms. |
2. Low Latency | Ensure P95 end-to-end message delivery latency under 200ms. |
3. CAP-Aware Messaging Guarantees | Prioritize availability, but enforce strong per-chat message ordering. |
4. Message Durability | Persist messages before acknowledging; no acknowledged message should ever be lost. |
5. Consistent in Multi-Device | P95 sync latency < 300ms across devices; <1% inconsistency in read state or notifications. |
Before diving into Entities & APIs, you may notice that we’ve spent a significant amount of time unpacking both functional and non-functional requirements. That’s intentional. In many existing materials, this critical phase is often rushed or oversimplified — but in real-world design and high-level interviews, clarity around what the system must do and how it must behave under pressure is what separates great solutions from generic ones. We aim to set a higher bar here, not just to define the problem well, but to invite thoughtful trade-off discussions, scalability considerations, and system behaviors that hold up in production. The better your foundation, the sharper your design decisions will be — and we’ll carry that mindset through the rest of this article.
Core Entities
| Core Entities | |
|---|---|
| Name | Description |
User | A registered participant who can chat and receive notifications. |
Chat | A channel or direct conversation with a list of members. |
Message | A text or media unit sent to a chat. |
Media | Media asset linked to a message (image, video, file). |
DeviceSession | Active WebSocket or push-notification endpoint for a user's device. |
Slack Threaded Messages: A Crucial Extension to the Message Entity
In Slack-like systems, Threads allow users to reply to a specific message — creating a nested conversation within a larger chat
To support this, the Message entity should include an optional field: parent_message_id (nullable)
If present, this indicates the message is a thread reply, and the original message becomes the thread root.
This design allows you to:
- Group replies under a thread view
- Fetch all replies for a given message (
/chats/{id}/messages?parent=xyz) - Trigger notifications only to participants of that thread
Staff+ engineers often raise this early — not because it's hard to implement, but because it has major UX, API, and performance implications down the line.
In our design, threads are modeled as messages with a parent_message_id, rather than a separate Thread entity. This keeps the model flexible while supporting both flat and threaded conversations.
Why We Use DeviceSession Instead of Just Clients
While some materials refer to "clients" (like mobile or web apps) when modeling messaging systems, we explicitly use the term DeviceSession to emphasize active, trackable connections between users and devices.
A client describes the platform (e.g., iOS app or web browser), but it doesn't distinguish between:
- Online vs offline state
- Multiple simultaneous logins
- Real-time delivery routes like WebSocket or push tokens
DeviceSession allows us to model:
- Which devices are currently online
- How to route events (e.g., via WebSocket or FCM)
- Cross-device consistency, like syncing read receipts or preventing duplicate notifications
This precision is especially important in Slack-like systems, where users often operate across multiple devices at once. Including DeviceSession as a core entity gives us the flexibility to support reliable real-time behavior at scale.