System Design
Most system design material is either a glossary you cannot learn from or a case study that assumes you already know the vocabulary. This is the whole path in order: the words first, then the arithmetic, then the moving parts, then twenty-nine full designs that use all of it. Work through it one topic at a time and tick each one off as you go.
3 of 376 topics have their article written so far. The rest are on the list, in this order.
Introduction to System Design
The vocabulary. Every later part assumes you can say what availability, durability and latency actually mean.
- 4System Design vs Software Architecture
- 5Functional Requirements
- 6Non-Functional Requirements
- 7Constraints and Assumptions
- 8Scalability
- 9Performance
- 10Latency
- 11Throughput
- 12Availability
- 13Reliability
- 14Durability
- 15Fault Tolerance
- 16Maintainability
- 17Extensibility
- 18Security
- 19Cost and Operational Complexity
- 20Understanding Service-Level Indicators, Objectives and Agreements
System Design Interview Process
How the 45 minutes actually run, and what the person across the table is scoring.
- 21How System Design Interviews Work
- 22What Interviewers Evaluate
- 23How to Clarify an Ambiguous Problem
- 24Defining the Scope
- 25Identifying Core Use Cases
- 26Gathering Scale Requirements
- 27Making and Communicating Assumptions
- 28Designing the Simplest Initial Architecture
- 29Finding Bottlenecks
- 30Discussing Trade-Offs
- 31Handling Interviewer Requirement Changes
- 32Managing a 45-Minute Interview
- 33How to Communicate While Whiteboarding
- 34How to Finish and Summarise a Design
- 35Common System Design Interview Mistakes
Back-of-the-Envelope Estimation
Arithmetic you can do out loud without a calculator, and the confidence to round hard.
- 36Why Capacity Estimation Matters
- 37Powers of Ten and Common Data Sizes
- 38Estimating Daily and Monthly Active Users
- 39Calculating Average and Peak QPS
- 40Estimating Read and Write Traffic
- 41Estimating Concurrent Connections
- 42Estimating Storage Requirements
- 43Estimating Memory and Cache Size
- 44Estimating Network Bandwidth
- 45Estimating Future Growth
- 46Avoiding False Precision in Interviews
- 47Complete Capacity Estimation Example
Networking and Communication
What happens between the client pressing a key and your server seeing bytes.
- 48How a Request Travels Through a System
- 49Domain Name System
- 50IP Addresses, Ports and Sockets
- 51TCP vs UDP
- 52HTTP and HTTPS
- 53HTTP/1.1 vs HTTP/2 vs HTTP/3
- 54TLS and SSL Termination
- 55Forward Proxy vs Reverse Proxy
- 56Load Balancers
- 57Layer 4 vs Layer 7 Load Balancing
- 58Load-Balancing Algorithms
- 59Health Checks
- 60Connection Pooling
- 61Keep-Alive Connections
- 62REST APIs
- 63GraphQL
- 64gRPC
- 65WebSockets
- 66Server-Sent Events
- 67Long Polling
- 68Webhooks
- 69Synchronous vs Asynchronous Communication
- 70Service Discovery
- 71API Gateways
- 72API Versioning
- 73Pagination: Offset vs Cursor
Data Storage and Databases
The largest part, because the database is the decision you cannot walk back.
- 74How to Choose a Database
- 75Relational Databases
- 76NoSQL Databases
- 77SQL vs NoSQL
- 78Key-Value Databases
- 79Document Databases
- 80Wide-Column Databases
- 81Graph Databases
- 82Time-Series Databases
- 83Search Databases
- 84Object Storage
- 85File Storage vs Block Storage vs Object Storage
- 86Database Schema Design
- 87Normalisation
- 88Denormalisation
- 89Primary and Foreign Keys
- 90Database Indexes
- 91Composite and Covering Indexes
- 92B-Trees and LSM Trees
- 93Database Transactions
- 94ACID Properties
- 95Database Isolation Levels
- 96Dirty Reads, Non-Repeatable Reads and Phantom Reads
- 97Optimistic vs Pessimistic Locking
- 98Database Replication
- 99Leader-Follower Replication
- 100Multi-Leader and Leaderless Replication
- 101Read Replicas
- 102Database Partitioning
- 103Horizontal vs Vertical Partitioning
- 104Database Sharding
- 105Choosing a Shard Key
- 106Hot Partitions
- 107Rebalancing Shards
- 108Multi-Tenant Database Design
- 109Database Connection Pooling
- 110Database Migrations at Scale
- 111Soft Deletion, Archiving and Data Retention
Caching
The cheapest performance win and the richest source of subtle bugs.
- 112What Is Caching?
- 113Why Caching Improves Performance
- 114Browser, CDN, Application and Database Caching
- 115Cache-Aside Pattern
- 116Read-Through Cache
- 117Write-Through Cache
- 118Write-Behind Cache
- 119Cache Invalidation
- 120Time to Live
- 121Cache Eviction Policies
- 122Cache Consistency
- 123Cache Stampede
- 124Cache Penetration
- 125Cache Avalanche
- 126Hot Keys
- 127Distributed Caching
- 128Redis in System Design
- 129When Caching Makes a System Worse
Message Queues and Event-Driven Systems
Decoupling work from the request that asked for it, and the delivery guarantees that follow.
- 130What Is a Message Queue?
- 131Why We Use Asynchronous Processing
- 132Message Queues vs Event Streams
- 133Producers, Consumers and Brokers
- 134Consumer Groups
- 135Message Ordering
- 136At-Most-Once Delivery
- 137At-Least-Once Delivery
- 138Exactly-Once and Effectively-Once Processing
- 139Idempotency
- 140Retries and Exponential Backoff
- 141Dead-Letter Queues
- 142Backpressure
- 143Kafka vs RabbitMQ vs Cloud Queues
- 144Publish-Subscribe Systems
- 145Event-Driven Architecture
- 146Event Sourcing
- 147CQRS
- 148Saga Pattern
- 149Transactional Outbox Pattern
- 150Change Data Capture
- 151Stream Processing
- 152Batch Processing vs Stream Processing
Distributed Systems Fundamentals
The theory that explains why the obvious design does not work across a network.
- 153What Is a Distributed System?
- 154Challenges of Distributed Systems
- 155Network Partitions
- 156Strong Consistency
- 157Eventual Consistency
- 158Read-After-Write and Causal Consistency
- 159CAP Theorem
- 160PACELC Theorem
- 161Quorum Reads and Writes
- 162Consensus
- 163Leader Election
- 164Distributed Locks
- 165Consistent Hashing
- 166Clock and Time Problems
- 167Logical and Vector Clocks
- 168Conflict Detection and Resolution
- 169Split-Brain Problems
- 170Gossip Protocol
- 171Data Locality
- 172Distributed Transactions
- 173Two-Phase Commit
- 174Handling Partial Failure
Architecture and Scaling
Shapes a system can take, and what each shape costs to run.
- 175Vertical vs Horizontal Scaling
- 176Stateless vs Stateful Services
- 177Monolithic Architecture
- 178Modular Monolith
- 179Microservices Architecture
- 180Monolith vs Microservices
- 181Service Boundaries
- 182Domain-Driven Design
- 183Service-Oriented Architecture
- 184Serverless Architecture
- 185Containers
- 186Container Orchestration
- 187Autoscaling
- 188Session Management in Distributed Systems
- 189Sticky Sessions
- 190Multi-Region Architecture
- 191Active-Active vs Active-Passive Architecture
- 192Edge Computing
- 193Cell-Based Architecture
- 194Control Plane vs Data Plane
- 195Multi-Tenant Architecture
- 196Build vs Buy Decisions
Reliability and Resilience
Assuming every dependency fails, and designing so the failure stays small.
- 197Designing for Failure
- 198Single Points of Failure
- 199Timeouts
- 200Retries
- 201Circuit Breakers
- 202Bulkhead Pattern
- 203Graceful Degradation
- 204Load Shedding
- 205Rate Limiting
- 206Throttling
- 207Redundancy
- 208Failover
- 209Disaster Recovery
- 210Recovery Point Objective and Recovery Time Objective
- 211Backup and Restore Strategies
- 212Idempotency Keys
- 213Handling Duplicate Requests
- 214Preventing Cascading Failures
- 215Chaos Engineering
- 216Zero-Downtime Architecture
Common Specialised Systems
Named building blocks that recur across interview problems.
- 217Content Delivery Networks
- 218Full-Text Search
- 219Inverted Indexes
- 220Search Ranking
- 221Autocomplete Systems
- 222Recommendation Systems
- 223News Feed Generation
- 224Fan-Out on Write vs Fan-Out on Read
- 225Real-Time Chat Systems
- 226Presence and Online Status
- 227Notification Systems
- 228Location-Based Services
- 229Geohashing
- 230Payment Systems
- 231Inventory and Reservation Systems
- 232Distributed Job Schedulers
- 233Web Crawlers
- 234Media Upload and Processing
- 235Video Streaming
- 236Collaborative Editing
- 237Analytics and Metrics Pipelines
- 238Logging Platforms
Security
The part candidates skip and interviewers remember.
- 239Authentication vs Authorisation
- 240Sessions vs Tokens
- 241JSON Web Tokens
- 242OAuth 2.0 and OpenID Connect
- 243Role-Based and Attribute-Based Access Control
- 244Encryption in Transit and at Rest
- 245Secrets Management
- 246API Security
- 247Rate Limiting and Abuse Prevention
- 248DDoS Protection
- 249Secure File Uploads
- 250Signed URLs
- 251Webhook Security
- 252Replay-Attack Prevention
- 253Tenant Isolation
- 254Personally Identifiable Information
- 255Data Retention and Deletion
- 256Audit Logging
- 257Threat Modelling
Observability and Operations
Knowing what the system is doing, and changing it without downtime.
- 258Logs, Metrics and Traces
- 259Structured Logging
- 260Correlation IDs
- 261Distributed Tracing
- 262Important System Metrics
- 263SLIs, SLOs and Error Budgets
- 264Alerting
- 265Dashboards
- 266Health, Readiness and Liveness Checks
- 267Capacity Planning
- 268Performance Testing
- 269Load, Stress, Spike and Soak Testing
- 270Profiling and Bottleneck Analysis
- 271Feature Flags
- 272Rolling Deployments
- 273Blue-Green Deployments
- 274Canary Deployments
- 275Rollbacks
- 276Infrastructure as Code
- 277Cost Optimisation
Applied AI System Design
What changes when a model is in the request path: cost, latency, and non-determinism.
- 278How AI System Design Differs from Traditional System Design
- 279Online vs Offline Inference
- 280Model-Serving Architecture
- 281CPU vs GPU Inference
- 282Model Batching
- 283Model Gateways and Routers
- 284Model and Provider Fallback
- 285Prompt and Model Versioning
- 286Streaming AI Responses
- 287Semantic Caching
- 288Embedding Pipelines
- 289Vector Databases in System Design
- 290RAG System Architecture
- 291Document Ingestion Pipelines
- 292Chunking and Document Versioning
- 293Hybrid Search and Reranking
- 294Agent Architecture
- 295Agent Orchestration
- 296Tool Execution and Permissions
- 297Workflow State and Checkpointing
- 298Human-in-the-Loop Systems
- 299AI Guardrails
- 300AI Evaluation Systems
- 301Hallucination and Faithfulness Evaluation
- 302Abstention and Confidence Handling
- 303AI Observability
- 304AI Cost and Latency Optimisation
- 305Multi-Tenant AI Platforms
- 306Data Privacy in AI Systems
- 307Designing Reliable Agent Workflows
Low-Level Design
Class-level design: the other interview, and the one that rewards precision.
- 308What Is Low-Level Design?
- 309Object-Oriented Design
- 310SOLID Principles
- 311Composition vs Inheritance
- 312Interfaces and Abstractions
- 313Dependency Injection
- 314Common Design Patterns
- 315Factory Pattern
- 316Strategy Pattern
- 317Observer Pattern
- 318Adapter Pattern
- 319Decorator Pattern
- 320State Pattern
- 321Repository Pattern
- 322State Machines
- 323Domain Modelling
- 324Concurrency and Thread Safety
- 325Extensible API and Class Design
- 326Testing a Low-Level Design
System Design Case Studies
Full problems, worked end to end. This is where the earlier parts get used.
- 327Design a URL Shortener
- 328Design a Rate Limiter
- 329Design Pastebin
- 330Design a Notification System
- 331Design a Distributed Job Scheduler
- 332Design a Web Crawler
- 333Design Cloud File Storage
- 334Design a Logging System
- 335Design a Metrics Platform
- 336Design a Social Media Feed
- 337Design a Chat Application
- 338Design a Video-Streaming Platform
- 339Design a Ride-Sharing Platform
- 340Design an E-Commerce Platform
- 341Design a Payment Platform
- 342Design a Ticket-Booking System
- 343Design a Collaborative Document Editor
- 344Design a Search Engine
- 345Design an Autocomplete Service
- 346Design a Multi-Tenant SaaS Platform
- 347Design an API Gateway
- 348Design a Distributed Cache
- 349Design a Message Broker
- 350Design a ChatGPT-Like Application
- 351Design a Multi-Tenant RAG Platform
- 352Design an AI Coding Assistant
- 353Design an AI Document-Processing Platform
- 354Design an Agent-Workflow Platform
- 355Design an AI Evaluation Platform
Low-Level Design Exercises
The classic object-modelling problems, each small enough to finish in one sitting.
- 356Design a Parking Lot
- 357Design an Elevator System
- 358Design a Vending Machine
- 359Design a Library System
- 360Design a Logger
- 361Design a Task Scheduler
- 362Design a Notification Framework
- 363Design a Payment State Machine
- 364Design a Ride-Booking System
- 365Design a Board Game
Final Interview Preparation
Rehearsal. A repeatable template and the drills that make it automatic.
- 366Creating a Reusable Whiteboard Template
- 367Practising Requirement Clarification
- 368Practising Capacity Estimation
- 369Practising APIs and Data Models
- 370Practising Critical Request Flows
- 371Practising Failure Scenarios
- 372Practising Trade-Off Explanations
- 373Conducting Timed Mock Interviews
- 374Handling Follow-Up Questions
- 375Evaluating Your Own Performance
- 376Final System Design Interview Checklist