System Design

Most system design material is either a glossary you cannot learn from or a case study that assumes you already know the vocabulary. This is the whole path in order: the words first, then the arithmetic, then the moving parts, then twenty-nine full designs that use all of it. Work through it one topic at a time and tick each one off as you go.

3 of 376 topics have their article written so far. The rest are on the list, in this order.

Your progress

0 of 376

Start at 1

What Is System Design? in Part 1

Part 10/20

Introduction to System Design

The vocabulary. Every later part assumes you can say what availability, durability and latency actually mean.

  • 4System Design vs Software Architecture
  • 5Functional Requirements
  • 6Non-Functional Requirements
  • 7Constraints and Assumptions
  • 8Scalability
  • 9Performance
  • 10Latency
  • 11Throughput
  • 12Availability
  • 13Reliability
  • 14Durability
  • 15Fault Tolerance
  • 16Maintainability
  • 17Extensibility
  • 18Security
  • 19Cost and Operational Complexity
  • 20Understanding Service-Level Indicators, Objectives and Agreements
Part 20/15

System Design Interview Process

How the 45 minutes actually run, and what the person across the table is scoring.

  • 21How System Design Interviews Work
  • 22What Interviewers Evaluate
  • 23How to Clarify an Ambiguous Problem
  • 24Defining the Scope
  • 25Identifying Core Use Cases
  • 26Gathering Scale Requirements
  • 27Making and Communicating Assumptions
  • 28Designing the Simplest Initial Architecture
  • 29Finding Bottlenecks
  • 30Discussing Trade-Offs
  • 31Handling Interviewer Requirement Changes
  • 32Managing a 45-Minute Interview
  • 33How to Communicate While Whiteboarding
  • 34How to Finish and Summarise a Design
  • 35Common System Design Interview Mistakes
Part 30/12

Back-of-the-Envelope Estimation

Arithmetic you can do out loud without a calculator, and the confidence to round hard.

  • 36Why Capacity Estimation Matters
  • 37Powers of Ten and Common Data Sizes
  • 38Estimating Daily and Monthly Active Users
  • 39Calculating Average and Peak QPS
  • 40Estimating Read and Write Traffic
  • 41Estimating Concurrent Connections
  • 42Estimating Storage Requirements
  • 43Estimating Memory and Cache Size
  • 44Estimating Network Bandwidth
  • 45Estimating Future Growth
  • 46Avoiding False Precision in Interviews
  • 47Complete Capacity Estimation Example
Part 40/26

Networking and Communication

What happens between the client pressing a key and your server seeing bytes.

  • 48How a Request Travels Through a System
  • 49Domain Name System
  • 50IP Addresses, Ports and Sockets
  • 51TCP vs UDP
  • 52HTTP and HTTPS
  • 53HTTP/1.1 vs HTTP/2 vs HTTP/3
  • 54TLS and SSL Termination
  • 55Forward Proxy vs Reverse Proxy
  • 56Load Balancers
  • 57Layer 4 vs Layer 7 Load Balancing
  • 58Load-Balancing Algorithms
  • 59Health Checks
  • 60Connection Pooling
  • 61Keep-Alive Connections
  • 62REST APIs
  • 63GraphQL
  • 64gRPC
  • 65WebSockets
  • 66Server-Sent Events
  • 67Long Polling
  • 68Webhooks
  • 69Synchronous vs Asynchronous Communication
  • 70Service Discovery
  • 71API Gateways
  • 72API Versioning
  • 73Pagination: Offset vs Cursor
Part 50/38

Data Storage and Databases

The largest part, because the database is the decision you cannot walk back.

  • 74How to Choose a Database
  • 75Relational Databases
  • 76NoSQL Databases
  • 77SQL vs NoSQL
  • 78Key-Value Databases
  • 79Document Databases
  • 80Wide-Column Databases
  • 81Graph Databases
  • 82Time-Series Databases
  • 83Search Databases
  • 84Object Storage
  • 85File Storage vs Block Storage vs Object Storage
  • 86Database Schema Design
  • 87Normalisation
  • 88Denormalisation
  • 89Primary and Foreign Keys
  • 90Database Indexes
  • 91Composite and Covering Indexes
  • 92B-Trees and LSM Trees
  • 93Database Transactions
  • 94ACID Properties
  • 95Database Isolation Levels
  • 96Dirty Reads, Non-Repeatable Reads and Phantom Reads
  • 97Optimistic vs Pessimistic Locking
  • 98Database Replication
  • 99Leader-Follower Replication
  • 100Multi-Leader and Leaderless Replication
  • 101Read Replicas
  • 102Database Partitioning
  • 103Horizontal vs Vertical Partitioning
  • 104Database Sharding
  • 105Choosing a Shard Key
  • 106Hot Partitions
  • 107Rebalancing Shards
  • 108Multi-Tenant Database Design
  • 109Database Connection Pooling
  • 110Database Migrations at Scale
  • 111Soft Deletion, Archiving and Data Retention
Part 60/18

Caching

The cheapest performance win and the richest source of subtle bugs.

  • 112What Is Caching?
  • 113Why Caching Improves Performance
  • 114Browser, CDN, Application and Database Caching
  • 115Cache-Aside Pattern
  • 116Read-Through Cache
  • 117Write-Through Cache
  • 118Write-Behind Cache
  • 119Cache Invalidation
  • 120Time to Live
  • 121Cache Eviction Policies
  • 122Cache Consistency
  • 123Cache Stampede
  • 124Cache Penetration
  • 125Cache Avalanche
  • 126Hot Keys
  • 127Distributed Caching
  • 128Redis in System Design
  • 129When Caching Makes a System Worse
Part 70/23

Message Queues and Event-Driven Systems

Decoupling work from the request that asked for it, and the delivery guarantees that follow.

  • 130What Is a Message Queue?
  • 131Why We Use Asynchronous Processing
  • 132Message Queues vs Event Streams
  • 133Producers, Consumers and Brokers
  • 134Consumer Groups
  • 135Message Ordering
  • 136At-Most-Once Delivery
  • 137At-Least-Once Delivery
  • 138Exactly-Once and Effectively-Once Processing
  • 139Idempotency
  • 140Retries and Exponential Backoff
  • 141Dead-Letter Queues
  • 142Backpressure
  • 143Kafka vs RabbitMQ vs Cloud Queues
  • 144Publish-Subscribe Systems
  • 145Event-Driven Architecture
  • 146Event Sourcing
  • 147CQRS
  • 148Saga Pattern
  • 149Transactional Outbox Pattern
  • 150Change Data Capture
  • 151Stream Processing
  • 152Batch Processing vs Stream Processing
Part 80/22

Distributed Systems Fundamentals

The theory that explains why the obvious design does not work across a network.

  • 153What Is a Distributed System?
  • 154Challenges of Distributed Systems
  • 155Network Partitions
  • 156Strong Consistency
  • 157Eventual Consistency
  • 158Read-After-Write and Causal Consistency
  • 159CAP Theorem
  • 160PACELC Theorem
  • 161Quorum Reads and Writes
  • 162Consensus
  • 163Leader Election
  • 164Distributed Locks
  • 165Consistent Hashing
  • 166Clock and Time Problems
  • 167Logical and Vector Clocks
  • 168Conflict Detection and Resolution
  • 169Split-Brain Problems
  • 170Gossip Protocol
  • 171Data Locality
  • 172Distributed Transactions
  • 173Two-Phase Commit
  • 174Handling Partial Failure
Part 90/22

Architecture and Scaling

Shapes a system can take, and what each shape costs to run.

  • 175Vertical vs Horizontal Scaling
  • 176Stateless vs Stateful Services
  • 177Monolithic Architecture
  • 178Modular Monolith
  • 179Microservices Architecture
  • 180Monolith vs Microservices
  • 181Service Boundaries
  • 182Domain-Driven Design
  • 183Service-Oriented Architecture
  • 184Serverless Architecture
  • 185Containers
  • 186Container Orchestration
  • 187Autoscaling
  • 188Session Management in Distributed Systems
  • 189Sticky Sessions
  • 190Multi-Region Architecture
  • 191Active-Active vs Active-Passive Architecture
  • 192Edge Computing
  • 193Cell-Based Architecture
  • 194Control Plane vs Data Plane
  • 195Multi-Tenant Architecture
  • 196Build vs Buy Decisions
Part 100/20

Reliability and Resilience

Assuming every dependency fails, and designing so the failure stays small.

  • 197Designing for Failure
  • 198Single Points of Failure
  • 199Timeouts
  • 200Retries
  • 201Circuit Breakers
  • 202Bulkhead Pattern
  • 203Graceful Degradation
  • 204Load Shedding
  • 205Rate Limiting
  • 206Throttling
  • 207Redundancy
  • 208Failover
  • 209Disaster Recovery
  • 210Recovery Point Objective and Recovery Time Objective
  • 211Backup and Restore Strategies
  • 212Idempotency Keys
  • 213Handling Duplicate Requests
  • 214Preventing Cascading Failures
  • 215Chaos Engineering
  • 216Zero-Downtime Architecture
Part 110/22

Common Specialised Systems

Named building blocks that recur across interview problems.

  • 217Content Delivery Networks
  • 218Full-Text Search
  • 219Inverted Indexes
  • 220Search Ranking
  • 221Autocomplete Systems
  • 222Recommendation Systems
  • 223News Feed Generation
  • 224Fan-Out on Write vs Fan-Out on Read
  • 225Real-Time Chat Systems
  • 226Presence and Online Status
  • 227Notification Systems
  • 228Location-Based Services
  • 229Geohashing
  • 230Payment Systems
  • 231Inventory and Reservation Systems
  • 232Distributed Job Schedulers
  • 233Web Crawlers
  • 234Media Upload and Processing
  • 235Video Streaming
  • 236Collaborative Editing
  • 237Analytics and Metrics Pipelines
  • 238Logging Platforms
Part 120/19

Security

The part candidates skip and interviewers remember.

  • 239Authentication vs Authorisation
  • 240Sessions vs Tokens
  • 241JSON Web Tokens
  • 242OAuth 2.0 and OpenID Connect
  • 243Role-Based and Attribute-Based Access Control
  • 244Encryption in Transit and at Rest
  • 245Secrets Management
  • 246API Security
  • 247Rate Limiting and Abuse Prevention
  • 248DDoS Protection
  • 249Secure File Uploads
  • 250Signed URLs
  • 251Webhook Security
  • 252Replay-Attack Prevention
  • 253Tenant Isolation
  • 254Personally Identifiable Information
  • 255Data Retention and Deletion
  • 256Audit Logging
  • 257Threat Modelling
Part 130/20

Observability and Operations

Knowing what the system is doing, and changing it without downtime.

  • 258Logs, Metrics and Traces
  • 259Structured Logging
  • 260Correlation IDs
  • 261Distributed Tracing
  • 262Important System Metrics
  • 263SLIs, SLOs and Error Budgets
  • 264Alerting
  • 265Dashboards
  • 266Health, Readiness and Liveness Checks
  • 267Capacity Planning
  • 268Performance Testing
  • 269Load, Stress, Spike and Soak Testing
  • 270Profiling and Bottleneck Analysis
  • 271Feature Flags
  • 272Rolling Deployments
  • 273Blue-Green Deployments
  • 274Canary Deployments
  • 275Rollbacks
  • 276Infrastructure as Code
  • 277Cost Optimisation
Part 140/30

Applied AI System Design

What changes when a model is in the request path: cost, latency, and non-determinism.

  • 278How AI System Design Differs from Traditional System Design
  • 279Online vs Offline Inference
  • 280Model-Serving Architecture
  • 281CPU vs GPU Inference
  • 282Model Batching
  • 283Model Gateways and Routers
  • 284Model and Provider Fallback
  • 285Prompt and Model Versioning
  • 286Streaming AI Responses
  • 287Semantic Caching
  • 288Embedding Pipelines
  • 289Vector Databases in System Design
  • 290RAG System Architecture
  • 291Document Ingestion Pipelines
  • 292Chunking and Document Versioning
  • 293Hybrid Search and Reranking
  • 294Agent Architecture
  • 295Agent Orchestration
  • 296Tool Execution and Permissions
  • 297Workflow State and Checkpointing
  • 298Human-in-the-Loop Systems
  • 299AI Guardrails
  • 300AI Evaluation Systems
  • 301Hallucination and Faithfulness Evaluation
  • 302Abstention and Confidence Handling
  • 303AI Observability
  • 304AI Cost and Latency Optimisation
  • 305Multi-Tenant AI Platforms
  • 306Data Privacy in AI Systems
  • 307Designing Reliable Agent Workflows
Part 150/19

Low-Level Design

Class-level design: the other interview, and the one that rewards precision.

  • 308What Is Low-Level Design?
  • 309Object-Oriented Design
  • 310SOLID Principles
  • 311Composition vs Inheritance
  • 312Interfaces and Abstractions
  • 313Dependency Injection
  • 314Common Design Patterns
  • 315Factory Pattern
  • 316Strategy Pattern
  • 317Observer Pattern
  • 318Adapter Pattern
  • 319Decorator Pattern
  • 320State Pattern
  • 321Repository Pattern
  • 322State Machines
  • 323Domain Modelling
  • 324Concurrency and Thread Safety
  • 325Extensible API and Class Design
  • 326Testing a Low-Level Design
Part 160/29

System Design Case Studies

Full problems, worked end to end. This is where the earlier parts get used.

  • 327Design a URL Shortener
  • 328Design a Rate Limiter
  • 329Design Pastebin
  • 330Design a Notification System
  • 331Design a Distributed Job Scheduler
  • 332Design a Web Crawler
  • 333Design Cloud File Storage
  • 334Design a Logging System
  • 335Design a Metrics Platform
  • 336Design a Social Media Feed
  • 337Design a Chat Application
  • 338Design a Video-Streaming Platform
  • 339Design a Ride-Sharing Platform
  • 340Design an E-Commerce Platform
  • 341Design a Payment Platform
  • 342Design a Ticket-Booking System
  • 343Design a Collaborative Document Editor
  • 344Design a Search Engine
  • 345Design an Autocomplete Service
  • 346Design a Multi-Tenant SaaS Platform
  • 347Design an API Gateway
  • 348Design a Distributed Cache
  • 349Design a Message Broker
  • 350Design a ChatGPT-Like Application
  • 351Design a Multi-Tenant RAG Platform
  • 352Design an AI Coding Assistant
  • 353Design an AI Document-Processing Platform
  • 354Design an Agent-Workflow Platform
  • 355Design an AI Evaluation Platform
Part 170/10

Low-Level Design Exercises

The classic object-modelling problems, each small enough to finish in one sitting.

  • 356Design a Parking Lot
  • 357Design an Elevator System
  • 358Design a Vending Machine
  • 359Design a Library System
  • 360Design a Logger
  • 361Design a Task Scheduler
  • 362Design a Notification Framework
  • 363Design a Payment State Machine
  • 364Design a Ride-Booking System
  • 365Design a Board Game
Part 180/11

Final Interview Preparation

Rehearsal. A repeatable template and the drills that make it automatic.

  • 366Creating a Reusable Whiteboard Template
  • 367Practising Requirement Clarification
  • 368Practising Capacity Estimation
  • 369Practising APIs and Data Models
  • 370Practising Critical Request Flows
  • 371Practising Failure Scenarios
  • 372Practising Trade-Off Explanations
  • 373Conducting Timed Mock Interviews
  • 374Handling Follow-Up Questions
  • 375Evaluating Your Own Performance
  • 376Final System Design Interview Checklist