9 rules for system design

System design is all trade-offs. There is no perfect choice, only a choice that fits the problem and the failure we can accept. I use these rules to keep a review focused on those decisions.

1. Start with questions

Before drawing boxes, ask what kind of system this is. Is it read-heavy or write-heavy? How long must it retain data? What guarantees does it need? How fresh must reads be, and how much replication lag is acceptable?

Then ask about traffic. Is the load steady or bursty? After a failure, should the system recover in seconds, minutes, or hours? The answers shape the design more than a familiar list of components does.

2. Ask about scale first

Do not start by listing non-functional requirements in the abstract. Ask for scale. Requests per second, payload size, stored data, growth rate, and peak traffic turn words such as "fast," "available," and "durable" into constraints we can reason about.

Without scale, every requirement sounds important and every technology sounds plausible. With scale, some options become obviously sufficient and others obviously wasteful.

3. Walk the storage

My note writes a shorthand sequence: process memory -> Redis -> persistent database -> blob storage. These are not interchangeable steps in a fixed ladder. I use the sequence as a prompt to compare workload-specific options: local memory, a cache or key-value store such as Redis, a transactional database, and object storage.

Each step changes latency, durability, capacity, cost, and operational work. Run the numbers before moving right. Estimate how much data exists, how quickly it grows, how often it is read, and what it costs to keep. A design review should show why the chosen point fits, not just name a popular database.

4. Treat durability as verification

Three replicas do not automatically mean durable data. They may faithfully copy corruption, a bad write, or an application mistake.

Durability also means checking that the copies are correct. Checksums are one example. Backups, restores, and repair paths need the same attention. The real question is not "How many copies exist?" It is "Can we prove the data is intact and recover it when a copy is wrong?"

5. Start simple, especially across networks

A local call is not the same as a network call. Crossing a network boundary adds latency, partial failure, and retry danger. The caller can time out while the receiver still completes the work, leaving both sides uncertain about what happened.

Start with fewer boundaries when the requirements allow it. Split a system when the trade-off becomes worthwhile, not because distributed architecture looks more serious. Every new boundary needs a timeout, retry policy, idempotency story, and recovery path.

6. Do not forget the client

The server is not the only place to improve a data path. A client may batch small requests, split large work into chunks, or compress payloads before sending them.

These choices can reduce request overhead and smooth load, but each has a cost. Batching can add delay. Chunking needs progress tracking and retry logic. Compression trades network bytes for CPU. Include the client in the design and make those trade-offs explicit.

7. Accept that some work is not atomic

Some business processes cannot happen as one database transaction. Payments are a clear example: authorization, capture, settlement, refund, and dispute are separate stages. Any stage can fail, arrive late, or need a retry or compensation.

This is where saga-like workflows appear. Instead of pretending the whole process is atomic, record each step, make retries safe, and define what compensating action means. Correctness comes from managing the sequence and its failures.

8. Optimize the read path for speed

The read path usually asks: how quickly can we return an acceptable answer? Caches, replicas, indexes, and precomputed results can help, depending on the freshness requirement.

But speed is conditional. A fast stale read may be fine for one feature and wrong for another. Decide what "acceptable" means before optimizing the path.

9. Optimize the write path for correctness

The write path asks a different question: did we accept the right change exactly as intended? Validation, ordering, deduplication, idempotency, and durable acknowledgement matter here.

A slow read is visible. A wrong write can spread quietly through replicas, caches, and downstream processes. That is why I bias read decisions toward speed and write decisions toward correctness. Not as a universal law, but as a useful default until the requirements say otherwise.

A good system design review does not end with the most sophisticated diagram. It ends with clear trade-offs: what the system promises, where it can fail, and why this is the simplest design that meets the numbers.