Back-of-the-Envelope Estimation

Before drawing boxes and arrows, good system designers do some quick arithmetic. How many requests per second? How much storage per year? Does the hot data fit in memory? How many servers, roughly? Back-of-the-envelope estimation answers these to within an order of magnitude in a few minutes, and that's usually enough to decide whether you need one database or fifty, a cache or not, a CDN or not.

It's a core part of system design interviews, and just as valuable in real design reviews and capacity planning. The skill is making reasonable assumptions, stating them explicitly, rounding aggressively, and checking results against intuition and known numbers.

TL;DR

Quick Example

Estimating a photo-sharing service with 50M daily active users:

Conclusions follow directly: object storage plus a CDN for images, a partitioned metadata store, and a modest cache tier.

Core Concepts

Latency Numbers Every Engineer Should Know

Approximate, modern hardware:

Takeaways: memory is ~1,000× faster than SSD random access, network round trips dominate request latency, and cross-region calls are expensive, so avoid them in synchronous paths.

Powers of Two and Units

Useful sizes: a UUID is 16 bytes (36 as a string), a timestamp 8 bytes, a typical tweet-sized text row ~0.3–1 KB, a compressed photo ~200 KB–2 MB, and a minute of compressed HD video ~10–50 MB.

Time Conversions

Estimating Traffic

  1. Start from users: DAU × actions per user per day.
  2. Divide by 10⁵ for average QPS.
  3. Multiply by a peak factor: 2–3× for global smooth traffic, 5–10× for spiky consumer apps, far more for flash events.
  4. Separate reads and writes, whose ratio drives caching and replication design.

Estimating Storage

storage = new items/day × size per item × retention days × replication factor × overhead

Include indexes (often 20–100% of data size in databases), metadata, and growth. Distinguish hot data (in memory or SSD) from cold (object storage, archive tiers).

Estimating Servers

Rough per-server throughputs (they vary enormously by workload):

servers ≈ peak QPS ÷ per-server capacity × headroom (1.5–2×), plus redundancy for failures.

Best Practices

State Assumptions Explicitly

Write down every assumption (DAU, actions per user, sizes, ratios). Assumptions are where estimates go wrong, and listing them invites correction. In interviews, confirm them with the interviewer.

Round Aggressively, Then Sanity-Check

Use 10⁵ seconds per day, 1 KB rows, and similar round numbers. After computing, ask: does 120 Gbps of egress sound plausible for this product? Does 6 PB per year match similar services? Is the implied cost reasonable?

Compute What Drives Decisions

Focus on numbers that change the architecture: peak write QPS (can one primary handle it?), hot data size (fits in memory?), egress (needs a CDN?), and storage growth (when to shard or tier?).

Revisit With Real Metrics

Estimates are starting points. Replace them with measured data from production metrics and load tests as soon as possible. See capacity planning and performance testing.

Common Mistakes

Designing for Average Instead of Peak

Average QPS hides daily peaks and bursts. Capacity must handle peak load (plus failure headroom), or the system falls over precisely when it's busiest.

Forgetting Replication and Overheads

Raw data size × 1 underestimates real storage by 3–5× once you add replicas, indexes, backups, and free space headroom.

False Precision

Reporting "13,888.9 QPS" implies confidence you don't have. Say "~15k QPS average, ~75k peak". The order of magnitude is what matters.

FAQ

How accurate do these estimates need to be?

Within an order of magnitude, sometimes within 2–3×. The goal is to choose the right architecture class (one server vs a fleet, database vs object storage, cache or not) and to spot infeasible designs early, not to predict exact capacity.

How do I convert daily requests to QPS quickly?

Divide by 10⁵ (≈ seconds per day). One million requests per day is roughly 10–12 per second on average, and one billion is roughly 10,000–12,000. Then apply a peak multiplier.

What read/write ratio should I assume?

Consumer content apps are heavily read-dominated (10:1 to 1000:1). Logging, telemetry, and IoT ingestion are write-dominated. Transactional business apps are often in between. Ask or reason from user behavior, since the ratio drives caching and replication choices.

Which latency numbers matter most?

Memory vs SSD vs network orders of magnitude, and cross-region round-trip times. They explain why caching in memory helps, why chatty service-to-service calls add up, and why synchronous cross-region calls should be avoided.

Related Topics

References