The Runtime Theory
Workload and Capacity

Start System Design With a Workload Model

A workload model describes request rates, payload sizes, read/write mix, burstiness, and latency goals.

The Runtime Theory Team5 min read#capacity-estimation#workload#queues
▸ On this page

The model

A workload model describes request rates, payload sizes, read/write mix, burstiness, and latency goals. Capacity estimates are useful when they expose assumptions and identify dominant resource costs; a single total-user count rarely predicts the load a system must handle.

A concrete walk-through

If one million users each make two requests per day, the average is far below the peak if activity clusters around a short window. Estimate average and peak requests per second, then multiply by bytes, storage retention, and downstream calls. State whether retries and background jobs are included.

Costs and failure cases

Back-of-the-envelope numbers are ranges, not promises. Cache hit rate, hot keys, payload distribution, and fan-out can change capacity sharply. Averages hide tail latency and bursts, so design headroom and measure the real workload before purchasing or sharding capacity.

Check your understanding

Estimate storage for event records given an arrival rate, average encoded size, and retention period. Then name two additional factors needed before sizing the database.

Further reading

AWS Well-Architected Framework

Not started

Sign in to save your learning progress.

Sign in to save