Interview prompt
Explain reliability comes from explicit failure behavior to an engineer who understands the surrounding system but has not used this technique. Walk from its contract to a concrete operation, then discuss where it fails or becomes expensive.
A strong answer
An unreliable component does not make a system unreliable if the system contains and recovers from its failure. Reliability design starts by naming failure modes, their blast radius, detection signals, and the behavior users should see when a dependency is unavailable.
If a recommendation service fails, a commerce application might continue checkout without recommendations. That fallback preserves a core transaction while degrading a secondary feature. A timeout bounds waiting; a circuit breaker can stop repeated calls during an outage and permit controlled recovery probes.
A complete answer also calls out the assumptions that control correctness. Redundancy helps only when failure modes are sufficiently independent and failover is tested. A fallback can return incorrect or unsafe data if its contract is vague. Availability targets also need a measurement window and an explicit definition of a successful request.
Close by describing one representative test or measurement. For a service that depends on a payment provider and a recommendation service, classify which outage should block checkout and what degraded mode is safe for each dependency.
Follow-up questions
Answer the follow-ups in the frontmatter. Use the linked article for the concept and the trace to make the explanation concrete.