The Runtime Theory
Failure Modes and Reliability

Reliability Comes From Explicit Failure Behavior

An unreliable component does not make a system unreliable if the system contains and recovers from its failure.

The Runtime Theory Team5 min read#reliability#availability#recovery
▸ On this page

The model

An unreliable component does not make a system unreliable if the system contains and recovers from its failure. Reliability design starts by naming failure modes, their blast radius, detection signals, and the behavior users should see when a dependency is unavailable.

A concrete walk-through

If a recommendation service fails, a commerce application might continue checkout without recommendations. That fallback preserves a core transaction while degrading a secondary feature. A timeout bounds waiting; a circuit breaker can stop repeated calls during an outage and permit controlled recovery probes.

Costs and failure cases

Redundancy helps only when failure modes are sufficiently independent and failover is tested. A fallback can return incorrect or unsafe data if its contract is vague. Availability targets also need a measurement window and an explicit definition of a successful request.

Check your understanding

For a service that depends on a payment provider and a recommendation service, classify which outage should block checkout and what degraded mode is safe for each dependency.

Further reading

Google SRE Book

Not started

Sign in to save your learning progress.

Sign in to save