The Runtime Theory
mediumSystemInternals#explain-the-model#reason-about-tradeoffs

Explain Reliability Comes From Explicit Failure Behavior

Explain the model, execution steps, complexity, and limits of reliability comes from explicit failure behavior.

TRT practice prompt — not a verified question from a named employer.

The Runtime Theory Team6 min read

Interview prompt

Explain reliability comes from explicit failure behavior to an engineer who understands the surrounding system but has not used this technique. Walk from its contract to a concrete operation, then discuss where it fails or becomes expensive.

A strong answer

An unreliable component does not make a system unreliable if the system contains and recovers from its failure. Reliability design starts by naming failure modes, their blast radius, detection signals, and the behavior users should see when a dependency is unavailable.

If a recommendation service fails, a commerce application might continue checkout without recommendations. That fallback preserves a core transaction while degrading a secondary feature. A timeout bounds waiting; a circuit breaker can stop repeated calls during an outage and permit controlled recovery probes.

A complete answer also calls out the assumptions that control correctness. Redundancy helps only when failure modes are sufficiently independent and failover is tested. A fallback can return incorrect or unsafe data if its contract is vague. Availability targets also need a measurement window and an explicit definition of a successful request.

Close by describing one representative test or measurement. For a service that depends on a payment provider and a recommendation service, classify which outage should block checkout and what degraded mode is safe for each dependency.

Follow-up questions

Answer the follow-ups in the frontmatter. Use the linked article for the concept and the trace to make the explanation concrete.

This answer walks

Practice follow-ups

  1. 01Which assumption is essential for the approach to be correct?
  2. 02What is the worst case, and how does it change the resource cost?
  3. 03How would you adapt the design if the input or workload became much larger?
  4. 04What boundary test would give you the most confidence in the implementation?

One dispatch a week

The trace behind each question, the tradeoff that explains it, and one technical dispatch per week — no noise.

One technical dispatch per week. No noise.

Not started

Sign in to save your learning progress.

Sign in to save