The Runtime Theory
Partial Failure and Timeouts

Distributed Calls Can Fail Halfway Through

In a distributed system, a caller may lose contact with a process that is still running, or receive no response after the remote side completed its work.

The Runtime Theory Team5 min read#partial-failure#timeouts#observability
▸ On this page

The model

In a distributed system, a caller may lose contact with a process that is still running, or receive no response after the remote side completed its work. A network partition, overloaded queue, crashed process, and slow dependency can look similar from the caller’s vantage point.

A concrete walk-through

A client starts a request with a deadline. If the deadline expires, it cancels local waiting and reports uncertainty; cancellation may or may not stop remote work. Propagating the remaining deadline downstream prevents each hop from independently consuming a full timeout budget.

Costs and failure cases

No timeout means resources can remain tied up indefinitely; an overly short timeout creates false failures and duplicate retries. Logs need request or trace identifiers to correlate attempts, but identifiers do not guarantee exactly-once execution. Design operations to tolerate ambiguous completion.

Check your understanding

A worker times out while charging a card, then retries. Explain why the first charge may have succeeded and how an idempotency record can resolve the ambiguity.

Further reading

Google SRE Book: Addressing Cascading Failures

Not started

Sign in to save your learning progress.

Sign in to save