Interview prompt
Explain distributed calls can fail halfway through to an engineer who understands the surrounding system but has not used this technique. Walk from its contract to a concrete operation, then discuss where it fails or becomes expensive.
A strong answer
In a distributed system, a caller may lose contact with a process that is still running, or receive no response after the remote side completed its work. A network partition, overloaded queue, crashed process, and slow dependency can look similar from the caller’s vantage point.
A client starts a request with a deadline. If the deadline expires, it cancels local waiting and reports uncertainty; cancellation may or may not stop remote work. Propagating the remaining deadline downstream prevents each hop from independently consuming a full timeout budget.
A complete answer also calls out the assumptions that control correctness. No timeout means resources can remain tied up indefinitely; an overly short timeout creates false failures and duplicate retries. Logs need request or trace identifiers to correlate attempts, but identifiers do not guarantee exactly-once execution. Design operations to tolerate ambiguous completion.
Close by describing one representative test or measurement. A worker times out while charging a card, then retries. Explain why the first charge may have succeeded and how an idempotency record can resolve the ambiguity.
Follow-up questions
Answer the follow-ups in the frontmatter. Use the linked article for the concept and the trace to make the explanation concrete.