👋 Hey {{first_name|there}},
Last week we walked one journey and wrote down the dependencies that weren't on anyone's register. This week is what you do with that page. Shorter exercise, harder question.
Why this matters
There's a sentence that gets said about ten minutes into a lot of incidents, usually with some relief, because it means the person on call isn't the one who broke it.
"It's not us, it's them."
Which is true, and doesn't help. The provider is down. So what is the customer looking at right now, and what were they meant to be looking at?
Most teams can't answer the second half of that. Not through carelessness; it's that nobody ever decided. Behaviour under third-party failure is hardly ever designed for. It emerges, out of whatever the timeout values and the retry policy and the error handler happen to add up to, and nobody has ever sat down and looked at the sum.
So the customer gets a spinner that never resolves. Or a 500 with a support ID in it. Or, occasionally, a confirmation for something that didn't happen.
The regulatory position on this settled some time ago, and the engineering position has been slowly catching up. You can't outsource the customer relationship, no matter what the contract says about the infrastructure.
Somebody will ask what our customers experienced during the outage. "The provider was down" answers a different question than the one being asked.
🧭 The shift
From: "That failure was outside our control."
To: "That failure was outside our control, and this is the experience we designed for it."
One of those is about blame. The other is about design.