👋 Hey {{first_name|there}},
You did the responsible thing. Five percent of users, two weeks, dashboards up, nobody complained. So it went to everyone, and that's when it started going wrong. The percentage was never the thing protecting you.
Why this matters
Progressive delivery is one of the few practices almost everyone agrees on. Ship to a slice, compare the slice to everyone else, and if error rate and latency hold, widen it. The comparison is what makes it work. You need a number that moves when something is wrong.
Last issue was about how this feature doesn't produce that number. It answers with a 200 status, it answers fast, and when it's wrong, it's wrong politely. Point a canary analysis at error rate and latency, and you'll get a green light on a feature that's producing confident nonsense for a quarter of the people who see it. The instrument that made staged rollout trustworthy is pointed at the wrong thing.
Selling B2B software adds a wrinkle on top of that. In a multi-tenant product, your customers don't differ mainly in how they behave. They differ in what their data looks like. One tenant has tidy structured records entered by a trained ops team; another has fifteen years of scanned PDFs and half-empty fields. The feature that works beautifully for the first can fall apart for the second, and no amount of watching a random slice of users will surface that, because a random slice is mostly your easy tenants. There are more of them.
🧭 The shift
From: "Roll it out to five percent and watch the error rate."
To: "Roll it out to tenants we chose on purpose, and watch what people do with the answers."
Both halves change. Who gets it stops being a dice roll and becomes a decision. And what you monitor stops being whether the system worked and becomes whether the output was any good, which the system cannot tell you and the user can.