👋 Hey {{first_name|there}},
Four issues on knowing what you depend on and whether it's working. This one is about the moment someone has to write it down for a regulator, and the clock governing that is stranger than most people assume.
Why this matters
Most engineering leaders can tell you DORA gives them four hours. Fewer can tell you four hours from what.
It isn't detection, and it isn't containment, which is the milestone almost every incident process I've seen is organised around. The initial notification is due within four hours of the incident being classified as major, with a backstop of twenty-four hours from the point you became aware of it.
Classification is a judgement. Somebody in your organisation looks at a thing that is currently on fire and decides it meets the threshold, and at that moment a regulatory clock starts running.
So the question worth asking isn't really about report-writing speed. It's how quickly you can reach a defensible decision about severity while the incident is still going on.
That decision gets made against seven criteria in the classification standard. Clients affected, geographic spread, data losses, criticality of the services involved, duration, economic impact, reputational damage. Two or more, and it's major.
Look at that list as an architect, and at least four of them are questions about data your systems either produce or don't. Number of clients affected. Geographic spread. Data losses. Duration measured against your recovery objective.
If those take a day and a half to assemble, the reporting deadline is the least of it. The incident can't be classified in any useful timeframe, and everything downstream of classification inherits the delay.
Somebody will eventually ask when you knew it was major. The honest version of that answer is usually a story about spreadsheets.
🧭 The shift
From: "We have four hours to report a major incident."
To: "We have four hours from the moment we can prove it was major, and proving it needs numbers we may not have."