On Monday evening, Outlook and Exchange Online stopped working. Not for one company, for a meaningful slice of the world. People couldn’t log in. Messages sat unsent. Search returned nothing. IT teams spent hours fielding calls from colleagues who couldn’t do their jobs, because the one tool most of us touch fifty times a day had simply gone quiet.
In England, Wales and Northern Ireland, that Monday was the late-summer bank holiday, so plenty of UK inboxes were quiet anyway and the practical impact here was softer than the headlines suggested. It’s worth pausing on that rather than skipping past it. The outage didn’t observe the UK holiday any more than it observed a weekend. Offices across the US, most of Europe and large parts of Asia were open and working as normal, and whatever incident response Microsoft was running that day, it was almost certainly running with fewer people watching from this side of the Atlantic. A skeleton crew on a holiday Monday is a smaller version of exactly the problem this piece is about: fewer eyes on the system, at the moment something in that system needed eyes on it.
When Microsoft explained what had happened, the answer was almost disappointing. No breach, no attacker, no sophisticated failure of engineering. Somewhere inside their infrastructure sat a certificate. It had a renewal date. Nobody renewed it in time. One missed date, and a piece of the internet’s email backbone went dark for hours.
There was nothing wrong with the technology
That’s the detail that matters most. This wasn’t a clever attack that exploited a weakness nobody had thought of. It was routine housekeeping that didn’t happen. And it didn’t happen inside one of the most sophisticated technology operations on the planet, run by a company with more engineers and more money than almost anyone else in the business of keeping systems alive.
Modern IT isn’t one system. It’s thousands of small, interlocking parts: certificates, API keys, service accounts, DNS records, automated renewal jobs checking other automated renewal jobs. Each one has its own quiet expiry date sitting somewhere in a config file or a calendar nobody quite owns. Most of the time, the automation works, the renewal happens, and nobody ever has to think about it. That’s precisely the problem. A control that runs silently and correctly for years earns no attention at all, right up until the one time it doesn’t run, and by then it’s not a maintenance task anymore, it’s an incident.
This keeps happening, and that’s the actual point
If this were a one-off, it would be a curiosity. It isn’t. In December 2018, an expired software certificate inside Ericsson’s network management tools took down mobile data and SMS for O2’s roughly 32 million UK customers for most of a day, with SoftBank in Japan hit by the same fault at the same time. In February 2019, an expired authentication certificate locked around 20 million people out of Microsoft Teams for three hours. And in the Equifax breach that exposed the data of some 147 million people, investigators found that a certificate on a device meant to inspect encrypted network traffic had been left expired for nineteen months, quietly blinding the very monitoring tool that might have caught the intrusion sooner.
Three different organisations, three different sectors, spread across almost a decade, all undone by the same unglamorous cause. Every one of them had the technical knowledge to prevent it. None of that knowledge mattered on the day it counted, because certificate expiry isn’t a skills problem. It’s an attention problem, in systems that have become too large and too interconnected for any one person, or often any one team, to hold the whole picture in their head.
Why this should worry business continuity people specifically
Here’s the bit that gets missed when this story is told as a Microsoft story rather than a resilience one. For a great many organisations, the plan for what to do when IT goes down itself depends on IT being up. Specifically, on email. The incident notification, the crisis bridge invite, the “we’re aware and investigating” message to staff and customers: all of it routed through the very inbox that has just stopped working.
It’s a gap I come across regularly when working through incident plans line by line: the response team gets activated properly, then everything downstream of that first call reverts to email by default, with no real fallback thought through for what happens if email itself is the thing that’s broken. Nobody plans for that scenario because it feels unlikely, right up until an unremarkable Monday evening when it’s exactly what happens, to millions of people at once.
The question worth asking
For anyone assessing a supplier, the useful question is no longer just “are you certified”. A certificate on the wall tells you a supplier passed an audit on a given day. It says nothing about whether the unglamorous, ongoing renewal work behind the scenes is actually being done, or who is watching it. Ask instead how they manage the lifecycle of certificates, keys and credentials, and whether they’ve ever tested what happens when that management itself fails.
Internally, the exercise is simpler and worth doing this week regardless of what any supplier says. If your primary means of telling people there’s a problem is itself down, what’s the second channel, and has anyone on the team actually used it, rather than just written it into a document and assumed it would work when needed.
Nobody can promise the next certificate won’t slip past whoever’s calendar it’s sitting in. Systems this complex will keep producing small, boring failures in places nobody was watching closely enough. What’s within anyone’s control is simpler: make sure that when it happens, it isn’t also the only way you had of finding out.



