The Fault Lasted a Millisecond. The Recovery Took Days.

On 8 September, a defect in the software that allocates radar identification codes to aircraft over the UK lasted, by NATS’ own account, a matter of milliseconds. Controllers lost the ability to tell which blip on their screens belonged to which flight without switching to manual coordination, and restrictions went on across the whole of UK airspace while the system was restarted and reloaded. Over 2,000 flights were affected, and it was more than two days before the last of the displaced aircraft, crews and passengers were back where they were supposed to be.

NATS published its preliminary report on 21 September. It found no evidence of any incorrect action by controllers, and the UK’s air traffic chief had already ruled out a cyberattack within days of the incident. Chief executive Martin Rolfe has been clear about the one thing he isn’t apologising for: ‘At no point last week was safety in question,’ he said, ‘our primary role is to keep our skies safe, and everyone who flies through them.’ Every aircraft already in the air when the fault hit landed safely, which is a genuine point of credit for NATS and its controllers. It is also, on its own, a different claim from saying the organisation recovered quickly.

What actually happened

The fault sat in the flight data system at the London Area Control Centre, in the part of the system responsible for allocating the codes that let controllers match a radar return to an actual aircraft. Lose that link and every controller in the room is suddenly working with less information, at higher workload, needing more coordination with neighbouring air traffic units to keep aircraft safely separated. Restrictions were in place for around six hours while the system was restarted and reloaded, and NATS says its own operations were back to normal that same evening.

The MBCO the report doesn’t mention

Continuity planning doesn’t run on a single Recovery Time Objective. Properly done, it runs on a Minimum Business Continuity Objective, the lowest level of service an organisation has decided it can operate at during a disruption, and often more than one MBCO for different scenarios or different parts of the business, with a set of RTOs underneath describing how quickly each activity, system and resource needs to be back to support it. The flight data system had an RTO measured in hours: restrictions lifted, data reloaded, normal service resumed by the same evening. But the level of service that airlines, crews and passengers needed before anyone could reasonably call the network recovered, its MBCO, depended on a longer chain of RTOs underneath that: aircraft repositioned, crews back within their duty hours, passengers rebooked. Judged against those, recovery took over two days. NATS’ preliminary report is detailed about the RTO it met. It says very little about the RTOs, plural, that actually determined how long this went on for everyone outside the control room, and nothing at all about what its data was reloaded to, or whether any of it needed manual reconciliation afterwards, which is its own Recovery Point Objective question for a system whose job is knowing which aircraft is which in real time.

A fallback that had never carried this much weight

Manual coordination between air traffic units is a documented procedure, not an improvisation. But there is a real difference between a fallback that has been walked through in a tabletop exercise and one that has just carried the country’s full live air traffic for hours under pressure. It is easy to write ‘single point of failure’ into a report and move on. This genuinely was one: there was no alternative route into the flight data those controllers needed, only a slower, harder way of working without it. Until a fallback like that is asked to run at full volume, for real, nobody actually knows whether it can hit its own RTO either.

Who actually picks up the bill

Ryanair has already called for Martin Rolfe to resign. That’s a reasonable position for an airline that lost money over the disruption to take, and accountability for a failure of this scale matters. But it’s worth separating two different questions the report raises: who is responsible for the fault, and who pays for the two days of disruption it caused.

The UK Civil Aviation Authority treats air traffic control failures like this one as ‘extraordinary circumstances’ under UK261 rules, which exempts airlines from the fixed compensation they would otherwise owe passengers for delays and cancellations. Airlines still owe passengers care: meals, hotels, rerouting. Claiming it back means keeping receipts and filing a claim, not receiving it automatically. Travel insurance fills some of the gap, unevenly: some policies cover air traffic control failures explicitly, plenty don’t, and two passengers on the same cancelled flight can end up with different outcomes depending on which policy they happened to buy.

None of that cost sits with NATS. If the honest planning assumption is that a fault like this one takes days, not hours, to clear once the whole network is accounted for, that’s a reasonable thing for NATS to plan around internally. It’s a harder position to defend in public while airlines, insurers and passengers are the ones actually carrying the cost of a recovery timeline NATS hasn’t had to put a number on.

The question for your own risk register

Every organisation with a business continuity plan has RTOs set for its critical systems, suppliers and processes, and an MBCO, or more than one, describing what ‘recovered’ actually needs to look like. NATS’ report is a useful prompt to check whether those RTOs, added together, actually deliver your MBCO in the time you’ve told people to expect, or whether there’s a gap between the critical system coming back online and the level of service your customers, regulators or board would recognise as recovered. Do you know what that second timeframe really is, not the one that looks best in a report? And if it’s longer than anyone would like, have you been clear with the people who bear the cost of that gap, whether that’s customers, insurers or your own board, about what to expect?

What this report is really useful for isn’t ruling out a cyberattack, or settling who should carry the can for it. Both matter less than the prompt it gives every other organisation: go and find out whether your own RTOs actually add up to the MBCO you’ve promised, and whether anyone outside the building would recognise that as recovered.

Share the Post:
Helen Molyneux, founder of Cambridge Risk Solutions, ISO 22301 and ISO 27001 Lead Auditor

Helen Molyneux is the founder and director of Cambridge Risk Solutions. A certified Lead Auditor for ISO 22301 and ISO 27001, she has spent nearly two decades helping organisations across the public and private sectors build genuine resilience — not just documented compliance. She writes from practice, not theory.

Work with us →