Posted by asplake 4 days ago
But we're talking about a system that is also older than most people.
It was first put into operation in 1967 in the U.S.A., and brought over to the U.K. in 1974. It's formally called 'NAS En Route Stage A', and is written in a 1950s language named JOVIAL. It originally ran on an IBM 9020. It runs on IBM 9020 compatible systems today.
For some reason, a couple of years ago someone added a lengthy unverifiable description of it to Wikipedia's IBM 9020 article, even though that very description said that NAS hadn't run on a 9020 for 34 years at the time of writing, but had been running on a 4381.
> While this request was being processed, the NAS received a message for a higher priority activity to be undertaken which resulted in the squawk code allocation being paused while the system processed the higher priority message. Switching between different activities in response to prioritised requests is a normal function of the system; however, when the processing of the squawk allocation request resumed, the software defect meant it did not resume correctly and the resulting output was corrupted.
> The reason this scenario has not occurred before is because:
> 1. The defect existed in a specific subsection of code within a software module, with an exposure window estimated as approximately one millisecond.
> 2. For the fault to occur, a higher-priority request had to arrive during that exact millisecond while the original request was part-way through updating a value.
> 3. Had the higher-priority request arrived even one millisecond earlier or later, the update would have completed normally.
> Post-incident investigation has identified that when processing of the squawk allocation request resumed, the data associated with it had been corrupted and affected some subsequent flight data updates.
How often does that original request happen per day? How often does the higher priority activiry take?
If the original request happens 864 times a day and the high priority request ten times, there's a 1 in 25 chance it will happen in a given year.
I assume the "manual request" is an aircraft squawking 7700 or similar, but why does the system need to interrupt an in-flight allocation in the first place? Any controllers here have insight?
One would think it would be sufficient to do something single threaded like
if(!highPriorityQueue.empty() {
highPriortyQueue.processOne();
} else if(!lowPriorityQueue.empty()) {
lowPriorityQueue.processOne();
}
or whatever, but they're not and I'm curious why.if they could guarantee hard bounds on how long low priority tasks take to complete they could implement your scheduling algorithm. but i think in reality the low priority tasks are either not boundable or they if they do have a provable bound the bound is too high.
It doesn’t make sense to have other non-code allocation things competing for queue space with code allocations.
Well, that's comforting to know.
BBC: "Flight chaos caused by software defect in space of a millisecond, report says"
Sky: "'Millisecond' software error caused air traffic outage that grounded thousands of flights"
The Guardian: "Flight chaos for hundreds of thousands was caused in ‘millisecond’ by software error"
Sounds like pure bad luck.
Maybe I'm being too harsh.. on the plus side the system has at least failed hard every time there's been a fault. Nobody has died. But it's been 3 times now in the past couple of years, and two of those times resulted in over 2000 flights cancelled and days of backlog, and misery for hundreds of thousands. It's really not acceptable.
You don't know that.
Seriously, for 2026 this is pure amateur hour with no excuse.
Considering the last issue they encountered, it looks like in more than one place, there is no error catching and graceful resolution for those errors.
I would assume a system of such importance to handle issues without hiccups and alert the operators of what did not work. Like “this input caused this problem”, not just crash.
The fact that it entered in maintenance mode still ruined a lot of people’s days.
I don’t think any flight crew or passenger cared about semantics back then.