- Delta Air Lines CEO Ed Bastian said the massive IT outage earlier this month that stranded thousands of customers will cost it $500 million.
- The airline canceled more than 4,000 flights in the wake of the outage, which was caused by a botched CrowdStrike software update and took thousands of Microsoft systems around the world offline.
- Bastian, speaking from Paris, told CNBC’s “Squawk Box” on Wednesday that the carrier would seek damages from the disruptions, adding, “We have no choice.”
Only for the sake of specific-ness: Crowdstrike forced the update, not the OS. :) and yeah, that’s generally unheard of. Like so unheard of that it’s a professional recommendation reversing occurrence based purely on how they could release a product that bypassed user expectations so aggressively and without any documentation that it was happening.
I work in the security sector with computers, and before all this I would have said “yeah, crowdstrike is a widely deployed product and if it fits your requirements it’s reasonable to use”. Now I would strongly recommend against it, not because of this incident, but because of the engineering, product and safety culture that thought it was okay to design a product this way without user controls or even documentation around any part of it. Their after incident report is horrifying in testing it communicates they weren’t doing.
I wouldn’t advise someone to use windows for a server, but that’s a preference thing, not a “hazard” thing. If they had a working windows setup I wouldn’t even comment on it.
What sounds like happened to Delta is that they were set-up roughly like other companies. Maybe a little loose on different setups at different airports. That’s a forgivable level of slop. Where they differed was in having a piece of software that couldn’t handle being entirely shut off, and then immediately loaded to 100% with no ease in.
Scheduling is a type of computer problem that’s very susceptible to getting increasingly difficult the bigger the number of things being worked with. Like exponentially more difficult, but it’s actually worse than exponential.
I know nothing about they’re system, but I can guess that it worked fine when it was running because it needed to make a small number of scheduling decisions at a time, and could look at the existing state of things as a decided “fact”. Start the system fresh, and suddenly it needs to compare the hundreds of airports, more hundred of planes and crews, and thousands of possible routes to each other and is looking at literally billions of possible schedules which it needs to sort through to pick the best ones.
Other airlines appear to have scheduling systems that were either developed using more modern techniques that can find “good enough” very efficiently, or the application was written to fail less easily or had better hardware so it could work faster.
For whatever reason, delta was the only one that had the key bit of software fail to come back up.
Delta has higher costs than the other airlines because there are regulations protecting travelers and ensuring they get appropriate refunds and accomodations if their flights are cancelled. Other airlines were able to shift people around and get going again before they had to pay out too much in ticket refunds, food, or hotels.
Delta is arguing that crowdstrike is responsible for the total cost of the incident, which would include all the refunds and hotels, since they caused it.
Crowdstrike recently responded that they think their liability is no greater than $10mil. They seem to be taking the position that they’re only responsible for the immediate effects, so things like diverting aircraft, needing to manually poke systems and all that.
“Yeah I t-boned you when I ran a red light, so I owe you for the damage to your car, but your car was a dangerous piece of crap so I’m not responsible for your broken legs, hospital bills or lost wages”.
I think the judge will find that running the red light means they are responsible for the extended consequences of their actions, even if they’re vastly in excess of what anyone would have predicted up front, but that the car was pretty dangerous so it was really only a matter of time so it’s not all on them.
If there’s one thing I’ve learned from reading about court cases, it’s that a civil suit like this will get really complicated with how they assess damages and responsibilities.
And yeah, there’s no perfect answer for computer system stability. You can never get perfect stability, and each 9 you add to your 99.9% uptime costs more than the last one. Eventually you have teams of people whose full time job is keeping the system up for an additional second per year. And even with that, sometimes Google still goes down because it’s all a numbers game.
I didn’t mean to ramble so long, but I have opinions and I get type-y before bed. :)