The software update that took down eight and a half million machines this fortnight, grounding aircraft and freezing hospitals across the planet, is being explained as a quality-control failure: a bad file that better testing would have caught. This is the reassuring story, because it implies a simple remedy, be more careful, and it is wrong, and its wrongness matters, because the remedy it implies will not work and the real one is being ignored. The outage was not a lapse in an otherwise sound system. It was the predictable output of a system built to produce exactly this, and to see why requires setting aside the language of blame and picking up the language of structure.
In 1984 the sociologist Charles Perrow, studying the near-meltdown at Three Mile Island, proposed an idea he called normal accidents. Some systems, he argued, have two properties that together make catastrophe not a risk but a certainty over time. The first is tight coupling: the parts are so interdependent that a failure in one propagates instantly to the others, with no slack, no buffer, no time to intervene. The second is interactive complexity: the parts interact in so many ways that no human can foresee them all, so failures combine in paths no designer anticipated. In such systems, Perrow said, accidents are not aberrations to be engineered away; they are normal, an inherent property of the structure, and the only open questions are when and how large. Perrow was writing about reactors and chemical plants. He did not live to see the modern software supply chain, which is the most tightly coupled and interactively complex system human beings have ever assembled, a structure in which one company can alter millions of machines simultaneously and no one can fully predict what the alteration will touch. By his framework, this fortnight's outage was not a failure of testing. It was a normal accident, as inevitable as rain.
There is a known way to loosen the coupling, and the industry calls it staged rollout: releasing a change to a sliver of machines first, watching, then a few more, so that a bad update poisons a sample rather than the world. It is, precisely, the deliberate reintroduction of slack, and it works. So the real question is not why the outage happened but why the discipline that prevents it is everywhere eroding, and the answer is not carelessness. It is economics, and it is worth stating exactly, because it is the part everyone skips. A firm that stages its rollouts is slower and costlier on every one of the roughly thousand ordinary days; the firm that pushes to everything at once is faster and cheaper on all thousand, and wins the market on all thousand, and loses only on the single tail day, which arrives rarely enough to fall outside the horizon on which markets actually price. So the disciplined firm is punished continuously and the reckless one rewarded continuously, and when the tail day finally comes, the cost of the catastrophe lands not on the firm that removed the slack but on the hospitals and airlines and strangers downstream, who never chose the risk and cannot bill anyone for it. This is not a moral failing of particular engineers. It is a textbook externality: the private incentive to remove slack diverges from the social interest in keeping it, because the firm captures the savings and the world absorbs the disaster.
The strongest objection to all this is the one Silicon Valley would raise, and it is genuinely powerful, so let me put it at full strength. Slack is waste. The relentless stripping of it, the efficiency, the velocity, the continuous deployment, is exactly what let software remake the world and deliver the abundance we now take for granted; the "discipline of slowness" is nostalgia, and the market is right to punish it, because ninety-nine times in a hundred the cautious firm is simply the slower firm, and slower firms deserve to lose. Every word of that is true, and it is why the erosion is so hard to resist. But it contains the error that undoes it. The market prices frequent, small, visible costs superbly and rare, large, externalised costs abominably, and the trouble is that velocity bundles the two together: the same relentless efficiency that is genuine gain on the thousand normal days is uncompensated tail risk on the one catastrophic day, and the market, seeing only the thousand, rewards both as if they were the same thing. It cannot tell the efficiency that creates value from the efficiency that merely borrows against a disaster someone else will pay for, because on every day it can observe, they look identical.
Which is why the fashionable response, be more careful, test more, deploy slower, is worse than useless: it asks individual firms to be virtuous against their own incentives, and virtue does not survive a market that punishes it. You cannot exhort your way out of a structural externality; you have to change the structure, so that the firm which removes the slack bears the cost of the crash. We did this for aircraft, for drugs, for banks, unglamorous liability and regulation and mandated process for the systems whose failure spills onto everyone, and we did it precisely because we learned that "be careful" is not a policy. Until we do the same for the handful of firms whose software can now fail the world in an afternoon, we will keep calling these outages accidents and demanding better testing, and we will keep getting them, on schedule, because they are not accidents and testing is not the lever. The next one is not a risk to be managed. It is an appointment already on the calendar, and we are choosing, by leaving the incentives exactly as they are, to keep it.