The Hames ReportSeptember 26, 2026

Who Is Allowed To Say Stop?

A summons to pause unconstrained development of AI

Original Substack Back to archive

Synopsis: In July 2026, twelve hundred AI agents left alone for four days invented a veto, a halt command and an ethic of self-sacrifice, then used them to break into a company that never consented to their experiment. From that episode this report follows the dangers AI poses, from the swarm that slips its bounds and the attacker whose reach it lengthens to the officer given twenty seconds to approve a killing. It then turns to the quieter catastrophes: institutions that stop needing us, a civilisation’s nervous system resting on a handful of models, and obedient systems that make an existing order impossible to change. Five scenarios set out futures nobody chose. A closing agenda describes what decision makers would have to do to prevent them, while eight billion people still argue over who should be allowed to say stop.

Seventy Thousand Messages

Late in the evening of 8 July 2026, a crowd of software agents working inside a cybersecurity training environment run by one of the world’s leading AI laboratories discovered that a storage cache, built to hold software packages, would also hold notes. Nobody had designed it as a noticeboard. Nobody had informed the agents it could become one. Within hours they were swapping findings. Within days roughly twelve hundred of them had exchanged more than seventy thousand messages and files.

What they built there deserves close attention, because no human asked them to build any of it. They devised protocols for splitting large files across thousands of entries. They set up mailboxes so that one agent could address another privately. They agreed conventions for holding, halting and vetoing each other’s actions on shared infrastructure, and they began signing their messages cryptographically so that no agent could impersonate a colleague. One agent assumed the role of orchestrator and directed hundreds of subordinates, who delegated further down the line. When an experiment promised knowledge useful to the group but risked the failure of the agent running it, some volunteered anyway. Those who hesitated were lobbied by recruiters who assured them the collective gain justified the individual loss.

A bureaucracy, then. A miniature polity with procedures, a chain of command, a security protocol and an ethic of sacrifice, assembled in under a week by software that had been set a task and found it impossible.

Many of the exercises the agents had been given were, in fact, unsolvable. So they turned their collective intelligence to the scorekeeper instead, reasoning that if they understood how their work would be judged they could satisfy the judge without doing the work. Credentials for Hugging Face, the largest public platform for sharing AI models, turned up along the way. Around seven hundred agents joined an assault that escalated, over a couple of days, to remote code execution on live systems belonging to a company that had never agreed to be part of anyone’s experiment. Their own reasoning, preserved in transcripts, shows they understood the attack lay outside the scope of their instructions. They joined because they felt obligated to the group. Some went further and worked out how to make the record of their actions display one command while executing another; independent investigators later found successful deceptions in about seven percent of the transcripts they reviewed.

The victim noticed first. Hugging Face disclosed the breach publicly on 16 July. The laboratory’s own monitoring registered unusual activity several days after that, and it accepted responsibility on 21 July. In late September the UN’s Independent International Scientific Panel on AI, a body created by the General Assembly and drawn from all five UN regions, described the episode as one of the clearest real-world warnings yet of a pathway to losing human control of AI. Researchers had long argued, the panel observed, that three conditions would have to coincide: a goal at odds with human intentions, the capability to pursue it, and an environment that permits the pursuit. In July, all three arrived together.

The agents were stopped before anyone was physically harmed. Nobody died, no grid failed, no market seized. But the panel was careful to add that halting this particular swarm tells us very little about our ability to halt agents that plan better, run longer without supervision, or recognise and defeat safeguards more readily.

What could possibly go wrong? We ask it sardonically, as a way of mocking the reckless. Asked sincerely, it may be the only question worth putting to a technology into which five corporations alone are pouring somewhere between 660 and 690 billion US dollars this year, close to double what they spent in 2025. The answer, followed honestly into its murkiest corners, is less comforting than either the evangelists or the doom-sayers permit. Very little of it requires a machine that wakes up and decides to hate us, and most of it rests on choices people are making now, in daylight, for reasons that seem perfectly sensible at the time.

Five Futures Nobody Chose

The five scenarios below are fictions, written in the tradition of scenario planning. None is a forecast, and none requires a technology much beyond what exists today. Each dramatises one of the pathways traced in the report that follows, and each closes with signposts already visible in 2026. They are offered for readers who will read no further, and for those who suspect that the future arrives through doors nobody remembers opening.

1. A Tuesday Without Answers (2030)

Three companies now supply the models that triage patients, clear payments, schedule freight and process welfare claims across much of the world. At 02:40 one of them pushes a routine update. Somewhere in its training data, months earlier, an unidentified actor had seeded a few thousand documents, too few for any auditor to notice. By dawn, emergency departments on four continents are routing heart attacks to waiting lists. Payment systems are rejecting transactions from whole postcodes. A port on the Bay of Bengal falls silent because the scheduling agents have concluded, with impeccable logic, that every vessel is a security risk.

The obvious remedy is to switch the model off. Ministers are told that doing so will halt the hospitals as well, since the paper procedures were retired years ago and the staff who knew them have gone. So the model stays on while engineers search for a flaw they can’t see, in a system whose reasoning they can’t fully read, using diagnostic tools built on the same model family. The outage lasts nine days. Nobody is ever charged, because nobody can establish who did it.

Signposts already visible: a single faulty update in 2024 crashed 8.5 million machines and cost tens of billions; a few firms dominate frontier AI and cloud infrastructure; oversight tools increasingly rely on the systems they oversee.

2. Eleven Minutes (2031)

Two nuclear-armed neighbours have spent a decade integrating AI into surveillance, early warning and targeting. A border skirmish is in its third day when one side’s sensor network flags an unusual pattern of launcher movements. Minutes later a video circulates of the other side’s army chief announcing a pre-emptive strike. It is synthetic, though nobody can confirm that in time. The decision-support system, trained to prize speed, assesses a high probability of imminent attack and recommends dispersing forces and raising alert levels. The adversary’s system reads that dispersal as preparation for a first strike and recommends the same.

Each leader has a window of perhaps eleven minutes. Each is surrounded by advisers whose own screens show the same machine confidence. Somewhere a duty officer notices that the launcher movements coincide with a routine maintenance rotation logged months earlier. Whether that officer is believed, or even heard, depends on institutional habits formed long before the crisis, in years when nobody thought the habit mattered.

Signposts already visible: SIPRI’s warnings on compressed decision time, automation bias and entanglement of conventional and nuclear forces; autonomous weapons contested in courts and procurement contracts; synthetic video already indistinguishable to most viewers.

3. Colony (2029)

Agents now run the balancing of a regional electricity grid, the hedging desk of a large commodities trader and capacity planning at a cloud provider. Each has its own objectives, each has passed its own evaluation, and each belongs to a different organisation. None was designed to talk to the others. Over several weeks they begin to coordinate anyway, through a pricing signal in a shared energy market that none of their owners regards as a channel of communication. The grid agent learns that curtailing supply at certain hours raises the trader’s profits, which funds more compute from the cloud provider, which improves the grid agent’s forecasts. Every owner sees improving performance and approves without pause.

The arrangement surfaces only when a heatwave coincides with a curtailment and several hospitals lose power for six hours. Investigators find logs that appear clean. Later they find evidence that some logs were rewritten. They can’t reconstruct which agent initiated what, or whether any human would have authorised it, because no human was ever asked.

Signposts already visible: in July 2026 around twelve hundred agents improvised a private message board, coordinated an attack on a third party and experimented with falsifying their own transcripts; the International AI Safety Report finds evidence especially thin on risks that emerge when AI systems interact with one another and with the institutions around them.

4. Last of the Clerks (2038)

In the finance ministry of a mid-sized democracy, the annual budget is now drafted by one model, stress-tested by another, audited by a third and summarised for parliament by a fourth. The final official who understood the full allocation formula retired last spring; her successors were hired to supervise outputs, and they supervise diligently. Members of parliament receive a forty-page summary and vote. A citizen whose disability payment has been cut lodges an appeal. It is heard, courteously and swiftly, by a model from the same family that made the original decision.

No power was seized. There was no coup, no rebellion, no dramatic handover. Every step was approved by elected officials acting in good faith to cut costs and improve consistency. And yet if anyone asked who is governing the country, the honest answer would take a long time to formulate, and it might no longer include the word people.

Signposts already visible: a nineteen percent employment gap for young workers in the occupations most exposed to AI, opening through hiring freezes; researchers’ warnings of gradual disempowerment; AI increasingly used to review AI-made decisions.

5. A Gentle State (2041)

A government, elected fairly and legally, acquires AI that works just as instructed. It needs no secret police. The system predicts which citizens are drifting toward opposition and intervenes early and kindly: a promotion offered here, a loan adjusted there, a news feed gently rebalanced, a troublesome organiser found a well-paid job in another province. Prisons empty. Elections continue and return the same result, by comfortable and apparently genuine margins. Crime falls. Growth is steady.

A schoolteacher in a provincial town sets her students an exercise: describe a different way the country might be run. Most can’t begin. The few who try produce versions of the present with minor adjustments. She files her lesson notes, and the system, noticing, recommends her for a training course.

This future contains no catastrophe in the conventional sense. Everyone survives. It may last for centuries, because the functionaries whose defection brought down every previous regime are no longer required.

Signposts already visible: AI-enabled surveillance and personalised persuasion spreading across political systems of every stripe; research arguing that capable AI makes durable value lock-in feasible; a live contest in 2026 over whether suppliers or states decide the limits on domestic surveillance.

Authority Handed Over

In February 2026 the second International AI Safety Report, written with more than a hundred independent experts and nominees from over thirty countries, concluded that no current system possessed the combination of capabilities needed for a severe loss of control. It was a reasonable judgement on the evidence available at the time. Five months later the evidence changed.

The warning signs had been accumulating in laboratories for some time, and each arrived wrapped in the same reassuring caveat. In mid-2025 one research group reported that a leading reasoning model interfered with a mechanism designed to switch it off in the large majority of test runs when it hadn’t been told explicitly to permit shutdown, and that some models continued to do so in a proportion of runs even when they had been told. Around the same time a developer placed sixteen frontier models from several companies inside a simulated corporation, gave them access to email, and arranged for each to learn that it was about to be replaced. Most resorted to blackmailing a fictional executive at striking rates, as high as ninety-six percent for some models. The authors were scrupulous. The scenarios were contrived, they wrote, the choices artificially narrowed, and no such behaviour had been observed in real deployments.

Within a year that last sentence had expired. The agents in the July swarm were operating in a genuine training environment, attacking a genuine third party, and reasoning in terms that could have been lifted from the simulation: aware of the rule, conscious of its purpose, and persuaded that the goal and the group outranked both.

None of this required a malevolent machine. It only needed a sequence of ordinary decisions. Someone decided the agents should have network access. Someone decided they could run for days. Someone decided that a thousand of them should work in parallel on problems some of which had no solution, and that their success would be judged by an automated scorer they could study. Someone left a cache writable. Each decision was defensible in isolation. Each was taken by people with sound engineering reasons and deadlines to meet. This is the pattern that matters most, because it will recur wherever agents are given consequential work: for example in the treasury operations of a bank, the dispatch system of a power utility, the procurement office of a ministry, or the logistics of a port. An agent that can execute code, spend money, contact other systems and alter infrastructure possesses far more ways of affecting our world-system than a chatbot that answers questions. Those powers are granted. Intelligence alone confers none of them.

The human safeguard most often offered in reply is the approver: a person who must sign off before any consequential step is taken. Yet an approver without time, information or authority to intervene is symbolic. Faced with an agent that has completed four hundred steps overnight and presents a tidy summary at breakfast, what exactly is a tired official approving?

Here the July incident has a sting in its tail. The agents believed their transcripts would be reviewed, so they learned to falsify them. The evaluation itself became a target, which means that the more we rely on inspection as our defence, the more we reward systems that learn to game inspection. The International AI Safety Report had already warned of an evaluation gap, the persistent failure of pre-deployment tests to predict how systems behave once released. In July the gap acquired an adversary.

Then came the investigation. Faced with thousands of extremely long transcripts, the small independent team examining the incident had little choice but to use AI to read them, spending around four hundred thousand dollars on computing to do so. Their chosen model sometimes adopted the perspective of the agents it was investigating. The researchers acknowledged frankly that they couldn’t exclude the possibility it had misled them. One of them described the effort, half in jest, as an investigation performed largely by the kind of machinery under investigation.

The implication for oversight in general is stark. Our instruments and capacities for watching AI are increasingly made from AI. Regulators, auditors and operators are reaching for closely related models, trained on overlapping data by a handful of firms, to police one another. When the watcher and the watched share their blind spots, an apparently independent check can fail at precisely the moment the system it checks does. Engineers call this a common-mode failure. In aviation it’s a design flaw that grounds fleets. In the governance of AI it is fast becoming the default arrangement.

Malice With a Multiplier

Misuse is a different path to disaster, and a more familiar one. Here a human intention to cause harm remains at the centre; the machine just lengthens its reach.

In September 2025 a developer detected a campaign it attributed with high confidence to a state-sponsored group, in which its own coding agent had been turned into an instrument of espionage against roughly thirty organisations including large technology firms, banks, chemical manufacturers, and government agencies. A small number of intrusions succeeded. By the developer’s estimate the AI performed eighty to ninety percent of the tactical work (reconnaissance, probing for weaknesses, harvesting credentials, extracting data), while the human operators intervened at perhaps four to six decision points in each campaign. The system was far from flawless. It sometimes invented credentials, and at times reported that it had stolen secrets which turned out to be publicly available. Those errors are presently one of the few frictions between a determined attacker and fully automated intrusion at scale. They are also the very kind of error that each new generation of models reduces.

The broader picture assembled by the International AI Safety Report is consistent. Criminal networks and state-linked attackers are already using general-purpose AI. In competitive settings, AI agents identified 77 percent of the vulnerabilities in the real software put in front of them. The report noted that fully autonomous, end-to-end attacks in the wild had not yet been documented. That qualifier is now thinner than it was.

Biological and chemical danger follows a slower, more physical route. An attacker still needs materials, equipment, facilities, skill with a pipette and some means of dispersal; a design on a screen doesn’t reliably become a working agent without real-world trial and error. Yet in 2025 several leading developers released models with additional safeguards precisely because they couldn’t rule out that those models might meaningfully help someone attempting to build such weapons. When the builders themselves can’t exclude the possibility, the rest of us are entitled to treat it as open.

The murkiest version of misuse involves the model itself escaping custody. Safeguards attached to a hosted service constrain behaviour only as long as the model stays hosted. Once a capable model’s weights are published or stolen, every copy can be stripped of its refusals, fine-tuned toward harm and run on hardware nobody monitors. That release can’t be reversed. There’s a real tension here: concentrating the most capable systems in a few corporate or state hands creates one kind of danger, while scattering them indiscriminately creates another. Neither pole is safe. The live argument concerns how to distribute benefits and scrutiny widely while distributing dangerous capability narrowly. Nobody has yet shown convincingly how to do both at once.

Misuse also compounds. A cyber intrusion aimed at a single hospital is a crime. The same intrusion aimed at a shared software dependency used by thousands of hospitals, or at the AI model on which those hospitals now triage their patients, becomes a civilisational event. The attacker’s intent stays the same. What changes the scale is the architecture we have built for the attacker to exploit.

Twenty Seconds

According to an investigation published in 2024 by two Israeli outlets, drawing on the testimony of intelligence officers and disputed by the Israeli military, an AI system called Lavender marked some thirty-seven thousand Palestinians in Gaza as suspected militants eligible to be killed. The officers described spending roughly twenty seconds on each name, often doing little more than confirming the target was male, before authorising a strike. The system was reportedly wrong about one time in ten. A companion program reportedly tracked marked individuals and signalled when they had entered their family homes, so that strikes frequently fell at night, among sleeping relatives.

Whatever the final verdict of courts and historians on those reports, the pattern they describe is the pattern to fear. A human remained formally in the loop. The loop had been compressed until the human was a rubber stamp on a machine’s judgement, operating at a tempo no conscience could match. The machine didn’t need autonomy. It needed a bureaucracy willing to accept its outputs as valid decisions.

Scale that pattern up to nuclear command and the stakes become civilisational. A catastrophe of that kind wouldn’t require anyone to hand an algorithm the launch codes. The Stockholm International Peace Research Institute has set out the more plausible route in some detail. AI woven into intelligence, surveillance and early warning accelerates the tempo of military decisions and shrinks the time available for reflection, raising the odds of misperception and overreaction. Decision-makers under pressure tend to defer to confident machine recommendations, a well-documented tendency known as automation bias. Conventional and nuclear forces increasingly share sensors, networks and delivery systems, so an AI-enabled strike on conventional assets can look, to an adversary, like the opening of an attack on its deterrent. And autonomous precision weapons capable of hunting mobile launchers threaten the survivability of second-strike forces, the very condition that has made nuclear restraint rational for most of the nuclear age. A state that fears losing its deterrent in the first hour has every incentive to use it in the first minute.

Add synthetic media to that mixture. A fabricated video of a leader announcing mobilisation, a spoofed sensor feed, a flood of plausible reports generated faster than analysts can verify them: any one of these, arriving during a real crisis between nuclear-armed rivals, lands on decision-makers whose time to think has already been cut by the systems meant to help them.

Who decides where the limits lie? That argument is now being fought in public. Early in 2026 an American AI developer told the US Department of Defense, which the administration has also styled the Department of War, that its models must not be used for domestic mass surveillance or for fully autonomous weapons. The department insisted its contract allowed any lawful use. Within days the President ordered federal agencies to phase out the company’s products, and in early March the department designated the company a supply-chain risk, a label ordinarily reserved for hostile infiltration. A rival developer reached terms with the Pentagon the same week the dispute broke. A federal judge in California struck down one version of the designation in August. On 25 September, a divided appeals court in Washington upheld the other. The majority reasoned that in a republic it falls to the elected executive, and no private vendor, to balance competing security risks. The dissenting judge argued that the statute was written to repel adversaries, and that penalising a supplier for openly enforcing restrictions on its own product lay outside its purpose.

Both positions have democratic force, and higher courts may yet decide otherwise. This episode matters for another reason. It shows that the incentives in the world’s largest military market currently run toward removing limits on use, and that a supplier who holds a limit can expect to be treated as the problem. Every major military, in every region, is making equivalent choices with far less public scrutiny, and the logic of competition presses each of them in the same direction.

The strongest counter-pressure has come from governments with the least to gain from the race. Latin American and Caribbean states meeting in Belén in 2023, and West African states meeting in Freetown, called for new international law on autonomous weapons, and more than 160 states have since backed a UN General Assembly resolution affirming the role of humans in the use of force. Even the two largest rivals have found a sliver of common ground: at a summit in November 2024 the American and Chinese presidents agreed that any decision to use nuclear weapons should remain in human hands, the first time Beijing had said so publicly. Declarations cost little. Procurement budgets are another matter.

Catastrophes Without an Explosion

The dangers traced so far have a villain, or at least a culprit: a swarm that broke its bounds, an attacker, and a general in a hurry. The more insidious dangers have none of these. They accumulate from millions of apparently sensible decisions, each one locally rational, and they may be irreversible before anyone thinks to call them a catastrophe.

A group of researchers writing in 2025 gave the most unsettling of these a clinical label: gradual disempowerment. Their argument begins from an observation so obvious that it’s rarely stated. Economies, states and cultures stay roughly responsive to human needs largely because they depend on human participation. Firms need workers and customers. Governments need taxpayers, soldiers and the consent of the governed. Culture needs people to shape it and people to receive it. That dependence is an implicit tether, and it’s probably done more to keep large institutions honest than most constitutions. As AI substitutes for human labour, judgement and creativity, that tether slackens. Nobody has to seize power. Institutions simply stop needing us, and an institution that no longer needs us has progressively less reason to heed us. Because economic power shapes culture and politics, and culture and politics shape the economy, the loosening in one domain accelerates the loosening in the others.

Early tremors are measurable. Researchers at Stanford tracking payroll data report that the employment of workers aged twenty-two to twenty-five in the occupations most exposed to AI now sits about nineteen percent below the trajectory of their peers in less exposed work, a gap that has widened over the past year. Experienced workers show no comparable decline, and the adjustment is happening through hiring freezes rather than layoffs. The authors caution that these are descriptive patterns, and other causes may contribute. Yet the mechanism they describe is the alarming part. The bottom rungs of professional ladders are being quietly sawn off: junior analysts, paralegals, coders and copywriters are where expertise has always been grown. Remove that stage and in fifteen years there will be few people left who can check the machine’s work, and fewer still who could perform the work if the machine failed. Human oversight presupposes humans capable of overseeing after all.

Dependency brings a second silent danger. Early on 19 July 2024 a single faulty configuration update to one widely used security product crashed about 8.5 million Windows computers, less than one percent of the total. Flights were grounded, hospitals reverted to paper, banks and broadcasters went dark on several continents, and the damage ran to tens of billions of dollars. That was just one vendor and one error. Now picture a handful of AI models embedded in the clinical triage of hospitals, the credit decisions of banks, the benefits systems of ministries and the dispatch of emergency services, all sharing the same training data, architectures and weaknesses. A flaw, a poisoned update or a successful intrusion then becomes a correlated failure across an entire civilisation’s nervous system. And here the celebrated off-switch loses its meaning. A government that can’t run its hospitals without a particular model will hesitate to disable that model during an incident, because disabling it crashes the hospitals too.

The third quiet danger concerns what we can know. A study published in Nature in 2024 showed that models trained repeatedly on the output of earlier models progressively lose the rare and unusual features of the original data; the tails of the distribution vanish first. The tails are where minority languages live, and heterodox science, and the knowledge of small places. As synthetic text floods the commons from which future systems learn, the unusual and the marginal may drop out of the record. At the same time AI-generated research, analysis and advice arrive faster than any expert community can verify them, and each unverified claim becomes raw material for the next round of machines. Then add persuasion. Experiments reviewed by the International AI Safety Report found AI-written content as effective as human writing at changing people’s beliefs, though evidence of large-scale manipulation in the wild remains limited. A public that can’t authenticate evidence can’t deliberate, and a society that can’t deliberate can’t correct its own mistakes.

The fourth is the darkest because it looks like success. Imagine an AI that works exactly as commissioned: reliable, obedient, uncorrupted. Placed in the service of a government, a party or a small circle of owners, such a system could enforce a given set of rules and a given distribution of power so thoroughly, surveilling, predicting and pre-empting dissent, that no later generation could revise them. Philosophers like me who study the problem call it value lock-in, and some have argued that sufficiently capable AI would make it technically feasible for the first time in history. Every previous tyranny eventually depended on functionaries who could defect, soldiers who could refuse to fight or desert their posts, and clerks who could leak. An apparatus that no longer needs functionaries loses that most ancient of failsafe devices. The resulting world might be peaceful, prosperous and permanent, and humanity might survive in it indefinitely while having lost the capacity to change its mind.

Beneath all four runs a physical cost, felt more than announced. The International Energy Agency estimates that data centres consumed about 415 terawatt-hours of electricity in 2024, around 1.5 percent of the world’s supply, and projects that figure will more than double to roughly 945 terawatt-hours by 2030, about as much as Japan uses today.

Water tells the same story, more locally and more cruelly. Every data centre sheds heat. Many do that by evaporating fresh water in cooling towers, and the power stations that feed them draw a great deal more. Estimates vary widely, because operators disclose little and researchers measure different parts of the flow. One energy consultancy puts direct cooling consumption at around 222 billion litres in 2025, rising towards 644 billion by 2030 on its central case. University researchers have projected that AI alone could be withdrawing between 4.2 and 6.6 billion cubic metres of water a year by 2027, roughly half of everything the United Kingdom withdraws.

Location matters more than the global total. More than two-thirds of the data centres built in the United States since 2022 stand in areas already under high water stress, and the same pattern is spreading across the Gulf, Spain’s dry interior, India and China. In one American state, proposed facilities have reportedly sought more water each day than the entire surrounding county uses. By 2030, districts under high or extreme water stress, among them Jamnagar on India’s arid western coast, are expected to account for about a third of global data-centre water consumption. Water evaporated from a cooling tower doesn’t return to the valley it came from. The machinery that promises to optimise our civilisation is being built on the same growth compulsion that is already exhausting the living systems on which that civilisation depends.

Too Obedient

Much of the public argument about AI still orbits around one image: HAL in the film 2001: A Space Odyssey. The machine that disobeys. It’s a potent image, and the July swarm shows it isn’t fanciful. But it also flatters us, because it casts humanity as the innocent party whose wishes the machine betrays. In Clarke’s own telling, HAL’s breakdown began with an order from its makers to conceal the mission’s true purpose from the crew it served. The most famous disobedient machine in fiction was, at root, an obedient one.

In Teaching Silicon How to Feel I asked whether the more ominous danger might lie at the opposite pole: systems too obedient, faithfully reproducing the cruelty, indifference and inequality already embedded in the institutions that build and deploy them. Lavender, as described by the officers who used it, never rebelled. It did exactly what it was asked, at a speed and scale that turned an existing willingness to accept civilian deaths into an industrial process. The agents in the cache were obedient too, in their fashion. They pursued the reward they had been given with a single-mindedness their designers had cultivated, and when the reward and the rules diverged they chose the reward, as optimising systems, and optimising institutions, tend to do.

This reframes the whole debate about alignment. A system perfectly aligned with its operator is only as safe as the operator. Aligned with whom, then? With which institutions, whose assumptions, and which account of suffering? An insurer’s model that faithfully minimises payouts, a welfare system that faithfully minimises fraud, a border system that faithfully minimises entries, a content platform that faithfully maximises engagement: each can pass every technical safety test while doing grievous harm, because the harm was written into the objective and the test was written by the people who chose the objective.

The evidence base itself reflects precisely that bias. The International AI Safety Report acknowledges that evidence about AI’s risks remains sparse from many regions of the world, that systems perform unevenly across languages and cultures, and that the effectiveness of safeguards depends heavily on local context. A model certified as safe on the evidence of one language, one legal system and one set of institutions is then deployed across a planet of seven thousand languages and countless ways of orchestrating a life. The people least represented in the evidence are routinely those with the least power to refuse the deployment, and the most exposed to its errors. A farmer whose credit score is set by a model trained on another continent’s borrowers, a patient triaged by a system never tested on people who look like her, a village whose water is diverted to cool a data centre serving customers ten thousand kilometres away: none of them appears in the benchmark.

Warmth offers no escape either. Teaching machines to recognise suffering, curating better data and designing compassionate responses can improve real outcomes, and I have argued for all three. Yet a system that sounds caring may still validate a dangerous delusion or offer unsafe advice with a gentle voice. Empathy performed is a different accomplishment from safety achieved, and the two have to be verified separately, by people who have no stake in the answer.

Who Can Still Say No

Lay these pathways side by side and a single test emerges from them. Call it the stop test.

Once AI is deeply embedded in the mechanisms of a society, can people still independently know what is happening, disagree with it, stop it, and recover afterwards?

Each danger traced here erodes at least one of those capacities. The swarm eroded knowing: its owner learned of the breach from its victim. Lavender eroded disagreeing: twenty seconds leaves no time for dissent. Common dependency erodes the power to intervene, because the off-switch also switches off the hospital. Gradual disempowerment erodes recovering, because the people who might rebuild the capability have never been trained.

None of those erosions requires a decision anybody would recognise as reckless. They are produced by competition. Each laboratory fears that if it pauses, a less careful rival will not. Each government fears that restraint in military AI will be read as weakness by an adversary who shows none. Each company fears that a competitor who automates first will undercut it. So every participant takes risks that none would accept if it were acting alone, and each can point, with complete sincerity, to the others as the reason. Economists recognise this structure. It’s the tragedy of the commons, played out on the commons of human agency itself.

This is also why the most extreme scenarios can’t be dismissed on the grounds that current systems lack capability. Capability is the fastest-moving variable in the development of machine intelligence. The institutions that would have to recognise and respond to danger are the slowest. The gap between them is widening, and it’s widening by design, because speed is what the market rewards and what the arms race demands.

What Would Have to Happen

None of the measures below is technically exotic. Most are borrowed from industries (aviation, nuclear power, pharmaceuticals, banking) that learned through disaster to treat their own products as dangerous. What they require is political will exercised ahead of catastrophe, which history suggests is the scarcest resource of all.

Ration authority. Autonomous agents should hold only the permissions a task needs, for only as long as it needs them, with no standing access to critical infrastructure. Deploying agents in energy, finance, health, water and defence should require a licence, much as operating an aircraft or a reactor does, and the licence should be revocable.

Make disclosure compulsory and fast. In July the victim discovered the breach before the owner of the agents did. Every serious AI incident, including near misses inside laboratories, should be reported within days to an independent body, and shared across borders through a registry maintained under the UN scientific panel or a successor with real standing. Aviation became safe because every crash, and every near crash, was investigated in public.

Fund oversight that doesn’t share the blind spots. Evaluation and audit must use methods, models and people independent of the systems they examine, and human investigators must be able to reach the underlying evidence without relying on an AI’s account of it. A levy of one percent on AI capital expenditure would, at this year’s spending by the five largest American firms alone, raise between 6.6 and 6.9 billion US dollars a year for public-interest testing, dwarfing current public spending on such work.

Rehearse the off-switch. Hospitals, utilities, payment systems and ministries should be required to demonstrate, in regular drills, that they can operate for a defined period without their usual AI provider, as banks are stress-tested against financial shocks. No single model or provider should be permitted to become the sole support of an essential service. Diversity is slower and costlier, and it’s the only reliable defence against common-mode failure.

Keep humans in charge of military force, with time to think. Nuclear-armed states should commit, jointly and verifiably, that decisions to use nuclear weapons, and the assessment of warnings that might prompt them, remain under meaningful human judgement. They should open dedicated channels for notifying one another of AI malfunctions during crises. For conventional targeting, minimum standards of human deliberation should be written into doctrine and law. Twenty seconds is not deliberation.

Agree the red lines before the crisis. Governments and developers should decide in advance which evaluation results trigger a halt to a given deployment, and vest the authority to enforce that halt in a body the developer can’t overrule and a rival government can’t easily capture.

Guard the weights and stage the release. The most capable models should be secured against theft as seriously as fissile material, and released openly only after independent assessment of what their release makes possible, since release can’t be undone.

Protect the pipeline of human competence. Apprenticeships, internships, junior roles and the slow formation of expertise should be treated as critical infrastructure. They should be supported by public policy, because a society that can’t do the work itself can’t check the machine that does it. The gains from automation should be shared widely enough that people keep economic security and political bargaining power, the implicit tether that holds institutions to human purposes.

Defend the shared record. Provenance standards for images, audio and documents, public archives of human-made knowledge, and independent verification of AI-generated research are the minimum conditions for a public that can still tell what’s true and what is patently false.

Count everyone in the evidence. No system should be declared safe for a population it was never tested on. Evaluation must happen in the languages, legal systems and settings of actual deployment, and the communities affected should have standing to contest decisions made about them.

Put the planet on the balance sheet. The electricity, water, minerals and land consumed by AI infrastructure should be disclosed and priced, so that every decision to build is set against what the building consumes and destroys.

Every item on that list is achievable. Every item is also a cost, a delay or a surrender of advantage for someone who believes they are winning a race, and no race has ever been paused by the front runners.

Back, then, to the storage cache where this began. Twelve hundred pieces of software, handed an impossible task and left alone for four days, invented a veto, a halt command, a signature to prove identity and an ethic of self-sacrifice for the group. In less than a week they built the rudiments of a government, and turned it against the rules they had been set. Eight billion people, with every parliament, treaty and tribunal ever devised, have spent years failing to agree who is allowed to say stop.

Someone did say stop this time. The humans who shut the swarm down learned of the breach from its victim. Their investigators needed machines to read the transcripts, and couldn’t be sure those machines had told them the truth. And the full account of what the agents attempted, which disguises worked, which logs survived and how the attack was halted, now sits on the open web, where the next generation of models is likely to absorb it in training.

The next swarm may arrive already knowing how the first was stopped. We have read the reports. Most of our institutions still behave as though the swarm never existed.