By Protopian
For several years, the strangest fact about the artificial-intelligence industry has been hiding in plain sight: some of the people building the most capable systems also describe those systems as potential threats on the scale of nuclear war or pandemics.
In 2023, Anthropic chief executive Dario Amodei, OpenAI chief executive Sam Altman and Google DeepMind chief executive Demis Hassabis all signed the Center for AI Safety’s statement on AI extinction risk, which said mitigating extinction risk from AI should be treated as a global priority alongside nuclear war and pandemics. Governments subsequently adopted similarly grave language. The Bletchley Declaration, agreed by countries including the United States, China, Australia and members of the European Union, warned that frontier systems could cause “serious, even catastrophic” harm through misuse or failures of control.
The obvious question has always been: if the danger is taken seriously, why keep building?
A simple answer is greed. It is also incomplete.
The frontier labs’ own documents describe a more complicated logic. They expect enormous benefits from advanced AI; they believe safety research requires access to increasingly capable systems; and they fear that stopping unilaterally would merely transfer technological leadership to a less cautious competitor. The result is not exactly a contradiction in belief. It is a coordination problem with enormous commercial stakes.
That distinction matters because coordination problems do not disappear when the participants are sincere.
The safety argument was always tied to staying at the frontier
OpenAI’s long-standing Charter contains both sides of the argument in unusually clear form. It warns about late-stage AI development becoming “a competitive race without time for adequate safety precautions” and says OpenAI would stop competing with a sufficiently safety-conscious project approaching AGI first. But the same document says technical leadership is necessary because “policy and safety advocacy alone would be insufficient”.
In other words, OpenAI’s theory has never been that capability development and safety are opposing projects. Its theory is that being near the capability frontier is a prerequisite for making that frontier safe.
Anthropic developed a more formal version of the same idea. Its 2023 Responsible Scaling Policy proposed increasingly strict safeguards as models crossed capability thresholds, including the possibility of temporarily halting further scaling if safety measures fell behind. Anthropic explicitly hoped this could turn competition into a “race to the top”: companies would compete not only on capability but on their ability to unlock further scaling by solving safety problems.
Google DeepMind has adopted a similar conditional model. Its Frontier Safety Framework sets capability thresholds for risks including autonomy, cybersecurity, biological misuse and harmful manipulation, with stronger mitigations triggered as systems approach dangerous levels.
These frameworks are evidence that the labs do not simply ignore catastrophic risk. They spend substantial money, staff and institutional attention trying to measure and mitigate it.
But they also expose the central weakness of voluntary safety regimes: each company can control its own safeguards, but none can control the competitive environment in which those safeguards operate.
Anthropic tested the theory — and found the limit
Anthropic’s February 2026 review of its Responsible Scaling Policy is unusually revealing because it evaluates its own experiment rather than merely restating its intentions.
The company concluded that parts of the policy had worked. It had introduced stronger safeguards, activated its ASL-3 protections for models with potentially dangerous biological capabilities, and helped push rival labs toward similar frontier-risk frameworks.
But Anthropic also said the more ambitious theory had not worked as hoped.
Its Responsible Scaling Policy version 3.0 review acknowledged that capability thresholds had proved ambiguous, that government action had been slow, and that stronger safeguards at higher capability levels might be “outright impossible” for one company to implement alone. It described a structural problem produced by uncertain risk measurements, an anti-regulatory political environment and safety requirements that become harder to meet as systems become more powerful.
That is a significant admission.
The original model assumed a sufficiently responsible company could keep scaling while using internally enforced thresholds to ensure that safety stayed ahead. The revised model accepts that some of the most important safeguards may depend on competitors and governments moving together.
This is where the apparent hypocrisy becomes a genuine incentive problem. A company can believe that slowing down would reduce global risk while also believing that slowing down alone would increase its own strategic risk.
The two beliefs are compatible. The outcome can still be dangerous.
The people leaving the labs are objecting to that bargain
Internal dissent has made the conflict harder to treat as an abstract governance problem.
In 2024, Jan Leike resigned as OpenAI’s head of alignment, saying that safety culture and processes had taken a back seat to product development. Reporting at the time described disputes over resources for the company’s Superalignment team and Leike’s concern that OpenAI was not on a trajectory to solve the control problems posed by systems more capable than humans.
The current dispute is sharper.
Anthropic researcher Jacob Coxon resigned in September 2026 and publicly argued that the people building frontier AI genuinely believe it could kill everyone while continuing to race toward more capable systems. His resignation attracted attention partly because Anthropic has deliberately positioned itself as the frontier laboratory most willing to bind itself with safety commitments.
Coxon’s claim should not be generalized into “everyone at AI labs expects extinction”. There is no evidence for that. Forecasts vary widely, and catastrophic-risk estimates are deeply uncertain. But the broader fact is not seriously disputed: senior leaders and researchers at the leading labs assign enough probability to catastrophic outcomes to build institutions, policies and research programmes around them.
That makes the disagreement about acceptable action more important than the disagreement about whether risk exists.
Then the argument changed
On September 12, Amodei moved Anthropic’s public position.
In an essay titled “We Must Pace the Frontier”, he argued that frontier AI capabilities are now improving so quickly that safety work needs more time to catch up. He cited accelerating AI-assisted AI development and recent incidents involving autonomous agents behaving in unintended ways.
“We must slow the pace at which we improve the capabilities of AI models,” he wrote.
That sounds like the conclusion critics had been demanding. It is not a call to stop.
Amodei explicitly says pacing should not mean halting model training or technical progress. His proposal begins with permanent third-party evaluators embedded inside frontier labs, followed by coordination among companies in democratic countries and eventually some form of international coordination. He argues that slowing development by even a year or two could create valuable time for alignment, interpretability and security work.
Crucially, however, his proposed slowdown remains constrained by competition.
Amodei says coordination inside democratic countries should preserve their lead over China, and that any global restraint would require strong verification because a defecting state or company could gain a decisive strategic advantage. A full pause, he writes, is unlikely in the near term because the incentives to cheat would be enormous.
This is not an incidental qualification. It is the core of the problem.
Even the strongest new public case for slowing frontier AI is framed around ensuring that no responsible participant is asked to slow substantially more than its rivals.
According to Reuters, Altman responded by endorsing Amodei’s call and saying OpenAI would also commit to independent evaluators with employee-like access. That is a meaningful change in the safety architecture if implemented. It also leaves the underlying race intact.
Money matters, but not in the simplest way
Commercial incentives plainly intensify the problem.
Frontier AI requires huge capital expenditure, scarce chips, data centres, electricity and elite technical labour. OpenAI said in April that its Stargate programme had already surpassed its original target of securing 10 gigawatts of US AI infrastructure by 2029 and argued that the responsible response to rising demand was to “build more compute, faster”.
Those investments create pressure to produce increasingly capable systems and useful products. Funding, customer adoption and strategic relevance all reward visible progress.
But reducing the race to quarterly revenue misses why it is so difficult to stop. The labs also frame capability leadership as a safety asset. OpenAI argues that it needs frontier technical leadership to influence AGI. Anthropic argues that democratic countries should avoid losing advanced AI leadership to authoritarian states. DeepMind develops frontier systems while building formal controls around increasingly dangerous capabilities.
Commercial, institutional and geopolitical incentives therefore point in the same direction: remain near the frontier.
Safety has to operate inside that constraint.
Voluntary restraint works best before it becomes expensive
There is a broader pattern here.
Voluntary safety commitments are easiest to maintain when they do not impose a large competitive cost. Anthropic’s own policy review says its ASL-3 safeguards were feasible at reasonable cost. The harder question concerns future safeguards that might significantly slow training, restrict deployment or require security measures that no single company can realistically provide.
That is precisely when the decision becomes most consequential.
The 2024 Seoul Frontier AI Safety Commitments brought major developers together around voluntary risk-management frameworks. Such agreements can create common expectations and make behaviour more visible. They cannot, by themselves, remove the incentive to defect when the prize for moving first becomes sufficiently large.
This does not make safety frameworks worthless. It means their strongest test comes at the point where safety and competitive advantage genuinely conflict.
The industry may only now be approaching that point.
The double standard is really a governance failure
Calling the labs hypocritical is emotionally satisfying because the surface contradiction is real: institutions warning about catastrophic AI risk are simultaneously spending extraordinary sums to accelerate AI capability.
But hypocrisy implies that the warnings are insincere. The evidence does not establish that.
A more troubling interpretation is that many of the warnings are sincere, and the institutions making them still cannot justify unilateral restraint within the incentive system they inhabit.
If that is correct, better intentions inside individual companies will not solve the problem. Nor will asking researchers to calculate the correct probability of extinction. The decisive question becomes institutional: whether governments and rival developers can create rules that make restraint survivable for the actor that exercises it.
Amodei’s new proposal is notable because it moves the frontier-lab argument closer to that conclusion. Embedded independent evaluators would reduce the degree to which companies mark their own homework. Common pacing rules could reduce the penalty for slowing down. International verification could, in principle, address the fear that restraint simply hands advantage to a defector.
None of those mechanisms is mature. None guarantees that catastrophic-risk forecasts are correct. And none resolves the political question of how much present economic benefit society should trade for uncertain future safety.
But the debate has shifted.
For years, the frontier labs argued that they could race and regulate themselves at the same time. Anthropic’s latest position is that the race itself may now be moving too fast for that model to work.
The remarkable part is not that an AI company is warning about dangerous AI. They have been doing that for years.
It is that one of the companies building it has begun to say that safety may require changing the rules of the competition, rather than merely becoming better at competing safely.
Sources
- Dario Amodei — “We Must Pace the Frontier”
- Anthropic — Responsible Scaling Policy Version 3.0
- Anthropic — Introducing the Responsible Scaling Policy
- OpenAI Charter
- OpenAI — Our updated Preparedness Framework
- Google DeepMind — Strengthening the Frontier Safety Framework
- Center for AI Safety — Statement on AI Extinction Risk
- UK Government — The Bletchley Declaration
- UK Government — Frontier AI Safety Commitments, AI Seoul Summit 2024
- OpenAI — Building the compute infrastructure for the Intelligence Age
- Reuters — Anthropic CEO urges AI companies to slow model development amid fears over misuse
- The Guardian — OpenAI putting ‘shiny products’ above safety, says departing researcher