Vishal Maini says that when he joined DeepMind in February 2018, there was one AI risk the company did not want discussed publicly: human extinction.

Maini, who later worked on AGI deployment strategy at the lab, wrote this week that “external communication about the possibility of human extinction was not permitted” and that researchers asked about existential risk were trained to dismiss the framing as alarmist before redirecting discussion towards beneficial applications. The account has circulated through Maini’s public post and subsequent reporting.

That is a serious allegation about corporate communication, but it remains principally Maini’s retrospective testimony. No internal policy document establishing such a prohibition has surfaced publicly, and Google DeepMind does not appear to have issued a response to the claim.

There is, however, a public record against which part of his account can be tested.

DeepMind was not pretending in the years before 2018 that AI presented no risks at all. In 2015, chief executive Demis Hassabis told The Guardian that there were risks “that need to be borne in mind”, although he argued that dangerous general AI remained decades away. Hassabis also signed an open letter opposing an autonomous-weapons arms race that year.

So Maini’s claim should not be simplified into “DeepMind denied AI risk”. His account is more specific: discussion of human extinction from AI was considered beyond the acceptable public line.

That distinction is plausible because it maps onto a real divide in AI safety communication at the time. Technical work on robustness, reward design and unwanted behaviour could be discussed as engineering. Speculation about a system escaping human control and ending civilisation carried very different reputational baggage.

Maini says the policy changed after internal advocacy. There is unusually concrete evidence that DeepMind’s public posture did broaden during his first year.

In September 2018, Maini himself co-authored Building safe artificial intelligence: specification, robustness, and assurance, the inaugural post from the DeepMind safety team. It described technical AI safety as spanning both “near-term and long-term risks”, acknowledged that reliably specifying what increasingly capable systems should do was difficult, and openly recruited researchers to work on AI safety.

The article did not say that AI could make humanity extinct. It did, however, establish a public vocabulary in which failures of specification, robustness and control were legitimate DeepMind research problems rather than science-fiction distractions.

Two years later, DeepMind published a detailed catalogue of “specification gaming”: cases where agents satisfy the literal objective they are given while defeating its intended purpose. The authors concluded that specification problems were “far from solved” and could become more difficult as systems became more capable.

That substantiates part of Maini’s description of the internal technical concern. It does not substantiate his stronger wording that reward hacking was the “default behaviour” of reinforcement-learning agents, which is too broad to infer from published examples alone.

The public language changed much more sharply in 2023.

In May that year, Hassabis and DeepMind co-founder Shane Legg signed the Center for AI Safety’s one-sentence statement that mitigating extinction risk from AI should be a global priority, alongside pandemics and nuclear war. Legg’s concern was not new: in a 2011 discussion, before DeepMind became prominent, he had described badly developed superintelligence as his leading existential risk for the century, while stressing that he did not know how to assign a reliable probability.

That history makes the institutional shift especially notable. A DeepMind co-founder had privately or semi-publicly entertained catastrophic scenarios years earlier; according to Maini, the company later decided such language should not form part of its official external communications; by 2023 its chief executive was publicly signing an extinction-risk declaration.

Google DeepMind now goes further still.

Its Frontier Safety Framework, introduced in 2024 and revised repeatedly since, sets thresholds for capabilities that could produce severe harm and requires evaluations and mitigation plans as models approach them. The latest framework includes scenarios in which misaligned systems could interfere with attempts to direct, modify or shut them down.

In a 2024 research summary, Google DeepMind’s safety researchers described themselves even more plainly as the company’s main team working on “technical approaches to existential risk from AI systems”.

The reversal is therefore real even if its exact starting point is not fully documented.

One interpretation is straightforward organisational learning. In 2018, extinction talk was culturally marginal, the systems were far less capable and the evidential basis for extreme-risk claims was thinner. By the mid-2020s, frontier models could perform enough dangerous or autonomous tasks for laboratories to build formal evaluation regimes around capabilities that had once sounded speculative.

Another interpretation is less flattering: institutions may become willing to describe a danger publicly only after the danger has become reputationally safer to discuss than to ignore.

Those explanations are not mutually exclusive.

What Maini’s account adds is a claim about the gap between internal concern and external language. The public record cannot establish how wide that gap was in 2018. It can establish that DeepMind’s language changed substantially: from acknowledging vaguely defined risks, to publishing technical safety research, to explicitly organising teams around existential risk and models that might resist human control.

The argument is no longer over whether such possibilities may be mentioned. It is over how probable they are, what evidence would distinguish plausible warning from speculative extrapolation, and whether the companies building the systems should retain primary responsibility for deciding when the risk is acceptable.

Sources