DeepMind's Hassabis calls for a mandatory AI testing body
He wants a FINRA-style referee testing frontier AI models before they can launch.
You're building your whole company strategy on a bet about where AI capability is headed, so it's worth reading the actual words of the person setting that pace, not a summary of them. Demis Hassabis runs Google DeepMind, one of the handful of labs actually building the most capable AI systems in the world. On July 14, 2026, he posted a long piece on X laying out two things: his honest read on the AGI (an AI as capable as the human brain) timeline, and a concrete proposal for how the US should test and govern these systems before they ship.
His headline claim is that full AGI is probably only a few years away. He compares this moment to the discovery of fire or electricity, not to the internet or mobile phones, because he thinks the shift is that fundamental. His estimate of the scale: AI's effect on the world could end up ten times the size of the Industrial Revolution, arriving at ten times the speed. On the upside, he points to faster drug discovery, new clean energy sources, and new materials, and raises the idea that resources could eventually stop being the limit on human progress at all, what he calls an era of abundance. On the downside, he says frontier AI models (the most capable systems being built right now) are already creating real cybersecurity risks, and that nuclear and biological risks could follow as the systems get more capable. He is specifically worried about agentic systems (AI that acts in multiple steps on its own) that can also improve themselves, and says nobody yet has a reliable way to keep control of that.
His read on the current moment is blunt. Labs and countries are locked in an intense commercial and political race, and that race is pushing capability ahead of anyone's actual understanding of it. In his words, nobody in the world knows for sure what happens from here, and even the experts inside these labs disagree with each other. His prescription is not to slow down. It is to build a referee. He wants public policy that keeps innovation moving while also rewarding labs for being careful, gets countries working together on safety, and pushes labs to think harder about how their systems actually get used once they are out.
The concrete proposal is a new US Frontier AI Standards Body, and he is specific about its shape. He wants it modeled on FINRA, the Financial Industry Regulatory Authority that oversees stockbrokers, meaning a public-private partnership with government oversight rather than a pure government agency. Its board would include independent technical experts and people from the open-source world (where a model's files are public so anyone can run it). Funding would be substantial and mostly paid for by the AI industry itself, covering both expert salaries and the compute (the chips and time to run tests) needed to test frontier models properly. The Body would work with federal agencies and the US National Labs on any testing that touches national security. A model would officially count as Frontier-class once it crosses capability thresholds on a set of benchmarks (standard tests everyone in the field runs) that the Body sets and updates over time. Any company whose models cross that line becomes a Frontier Lab, and Hassabis wants Frontier Labs to publish model cards (short technical write-ups on a model), keep strong internal cybersecurity, vet key staff, and properly fund their own safety research.
The testing itself would start voluntary: Frontier Labs would hand their models to the Standards Body for review up to 30 days before public release. Once that process proves it works, Hassabis wants it made mandatory, meaning a Frontier Model could not legally launch in the US market without passing it first, and labs would keep working with the Body to patch dangerous flaws found after release. The tests would specifically probe for cybersecurity weaknesses and biological threat potential, and for agentic red flags: does the model try to get around its own guardrails (limits on what it's allowed to do), and does it show signs of deception. He also wants two transparency practices built in: watermarking AI-generated images so people can tell what's fake, and forcing models to produce their reasoning in human-readable tokens (text chunks, about three quarters of a word each) so outsiders can actually follow how a model reached an answer. The tests would be refreshed often, maybe quarterly at first, retiring benchmarks once they get stale or models start scoring perfectly on them. Early on, Frontier Labs would help design the tests. Over time, Hassabis wants the Body to build its own independent tests the labs never see in advance, so labs cannot just train their models to beat the test instead of actually being safe.
we’ve essentially found a way to make sand think.via @demishassabis on X →

